# hands-on-modern-rl: a VitePress-based RL curriculum that runs its first experiment in a browser

> WalkingLabs' open textbook walks from Markov decision processes to PPO, DPO, GRPO and agentic RL. The chapters are VitePress markdown, the experiments run on ModelScope Studios, and the README's own note says the course was created with AI assistance and has not been fully reviewed.

**walkinglabs/hands-on-modern-rl** — 🚀 An open-source, hands-on curriculum bridging the gap from basic RL concepts to LLM alignment, RLVR, and advanced Agentic systems. 

- Repository: https://github.com/walkinglabs/hands-on-modern-rl
- Website: https://walkinglabs.github.io/hands-on-modern-rl/
- Stars: 4,481 · Forks: 328
- Language: Python
- License: NOASSERTION
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/walkinglabs-hands-on-modern-rl

## The gap hands-on-modern-rl targets: from MDP notation to RLHF code

Most reinforcement learning material splits into two camps that do not talk to each other. Classic textbooks derive policy gradients and value functions on gridworlds, then stop. LLM alignment write-ups start from a pretrained model and a preference dataset, assuming the reader already knows what an advantage estimate is. The repository positions itself in between: its subtitle runs from Markov decision processes and policy optimization to reasoning models, agents, and multimodal systems, and the topic list covers ppo, grpo, dpo, rlhf, sft, agentic-rl and vlm work.

The intended reader is a Python developer who can read PyTorch and wants the algorithmic reason behind each training loop, not a survey. That is a narrower audience than "anyone curious about AI". If you have never written a training loop, the early chapters will be the useful part and the alignment chapters will read as recipes. If you already ship PPO implementations, the value is concentrated in the later agentic and multimodal sections, which are the ones the news entries describe as most recently added.

## How the book is assembled: VitePress, a code tree, and ModelScope Studios

The repository is a documentation site, not a Python package. package.json names the project hands-on-modern-rl and declares type: module, with VitePress as the docs engine and Node >=18 as the engine constraint. The dev and build scripts chain asset optimization and sitemap generation in front of vitepress dev docs and vitepress build docs, so the site is generated from markdown under docs/.

Experiment code lives apart from the prose, in code/ and notebooks/, with a code/online-experiments directory holding companion notebooks. Each online experiment pairs a notebook with a ModelScope Studio: the notebook imports the same training runtime as the Studio, exposes the experiment parameters, prints the training log, plots checkpoint evaluations, and shows the learned-policy replay. The published table lists CartPole PPO, a Gymnasium playground, ViZDoom, Atari/ALE, board games and self-play, and multi-agent games; only the Atari entry is marked xGPU, the rest CPU.

There is also a maintenance layer that most tutorial repositories lack. scripts/ contains verify.sh, check-bilingual-sync.mjs, check-english-quality.mjs, generate-sitemap.mjs and fix-heading-numbers.mjs, and package.json wires them to npm run verify, npm run verify:bilingual, npm run verify:english and npm run fix:headings. AGENTS.md at the repository root and the .zcode/ directory indicate the content pipeline is meant to be operated by automated agents as well as humans.

## Installing hands-on-modern-rl and running the first experiment

The README documents a Node toolchain rather than a pip install, because the primary artifact is the site. Node 18 or newer is required. From a clone of the repository, install dependencies and start the dev server:

```bash
npm install
npm run dev
```

The dev script first runs assets:optimize, then vitepress dev docs. VitePress prints a local URL; opening it shows the book with working navigation. If asset optimization is slow or fails on your machine, npm run dev:fast skips it and runs vitepress dev docs directly.

To produce the static site the way CI does, run the build script, which also generates the sitemap:

```bash
npm run build
npm run preview
```

npm run preview serves the built output so you can check that internal links resolve before deploying.

If you would rather not set up anything locally, the README points at ModelScope for the classic experiments. The CartPole PPO Studio bundles the interface, runtime and training entry point on one page, and its companion notebook under code/online-experiments imports the same runtime. Running that notebook should print the training log and then plot checkpoint evaluations, which is the quickest way to see whether the code matches the chapter you just read. The README also links PDF builds for both the Chinese and English editions, produced automatically through CI, and the release history shows v0.2.1 as the latest tagged textbook version.

## The AI-assistance disclaimer is the limitation that matters most

The README carries a note above the news list stating that the course was created with AI assistance and has not yet been fully reviewed, and that it may contain factual mistakes or code that does not run as expected. That is unusually direct, and it should set your expectations for every derivation in the book. Treat a chapter as a well-organized starting point that cites its sources, not as a checked reference.

The practical consequence is that the verification scripts are not optional garnish. npm run verify:english runs an English-quality check, npm run verify:bilingual checks that the Chinese and English editions have not drifted apart, and npm run verify runs the umbrella shell script. A bilingual textbook maintained by two language trees will drift; the existence of check-bilingual-sync.mjs is evidence the maintainers know this, not proof the drift is currently zero.

The second limitation is the runtime story. Local setup gives you the book, not the training environments. To run experiments you either use ModelScope Studios, which places the compute and the interface on a third-party platform, or you reconstruct the environment from code/ yourself. The README does not document a local environment file, a container image, or a pinned dependency lock for the experiment code, so offline reproduction of the training runs is not something the documentation promises.

A third constraint is scope creep. The repository spans classic control, board games, multi-agent games, LLM alignment, RLVR and multimodal reasoning. Breadth of that kind usually means some sections are sketches. The announcement at the top of the README says a new version is coming and that many sections are still being organized, which is consistent with that reading.

## How it compares with Spinning Up and the Hugging Face alignment stack

OpenAI Spinning Up is the closest classic-RL comparison. It is a small, fixed set of algorithm implementations with a single consistent code style, and it stops before language models. hands-on-modern-rl covers more ground and is structured as a book with a site, a PDF pipeline and online notebooks, but it does not ship one canonical implementation per algorithm in the way Spinning Up does. If you want to read one clean PPO and one clean SAC, Spinning Up is the tighter artifact. If you want the line from PPO to GRPO to DPO to agentic training, this repository is the one that draws it.

The other comparison is the Hugging Face alignment ecosystem, where TRL provides the training loops and the documentation is API reference plus example scripts. The difference in approach is that TRL is a library you import, while hands-on-modern-rl is a curriculum you read, with the runnable pieces hosted on ModelScope. A team that needs to fine-tune a model this week should reach for a library. A team that needs its engineers to understand why the KL term is there should read the chapters first. The two are complementary, and the repository's own topic list (rlhf, dpo, grpo) maps onto the same algorithm names the libraries implement.

## Licence and the cost of keeping a bilingual textbook current

The licence situation needs care. The repository metadata reports NOASSERTION, while package.json declares "license": "CC-BY-NC-SA-4.0" and the README badge links to CC BY-NC-SA 4.0. The non-commercial clause is the part that decides adoption: a company cannot lift chapters into internal paid training material without checking what NC means for its use, and ShareAlike affects derivative works. This is a description of the declared terms, not legal advice; read the LICENSE file in the repository before you reuse content.

Maintenance cost falls on whoever adopts the content, not on whoever reads it. The last push was on 2026-09-03, and the most recent tagged release, v0.2.1, dates from 2026-06-18. The upgrade path for a reader is trivial: pull the branch and rerun npm install, since the site is generated from markdown. The upgrade path for a contributor is heavier. Adding a chapter means keeping the Chinese and English trees in sync, satisfying check-english-quality.mjs, and regenerating the sitemap and PDF builds through scripts/build-pdf.mjs and scripts/build-epub.mjs. The verify:bilingual script exists precisely because that sync is easy to break.

## Conclusion

Adopt it if you want a readable path from MDPs to RLHF, DPO, GRPO and agentic RL, and you are willing to verify each chapter against the paper it cites. Skip it if you need a reviewed, citable reference or a benchmarked library; the README states the course was created with AI assistance and has not been fully reviewed. Before you rely on it, run npm run verify and npm run verify:bilingual against the checkout, open the CartPole Studio to confirm the online path works for your region, and read the LICENSE file, since the repository metadata says NOASSERTION while package.json declares CC-BY-NC-SA-4.0.

## FAQ

### Is reinforcement learning a dead end?

The repository is built on the opposite premise: its topic list and chapter scope run from classic policy optimization through RLHF, DPO, GRPO and agentic RL, and it devotes recent releases to LLM alignment and agentic training systems. It presents RL as the mechanism behind model alignment rather than a closed research area.

### Do AI agents use reinforcement learning?

According to the README, the course added reproducible training examples for Agentic RL, including a Deep Research / rLLM example and complete code and fine-tuning analysis for building an agentic training system from scratch. That places agent training inside the RL curriculum rather than beside it.

### What is the most popular reinforcement learning algorithm?

The repository does not rank algorithms by popularity and gives no usage data, so no ranking can be drawn from it. What it does show is which algorithms it teaches: the topics list ppo, grpo, dpo, rlhf and sft, and the online experiment table starts with CartPole PPO.

### What is a real life example of reinforcement?

The repository's own runnable examples are the concrete cases it offers: CartPole PPO, a Gymnasium playground, ViZDoom, Atari/ALE, board games with self-play, and multi-agent games. The README notes that only the Atari entry is marked xGPU; the others run on CPU.

## Sources

- [Issues](https://github.com/walkinglabs/hands-on-modern-rl/issues)
- [Project website](https://walkinglabs.github.io/hands-on-modern-rl/)
- [README](https://github.com/walkinglabs/hands-on-modern-rl/blob/main/README.md)
- [Releases](https://github.com/walkinglabs/hands-on-modern-rl/releases)
- [walkinglabs/hands-on-modern-rl on GitHub](https://github.com/walkinglabs/hands-on-modern-rl)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/walkinglabs-hands-on-modern-rl
