alessiodm/drl-zh: A Notebook Course Where You Write the RL Algorithms
Deep Reinforcement Learning: Zero to Hero!
At a glance
- What is it?
- drl-zh is an MIT-licensed deep reinforcement learning course built as Jupyter notebooks, with TODO stubs in the exercise track and complete code under solution/. It runs best from a Docker workspace that also ships a VS Code AI companion.
- Who is it for?
- Adopt drl-zh if you already write PyTorch and want to implement DQN, PPO, SAC, MCTS and Dreamer yourself rather than call a library. Skip it if you need a packaged training framework or you are not ready to debug your own training loops, since the exercise notebooks ship with TODO sections by design.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 127 days ago.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What drl-zh solves, and who it is actually for
Most reinforcement learning material splits into two unsatisfying halves. Papers and textbooks explain the math but leave you with nothing runnable. Libraries such as Stable-Baselines3 give you a trained agent in a few lines, which is useful at work and useless for understanding why the agent trains at all. drl-zh takes the third path: the root notebooks are an exercise track where the code is intentionally replaced with guided TODO sections, and the solution/ directory holds the complete runnable versions of the same notebooks.
The README is explicit about the intended reader: you should be comfortable with Python, PyTorch basics, and the usual math behind ML, meaning probability, statistics, linear algebra and derivatives. The notebooks teach the RL, but they assume you can read and modify real training code. That is a real filter. If you have never written a training loop by hand, the first TODO in the DQN notebook will be a wall rather than a lesson. If you have, the curriculum runs from MDPs and tabular RL through DQN, REINFORCE, actor-critic methods, DDPG, TD3, SAC and PPO, then into RND curiosity, multi-agent RL, offline RL with BC and IQL, Monte Carlo Tree Search and AlphaZero-style policy/value learning, RLHF with PPO, DPO and GRPO, Decision Transformers, a NanoVLA notebook, MBPO and Dreamer, and finally MAML and FOMAML.
The two-track notebook layout and the 00 to 18 curriculum
The repository is a flat set of numbered notebooks, 00_Intro.ipynb through 18_EOF.ipynb, plus assets/, cpu/, extension/, solution/ and util/ directories. The numbering is not decoration: the README states that the foundations are meant to be done in order, while the advanced notebooks are self-contained and the numbering gives a default path from exploration to the course capstone.
The tracks map cleanly onto notebook ranges. 00 to 07 covers foundations: MDPs, tabular RL, DQN, REINFORCE, actor-critic methods, DDPG, TD3, SAC and PPO. 08 to 10 breaks assumptions with RND curiosity, multi-agent RL and offline RL. 11 is planning: Monte Carlo Tree Search, self-play and AlphaZero-style policy/value learning. 12 and 13 are the modern AI stack, covering RLHF with PPO, DPO, GRPO, Decision Transformers and NanoVLA. 14 is production material: TensorBoard, checkpointing, debugging, multiple seeds, Ray and Optuna. 15 and 16 are world models, MBPO with SAC followed by DR3AM/Dreamer with RSSM latent imagination. 17 and 18 close with MAML, FOMAML and fast adaptation.
The two-track split is the design decision worth noting. Because the root notebook keeps the TODO stubs while solution/ holds a complete version, you can unblock yourself without leaving the course, but it also means you can finish the whole curriculum having read a lot of working code and written very little. Nothing in the repository forces you to attempt the TODO before opening the answer. That is a discipline you have to supply.
Installing drl-zh with Docker and opening the first notebook
The README recommends Docker because it produces one reproducible workspace containing code-server, the notebooks, a Python interpreter pinned to >=3.13,<3.14, the Jupyter kernel, dependencies and the AI Companion. Install Docker and Git first, then clone the repository and cd into it. On Linux and macOS the README gives this step so that files created inside the container end up owned by your host user:
printf "UID=$(id -u)\nGID=$(id -g)\n" > .envThose values feed the UID and GID build arguments declared in docker-compose.yml. With the .env file in place, start the default environment:
docker compose up --build -dThis builds the Dockerfile, which defaults to BASE_IMAGE=nvidia/cuda:12.9.1-runtime-ubuntu24.04 with TORCH_VARIANT=cuda. Once the container is up, open http://localhost:8080 in a Chromium-based browser, select the Python (drl-zh) kernel, then open 00_Intro.ipynb and start filling in TODOs. The compose file maps both 8080 and 8888 to the host, so the Jupyter port is reachable too.
If you have an NVIDIA GPU, the README gives this variant, which layers the GPU override on top of the base compose file:
docker compose -f docker-compose.yml -f docker-compose.gpu.yml up --build -dFor a smaller CPU-only image, the equivalent command uses the CPU override, which the Dockerfile comments describe as roughly 2 GB smaller because torch comes from the pytorch-cpu index:
docker compose -f docker-compose.yml -f docker-compose.cpu.yml up --build -dIf you would rather not use Docker, MANUAL.md is the documented path for a native setup with Python, Poetry, VS Code and the Companion.
The DRL-ZH AI Companion, and why it needs your own API key
The Docker workspace includes a VS Code extension built specifically for this course. According to the README, it knows which notebook and which TODO you are working on, offers Socratic hints instead of spoilers, and supports text or voice mode. It is compiled during the image build in a separate Node stage: the Dockerfile runs npm ci against extension/package.json and extension/package-lock.json, then npm run build, npm prune --omit=dev and npx vsce package to produce companion.vsix. The pruning step is deliberate, because the extension is not bundled and its runtime dependencies must ship inside the vsix.
The constraint that matters is bring your own LLM key. Gemini is the default provider, with OpenAI, Anthropic and Groq also supported. There is no hosted inference in the box, so the companion is not a free tutor and the course does not work offline in that mode. The extension source lives in extension/ if you want to check what it sends before you point it at a key.
Version pinning: Python 3.13, ale-py and the pygame fork
pyproject.toml pins python = ">=3.13,<3.14", and the accompanying comment explains why: ale-py has no wheels for 3.14 yet, while every other dependency in the file supports it. That single constraint propagates into the container, so if you build a native environment on a newer interpreter you will be fighting the resolver rather than the notebooks.
Two other choices in the dependency list are worth reading closely. The project depends on pygame-ce rather than pygame, with a comment noting that pettingzoo 1.26 and gymnasium 1.3 both moved to the community fork, and that it installs into the same pygame import namespace. And MPE environments were split out of PettingZoo 1.26 into their own package, which is why mpe2 appears separately alongside pettingzoo. The gymnasium dependency pulls the atari, box2d, mujoco and other extras, and the Dockerfile installs build-essential and swig because gymnasium[box2d] builds Box2D from source. Expect a long first image build; the compose file mounts a poetry-cache volume so that cost is not paid twice.
Where drl-zh is the wrong tool
This is a course, not a training framework. There is no published API, no model registry and no inference server. If your goal is to get a working PPO agent on a custom environment this week, adopting drl-zh means writing the agent yourself first, which is the opposite of what you want. A library built for that job will beat it on every axis except understanding.
The hardware story is also narrower than the notebook list suggests. The default compose file requests shm_size: '2gb' and builds from a CUDA base image, so the default path assumes a GPU-capable host. The CPU override exists and the Dockerfile describes it as lean, but a CPU-only image will not make the world-model notebooks or the RLHF notebooks pleasant to run. Nothing in the repository gives timings, so treat any expectation about how long a notebook takes to train as unverified.
Finally, the AI Companion is not self-contained. It requires an external LLM key, and the README does not document how it behaves if the key is missing or the provider is unreachable. If you are working in an environment where sending notebook context to a third-party model is not acceptable, the companion is dead weight and you are left with the notebooks alone, which is still a complete course.
drl-zh versus a maintained RL library
The honest alternative is a production RL library such as Stable-Baselines3 or the RLlib side of Ray. The difference is not quality, it is what the artifact is. A library exposes algorithms as objects you configure and call: you pass an environment and hyperparameters, and the library owns the training loop, the replay buffer, the update schedule and the checkpointing. drl-zh does the reverse. It hands you the loop and asks you to write the update rule, and notebook 14 is where the production concerns (TensorBoard, checkpointing, debugging, multiple seeds, Ray, Optuna) show up, near the end of the curriculum rather than at the start.
The practical consequence is that drl-zh teaches you what those libraries are doing inside, which makes you a better user of them, but it will not replace them on a deadline. Ray and Optuna both appear in pyproject.toml as dependencies, so the course does touch the surrounding tooling, and the production notebook is the place to look if you want to see how the two worlds connect.
Licence, maintenance and what upgrading costs you
drl-zh is MIT licensed, and the LICENSE file is at the repository root. MIT permits commercial use, modification and redistribution provided the copyright notice and permission notice are retained. That matters if you plan to lift a solution notebook into a work project: the licence does not stop you, but you should keep the notice with the code. This is not legal advice, and if you are distributing modified versions inside a company, have someone check the notice requirements.
The repository is not archived, and the last push was on 2026-05-26. Releases are tagged rather than continuous: v0.1.0 for classic methods and v1.0.0 for the second edition with advanced topics, both dated 2025-08-28, and v1.5.0 for the third edition, modern techniques, dated 2026-05-26. Note that pyproject.toml declares version 2.0.0 while the newest release tag is v1.5.0, so the package version and the release tag do not line up.
Upgrade cost is dominated by the Python pin. Because the project holds Python at 3.13 until ale-py ships 3.14 wheels, moving forward means either waiting on that dependency or editing pyproject.toml and accepting that the Atari environments may not build. Docker rebuilds are the cheaper path: the compose file deliberately does not volume-mount the extensions directory, so a rebuild picks up fresh extension versions instead of a stale volume.
Editorial conclusion
Adopt drl-zh if you already write PyTorch and want to implement DQN, PPO, SAC, MCTS and Dreamer yourself rather than call a library. Skip it if you need a packaged training framework or you are not ready to debug your own training loops, since the exercise notebooks ship with TODO sections by design. Before committing, check that Docker and Git are installed, confirm your Python version matches the >=3.13,<3.14 range pinned in pyproject.toml, and open solution/ to see how much of each notebook is left for you to write.
Frequently asked questions
Can you give me an example of Deep reinforcement learning in the drl-zh course?
The curriculum lists DQN in the foundations track, notebooks 00 to 07, alongside REINFORCE, actor-critic methods, DDPG, TD3, SAC and PPO. Later tracks cover AlphaZero-style policy/value learning with Monte Carlo Tree Search and RLHF with PPO, DPO and GRPO.
Is reinforcement learning a dead end, given what drl-zh covers?
The repository does not argue the point either way. The curriculum extends past classic methods into Decision Transformers, NanoVLA policies, world models with Dreamer and meta-learning with MAML, which is where the project places its third edition.
Is reinforcement learning considered AI in drl-zh?
The repository does not define the term. It is tagged deep-learning, machine-learning, reinforcement-learning and deep-reinforcement-learning, and the README describes it as a hands-on deep reinforcement learning course.
Which is better, deep learning or reinforcement learning, for someone starting drl-zh?
drl-zh does not compare them. Its prerequisites ask for PyTorch basics plus probability, statistics, linear algebra and derivatives, so it treats deep learning as prior knowledge and spends the notebooks on the RL side.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/alessiodm-drl-zh)