# rainbow-is-all-you-need: DQN to Rainbow in nine .py files

> This tutorial rebuilds DQN and each of the components that make up Rainbow as nine marimo notebooks that are plain Python files at the repository root, so a git diff between chapter 3 and chapter 4 shows real code. It runs on gymnasium classic-control environments, not the Atari benchmarks the papers report, and the environment setup is mise plus uv with two make targets.

**Curt-Park/rainbow-is-all-you-need** — Rainbow is all you need! A step-by-step tutorial from DQN to Rainbow

- Repository: https://github.com/Curt-Park/rainbow-is-all-you-need
- Stars: 2,035 · Forks: 353
- Language: Python
- License: MIT
- Published: 2026-09-30 · Updated: 2026-09-30 · Language: en
- Canonical page: https://hysenlabs.com/projects/curt-park-rainbow-is-all-you-need

## Nine .py files at the root, which is the argument for marimo

The repository listing is short enough to read at a glance, and the layout is the tutorial's central claim.

There are nine files at the root, one per chapter: 01_dqn.py, 02_double_q.py, 03_per.py, 04_dueling.py, 05_noisy_net.py, 06_categorical_dqn.py, 07_n_step_learning.py, 08_rainbow.py and 09_rainbow_iqn.py. There is also a __marimo__ directory, which is marimo's per-notebook metadata. There is no src/, no package, no tests directory and no notebook subfolder.

Those nine files are marimo notebooks, and the README says what that means: built with marimo, a reactive Python notebook that runs as a pure .py file with better reproducibility, git diffs, and interactive UI.

The git diff claim is the one that matters most, and it is the reason to care about the choice at all. The standard objection to teaching with notebooks is that they are opaque. A .ipynb file is a JSON container whose cell boundaries are arbitrary, whose outputs are stored alongside the code, and whose diffs are unreadable. Two students who run the same tutorial can end up with files that differ by megabytes of base64 image output and are semantically identical.

Here the notebook is a Python module. A cell is a decorated function. The file imports, it type-checks, your editor's go-to-definition works, and when the author changes the replay buffer in chapter 3, the diff for chapter 4 is a diff of Python. For a tutorial whose entire pedagogy is incremental, that is not a convenience. It is the mechanism that makes incremental legible.

The reactivity claim is the other half. marimo notebooks re-run cells in dependency order rather than top to bottom, which means state is declared rather than positional. In a top-to-bottom notebook, editing a cell above another one silently invalidates everything below it and you find out by re-running. Here the dependency graph is explicit, so a change propagates to exactly what depends on it.

Each chapter also has a Preview link to molab at molab.marimo.io, which is marimo's hosted notebook service, and the README says you can click Run in molab on the preview page to open an interactive session where you can edit and run the notebook. So there is a zero-setup path for a reader who only wants to read and poke at it, and a clone-and-run path for a reader who wants to change things. That pairing is the right default for a tutorial: try it in the browser, fork it locally if you care.

The consequence for an evaluator is worth stating plainly. You are not adopting a library. You are reading nine files, in order, each of which is the previous one plus one idea. Whether that is what you want is the entire adoption decision.

## mise, uv, ruff and marimo, behind two make targets

The environment setup is four tools deep and it is worth naming, because the choices are recent and the payoff is that a tutorial from 2018 and a tutorial from last month install the same way.

The Prerequisites section is:

```bash
# Install mise
curl https://mise.run | sh

# Clone the project
git clone https://github.com/Curt-Park/rainbow-is-all-you-need.git
cd rainbow-is-all-you-need
```

After the clone, two make targets do the rest. And the Makefile is six targets long:

```makefile
init:
	mise trust && mise install

setup:
	uv sync

run:
	marimo edit $(notebook)

lint:
	uv run ruff check .

format:
	uv run ruff format .
```

Each target pulls its weight. mise handles the Python interpreter version, driven by a .mise.toml at the repository root, and mise trust is the step where you authorise mise to execute the tool configuration in a cloned repository, which is a real security mechanism rather than a formality. uv sync installs the dependencies into a project environment from uv.lock, which is committed. So the interpreter version and the exact dependency versions are both pinned by files in the repository rather than by whatever the reader's machine happens to have.

That is a meaningful change from the tutorial norm. The traditional pattern is a requirements.txt with minimum versions, a note to create a virtualenv, and a paragraph about installing PyTorch correctly for your CUDA version, which is the single most common reason a reinforcement learning tutorial does not run for a given reader. Here the torch floor is declared as torch>=2.7.0 and the resolver handles the rest, and uv.lock means two readers on two machines get the same environment.

run takes a notebook variable, so the documented invocation is make run notebook=01_dqn.py, which expands to marimo edit 01_dqn.py. Every chapter is therefore one command with a different filename, and there is nothing to remember between them.

The last target is the one that would worry a cautious reader. clean runs git clean -xdf, which removes every untracked and ignored file in the working tree. For a repository that is nine notebooks and no build output, that is a reasonable reset. It is also the kind of target that should not be copied into a template, because in a project with a build directory or a local configuration it will delete work.

The four tools are mise, uv, ruff and marimo. Three of them did not exist when the papers this tutorial reproduces were published, and the setup is now two make targets long. That is the strongest argument in the repository for why tutorial rot is avoidable.

## Classic control, so the papers' numbers are not reproducible here

The dependency list settles what kind of tutorial this is, and one entry settles it.

```toml
dependencies = [
    "gymnasium[classic-control]>=0.28.1",
    "marimo>=0.21.0",
    "matplotlib>=3.5.1",
    "moviepy>=1.0.3",
    "numpy",
    "torch>=2.7.0",
]
```

gymnasium with the classic-control extra. That is CartPole, MountainCar, Acrobot and the rest of the classical control suite. The original DQN paper, the Nature paper by Mnih and colleagues that is reference 01 in this repository, reports on forty-nine Atari games. The Rainbow paper, reference 08, reports on Atari too, along with the ALE arcade learning environment. The Implicit Quantile Networks paper, reference 09, reports on Atari again.

So the environments here are not the environments in the papers, and no number you get from these notebooks is comparable to a published result. That is not a flaw in the tutorial, it is the trade that makes the tutorial possible. Training DQN to play an Atari game competently takes hours of GPU time and produces a reward curve nobody can read in an afternoon. Training DQN to solve CartPole takes a couple of minutes on a laptop, and the reward curve goes up on screen.

The point of a component-by-component tutorial is that you can see what one change did. Prioritised experience replay on CartPole visibly changes the sample efficiency. A dueling head visibly changes the value estimates. Noisy nets visibly replace an epsilon-greedy exploration schedule. Those effects are the lesson, and they are legible at the scale of a short notebook run. The cost is that you are not measuring anything publishable, and the README does not pretend otherwise.

The other two dependencies explain what you see. matplotlib is for the learning curves, and moviepy is for rendering video of the agent, which is the payoff for a classic-control environment specifically: you can watch a pole balance or fail. A distributional RL implementation visualised as a return distribution is much easier to grasp when the underlying task is something you already understand.

One small inconsistency is worth noting. Every other dependency has a floor version and numpy does not. numpy is the one package in this list where a major version bump is most likely to matter, and it is the one with no constraint. Given that uv.lock is committed, the actual environment is reproducible regardless, so the omission affects fresh resolution rather than locked installs.

The companion project is the other half of the framing. If you want policy gradient methods instead, the README points at PG is All You Need, written by the same second contributor who appears in the contributors table. So the two repositories between them cover the two families of deep reinforcement learning that a practitioner actually reaches for first.

## Chapter 09 is IQN, which is not in Rainbow

The title is accurate for eight of the nine chapters, and the ninth is the interesting one.

The contents list runs DQN, DoubleDQN, PrioritizedExperienceReplay, DuelingNet, NoisyNet, CategoricalDQN, N-stepLearning, Rainbow, and then Rainbow IQN. Compare that with the Rainbow paper's own list of components, which is double Q-learning, prioritised experience replay, dueling networks, multi-step learning, distributional learning and noisy networks. Six components, and chapters 2 through 7 are exactly those six in a sensible teaching order.

Chapter 8 is the combination, which is what the paper calls Rainbow: all six together. That is the point of the repository and it lands where the title says it should.

Chapter 9 is not part of that. Rainbow IQN is Implicit Quantile Networks, reference 09, W. Dabney and colleagues, arXiv 1806.06923, 2018. It is a different way to do distributional learning: instead of projecting a fixed set of categorical atoms as CategoricalDQN does, IQN infers quantile locations from the state through learned cosine embeddings, so the number of quantiles is decoupled from the network architecture.

Placing it after the combination is a defensible editorial decision and an informative one. A reader who finishes chapter 8 has the paper. A reader who finishes chapter 9 has the paper plus the most direct successor to its distributional component, with the machinery already built, which means the difference between the two implementations of distributional RL is the only thing left to understand. That is a much better way to learn it than starting from an IQN paper with no C51 baseline in your head.

The reference list is the other thing worth reading, because it is exactly nine entries for exactly nine chapters, which is a discipline most tutorial repositories do not maintain. The mapping is direct. Chapter 1 is the Nature DQN paper, chapter 2 is van Hasselt on double Q-learning, chapter 3 is Schaul on prioritised replay, chapter 4 is Wang on dueling networks, chapter 5 is Fortunato on noisy networks, chapter 6 is Bellemare on the distributional perspective, chapter 7 is Schaul et al. on multi-step returns.

And chapter 7's reference is not from 2017. It is R. S. Sutton, Learning to predict by the methods of temporal differences, Machine Learning 3(1):9-44, 1988, linked to the erratum PDF hosted on incompleteideas.net. So a tutorial whose subject is a 2017 combination of six tricks from 2015 to 2017 traces one of them back to the 1988 paper that invented the update rule the whole family is built on. That single citation does more for a reader's grounding than several paragraphs of prose would.

## The ruff config exempts l and u, and the reason is in a comment

The lint configuration is four lines longer than a default because someone wrote down why each exception exists, and one of the reasons is a good story.

```toml
[tool.ruff]
line-length = 100
target-version = "py312"

[tool.ruff.lint]
select = [
    "E",   # pycodestyle errors
    "F",   # pyflakes
    "I",   # isort
    "W",   # pycodestyle warnings
]
ignore = [
    "E501",  # line too long (handled by formatter for code; prose in mo.md() is exempt)
    "E741",  # ambiguous variable name (l/u are standard notation from the Categorical DQN paper)
]

[tool.ruff.lint.isort]
known-first-party = ["segment_tree"]
```

E741 is the ambiguous-variable-name rule, and it fires on single-letter names that read badly: l for look-alike-one, I for look-alike-capital-i, O for look-alike-zero. The project's exemption says l and u are standard notation from the Categorical DQN paper.

That is the correct reason and it is the right way to record it. In a distributional reinforcement learning implementation, l is the return sample and u is the atom location, and those names come from the paper's own formulation. Renaming them to something unambiguous would make the code diverge from the source it is teaching, and a reader comparing the two side by side would have to hold a translation table in their head. So the linter is told about the domain instead of being switched off wholesale.

That is a small piece of engineering taste worth noticing. The alternative, disabling E741 for the file or the project, would have hidden a real class of naming problem. The exemption is narrow, commented, and justified by reference to the paper being implemented.

The E501 exemption has the same shape. Line-too-long is normally deferred to the formatter, which is standard practice, and the comment adds the detail that prose inside mo.md() calls is exempt, because a markdown call carries sentences and cannot be wrapped at 100 columns without destroying the text. So the exception is scoped by cause rather than by rule.

The select list is E, F, I and W, which is pycodestyle errors, pyflakes, isort and pycodestyle warnings. Notably absent is the more opinionated set. There is no complexity check, no bugbear, no annotation-typing requirement. For nine teaching files where the clarity of the algorithm matters more than the abstraction discipline of a library, that is the right scope.

The known-first-party entry is the other detail. segment_tree is declared as first-party for import sorting, which means the repository has a first-party module by that name. Given that chapter 3 is Prioritised Experience Replay, a segment tree is the data structure PER needs, and putting it in a first-party module rather than inline in the notebook is the one place this project has done what a library would do.

## The topic tags still say colab-notebook and nbviewer

The repository topics are android-free, uncontroversial except for two entries that describe a tool this project no longer uses.

The topics are colab-notebook, dqn, gym-environment, nbviewer, pytorch, rainbow, and reinforcement-learning.

colab-notebook and nbviewer both describe the Google Colab and nbviewer ecosystem. The README's build section says the notebooks are built with marimo, that the per-notebook metadata lives in a __marimo__ directory, that the run target is marimo edit, and that every chapter has a Preview link to molab.marimo.io. There is no Jupyter notebook in the repository, no .ipynb file, and no Colab badge in the README.

So the project migrated from one notebook platform to another and the topic tags were not updated. The practical effect is small but real: someone searching GitHub by topic for colab-notebook will land here, expect a Colab link, and find a marimo preview instead. And someone searching for nbviewer expecting a rendered notebook view will find neither.

It is worth mentioning because topics are one of the few pieces of metadata a project sets once and then forgets, and because this is otherwise a very well-kept repository. The lint config has a comment justifying each exception, the environment is pinned with a lockfile, the reference list matches the chapter list one to one, and the README points at a sibling tutorial for the algorithm family this one does not cover. The tags are the one place the housekeeping lapsed.

Two other pieces of repository metadata are worth reading together. There are no GitHub releases, and pyproject.toml declares version 0.1.0. So there is no version history, no changelog and no tag to pin, and the version number in the manifest has never been tied to a release artefact. For a repository whose deliverable is nine files you read rather than a package you install, that is defensible. The name, the description and the readme field in the manifest all describe a tutorial, and the dependencies are pinned in uv.lock, so a reader who clones today gets the environment the author tested.

There is also no declared homepage. The canonical locations are the repository itself and the molab preview links, one per chapter, which is a reasonable distribution model for a tutorial and means the URLs people share are chapter-specific rather than project-level.

The last push was on 2026-06-20. That is recent enough that the project is not dormant, and it is worth noting that the maintenance activity here is the kind that produces a commit rather than a release, which is what a tutorial repository should look like when it is healthy.

## What a diff-driven tutorial gives you, and what it does not

The honest case for this repository is a claim about how incremental explanations should be stored, and it is a stronger claim than it first looks.

A conventional tutorial has a structural problem. Chapter 4 of a conventional tutorial is a self-contained document that re-explains chapters 1 to 3 in abbreviated form, and the abbreviation is where errors accumulate. When the author fixes a bug in the DQN implementation, they have to find every place it was restated, and they will miss some. When a reader works through the chapters in order, they have already internalised the earlier version, so the inconsistency shows up as confusion rather than as a bug report.

Storing each chapter as a complete Python file in one repository, where each is the previous one plus one idea, changes the failure mode. The duplication is not eliminated, it is made visible. A git diff between 03_per.py and 04_dueling.py shows exactly what the dueling head changed. If the author later fixes the replay buffer, the diff for every later chapter shows the fix flowing forward, and a reviewer can see whether they propagated.

That is why the marimo choice matters beyond the file extension. The reactivity claim means a cell can be a function with a name, the dependency graph is real rather than positional, and the file behaves like a module. A notebook format where the file is the artefact and the file is valid Python is the precondition for the diff-driven model. Everything else follows.

The nine-file limit is the discipline. There is no chapter 10 yet, no appendix of exercises, no dataset loader abstraction, no evaluation harness. Each file is one algorithm, and the repository is the set.

What the model does not give you is a maintained library. There are no releases, the version is 0.1.0, and nobody is going to fix a bug in 06_categorical_dqn.py for you. If you want to use a rainbow implementation in your own work, this is not it. If you want to understand what each of the six components does and how they combine, nine files that are diffs against each other, runnable in a browser without installing anything, is a better resource than a paper and a code repository separately.

The one thing a reader should do is start at 01_dqn.py and read forward as diffs, rather than reading each file as a standalone program. The tutorial is not nine independent implementations. It is one implementation, nine times, with one thing added each time, and that is the entire pedagogical claim.

## Conclusion

Use rainbow-is-all-you-need if you want to understand what each of Rainbow's seven components actually does, in isolation, in an order where each chapter is a diff away from the last, and if you are the kind of reader for whom CartPole learning in two minutes is a feature rather than a limitation. Do not use it to reproduce the paper results, because the environments are gymnasium classic-control rather than Atari, and no headline number here is comparable to Hessel et al. Do not expect a maintained library, because there are no releases, the pyproject version is 0.1.0, and the artefact is nine files you read rather than a package you install. Verify four things. Confirm your Python is 3.12 or newer, since requires-python says so and the ruff target-version is py312. Install mise and uv first, because make init delegates to mise trust and mise install and make setup delegates to uv sync, so neither target works with a hand-rolled virtualenv. Read 01_dqn.py before anything else, since the whole series is a series of diffs against it. And note that chapter 9 is Implicit Quantile Networks, which is not part of the original Rainbow combination, so the title describes chapters 1 to 8 rather than the repository. The deciding fact is that the notebooks are source files, which makes a tutorial that has resisted the usual notebook rot a durable artefact rather than a set of opaque blobs.

## FAQ

### What is rainbow-is-all-you-need and what does it cover?

It is a step-by-step tutorial from DQN to Rainbow, with each chapter containing theoretical background and an object-oriented implementation. The nine chapters are DQN, DoubleDQN, PrioritizedExperienceReplay, DuelingNet, NoisyNet, CategoricalDQN, N-stepLearning, Rainbow, and Rainbow IQN, each as a single marimo notebook stored as a .py file.

### How do I run the rainbow-is-all-you-need notebooks?

Install mise, clone the repository, then run make init and make setup. make init runs mise trust and mise install, and make setup runs uv sync against the committed uv.lock. Then run a chapter with make run notebook=01_dqn.py, which expands to marimo edit on that file. You can also open any chapter in a browser through its molab preview link without installing anything.

### Why are the notebooks .py files rather than Jupyter notebooks?

The project is built with marimo, which the README describes as a reactive Python notebook that runs as a pure .py file with better reproducibility, git diffs and interactive UI. The practical effect is that each chapter is a valid Python module, so the difference between one chapter and the next is a readable git diff, and a cell re-runs based on a declared dependency graph rather than on position.

### Will running these notebooks reproduce the results in the Rainbow paper?

No. The dependency is gymnasium with the classic-control extra, so the environments are CartPole, MountainCar and the rest of the classical control suite. The DQN, Rainbow and IQN papers all report on Atari, so no number from these notebooks is comparable to a published result. The trade is deliberate: the effect of each individual component is visible in a short run rather than after hours of training.

### What licence is rainbow-is-all-you-need under?

MIT, with a LICENSE file at the repository root. The project requires Python 3.12 or newer, has no GitHub releases, and declares version 0.1.0 in pyproject.toml, so there is no release history to pin. The environment itself is reproducible because uv.lock is committed.

### Is chapter 09 part of Rainbow?

No. Chapters 2 through 7 are the six components the Rainbow paper combines, and chapter 8 is the combination itself. Chapter 9 is Rainbow IQN, Implicit Quantile Networks by Dabney and colleagues, arXiv 1806.06923, which is a later alternative to the CategoricalDQN approach in chapter 6. It is placed after the combination so the difference between the two distributional methods is the only thing left to learn.

## Sources

- [Curt-Park/rainbow-is-all-you-need on GitHub](https://github.com/Curt-Park/rainbow-is-all-you-need)
- [Issues](https://github.com/Curt-Park/rainbow-is-all-you-need/issues)
- [License: MIT](https://github.com/Curt-Park/rainbow-is-all-you-need/blob/master/LICENSE)
- [README](https://github.com/Curt-Park/rainbow-is-all-you-need/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/curt-park-rainbow-is-all-you-need
