# rlcode/reinforcement-learning: one Python file per algorithm, from policy iteration to PPO

> A minimal PyTorch rewrite of a 2017 teaching repository. Nine single-file algorithms, a uv-based setup, and published training numbers on Apple silicon. Useful for reading, less so for anything you intend to scale.

**rlcode/reinforcement-learning** — Minimal and Clean Reinforcement Learning Examples

- Repository: https://github.com/rlcode/reinforcement-learning
- Stars: 3,668 · Forks: 734
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/rlcode-reinforcement-learning

## What rlcode/reinforcement-learning is for

The repository describes itself as "easy-to-read code examples" with "one file for each algorithm". That framing is accurate and it also sets the ceiling. This is not a library you import into a larger system; it is a directory of scripts you read and run. The README lists nine algorithms across three directories: six in 1-grid-world (policy iteration, value iteration, SARSA, Q-learning, deep SARSA, REINFORCE), three in 2-cartpole (DQN, A2C, PPO), and two more in 3-atari (DQN and PPO), with a fourth directory, 4-atari-hard, holding PPO with RND, Go-Explore and a robustification script. The audience is someone who has read about temporal-difference learning and wants to see the update rule expressed as Python rather than as a paragraph. The README notes that each algorithm file now opens with a paper citation and the core update equation, which is the single most useful thing in the repository for that reader. If you need a maintained training stack with checkpointing conventions, distributed rollouts and a stable API, this is the wrong starting point and the README does not pretend otherwise.

## The architecture is the file layout

There is no package, no shared base class and no plugin registry. Each script contains its environment loop, its network definition and its optimizer step, which means you can read one file top to bottom and understand the whole algorithm. The README calls the layout "flat", giving 1-grid-world/3-sarsa.py as the example, and notes this replaced an older nested arrangement such as 1-grid-world/4-sarsa/sarsa_agent.py. That trade is deliberate: a flat file duplicates the training loop across algorithms, so a bug in the loop has to be fixed in nine places, but nothing is hidden behind an abstraction you did not ask for. The data flow is the conventional one for each family. Grid World scripts operate on a small tabular state space and print or render a policy. CartPole scripts use PyTorch networks with gymnasium environments and can render through pygame, which the README lists as a replacement for tkinter so that no system Tk is required. The Atari scripts add wrappers and, optionally, Weights & Biases logging. Two of the Atari-hard files are not reinforcement learning at all in the usual sense: 2-go-explore.py is described in the README's own terms as "a search result, not an RL score", and the robustification script is a backward algorithm that bootstraps from a demonstration.

## Installing it and running your first algorithm

The setup is short because the project uses uv rather than a requirements file. The README states the requirement as Python 3.11 and uv, and pyproject.toml pins requires-python to "==3.11.*", so a 3.12 interpreter will not satisfy the resolver. Clone the repository and sync:

```bash
git clone <this repo>
cd reinforcement-learning
uv sync
```

uv reads pyproject.toml, which declares torch, torchvision, gymnasium[atari], ale-py, numpy, matplotlib, pygame, opencv-python-headless, wandb, moviepy and envpool, and writes a lockfile. The first sync pulls a large amount of compiled material, so expect it to take a while. Note that envpool is a dependency of the project even for readers who only intend to run the tabular scripts.

The fastest thing to run is a Grid World algorithm, because it needs no rendering and finishes almost immediately:

```bash
cd 1-grid-world && uv run python 3-sarsa.py
```

Swap 3-sarsa.py for 1-policy_iteration.py, 2-value_iteration.py, 4-q_learning.py, 5-deep_sarsa.py or 6-reinforce.py to compare algorithms on the same problem, which is the point of keeping them in one directory. For a neural example, the README gives three CartPole invocations: training, training with rendering, and replay of a saved checkpoint.

```bash
cd 2-cartpole && uv run python 1-dqn.py
cd 2-cartpole && uv run python 1-dqn.py --render
cd 2-cartpole && uv run python 1-dqn.py --test
```

Rendering is described as slower, which follows from the fact that it is drawing every frame. The --test path expects a trained checkpoint, and the README does not document where that checkpoint is written or what happens if it is missing, so check the script before relying on it.

Atari runs can stream metrics to your own Weights & Biases account. The README presents this as opt-in twice over: login once, then pass --wandb, and "Omit --wandb and the script runs without ever touching the network."

```bash
uv run wandb login
cd 3-atari && uv run python 2-ppo.py --env breakout --wandb
cd 3-atari && uv run python 1-dqn.py --env breakout --wandb
```

The runs land in your own rl-atari-ppo or rl-atari-dqn project. The --env flag takes breakout in both examples; the README does not enumerate the other accepted values.

## The published numbers and what they do not cover

The README includes a benchmark table, which is unusual for a teaching repository and worth reading carefully. The measurements were taken on a MacBook Pro 14" with an Apple M3 and 8 GB of unified memory, macOS 26.2, Python 3.11 and PyTorch 2.11 on the MPS backend. On Breakout with sticky actions at 10M agent steps, DQN is listed at about 9 hours with a final mean per-game return of 93.5 and 5.27 GB peak RAM, while PPO is listed at about 3.8 hours with 261.9 and 1.98 GB. Both have 1.69M trainable parameters. The README is explicit that these are "single seed per row", that CPU and GPU percentages come from Activity Monitor after roughly five minutes of stabilization, and that sticky actions with repeat_action_probability=0.25 make these scores lower than the deterministic v4 environments cited in older papers. That last note matters more than the numbers themselves: comparing 261.9 here against a PPO figure from a paper using a different action-repeat setting is not a comparison. The Montezuma's Revenge table is even more careful, labelling its two protocols "not cross-comparable" and giving the Go-Explore result as 31,000 replay-verified under deterministic search at 500M frames, against roughly 3,120 for PPO with RND under sticky-action RL at 65M frames. The README also records that the robustification script produced no from-reset score, because its curriculum plateaued around 22% on one machine, and it names the cause as a scale ceiling relative to the original work's hundreds to thousands of environments. Publishing a failed run alongside the successful ones is the most credible thing in the document.

## The pruning is the biggest limitation

The README's Updates section states that the project was "pruned to 9 core algorithms" and that Monte Carlo, DDQN, A3C, Atari and mountaincar were dropped, with PPO added. If you came to this repository expecting an A3C implementation, the topic list still advertises a3c and actor-critic, but the algorithm is gone from the code. That is a real mismatch between the repository metadata and its contents, and it is the kind of thing that costs an afternoon. The same section lists the modernization: Keras with TensorFlow 1.0 became PyTorch 2.11, gym 0.8 became gymnasium 1.2, tkinter became pygame, and requirements.txt became pyproject.toml with uv. Those are the right moves, but they also mean every third-party tutorial, blog post and Stack Overflow answer written against the 2017 version now points at code that no longer exists in this form. Beyond the pruning, the single-seed benchmarks mean you cannot use the published scores to decide whether an algorithm is better than another. And the hard-exploration directory is a different category of artifact: a Go-Explore script that archives cells and replays trajectories has almost nothing in common with the policy-gradient files next to it, so a reader who runs it expecting a reinforcement learning example will be confused by what the README itself calls a search result.

## Alternatives and where they differ

The obvious comparison is Stable-Baselines3, which also provides PyTorch implementations of DQN, A2C and PPO. The difference is not quality but purpose. Stable-Baselines3 exposes algorithms as classes you construct, configure and train through a shared interface, with vectorized environments, callbacks and a documented model-saving format; the algorithms are the product. Here the algorithms are the subject, and the interface is a script you edit. That makes this repository better for understanding what the update equation does and worse for running a hundred configurations overnight. A second comparison is CleanRL, which shares the single-file philosophy but targets reproducible benchmark results with fixed seeds and logged runs rather than readability as the primary goal. If your aim is to cite a number, CleanRL's structure serves you better; if your aim is to see the loss function and the environment loop in one screen, this repository is the more direct read. The README's own benchmark section, with its single-seed caveat and its explicit note about sticky actions, suggests the author is aware of which side of that line the project sits on.

## Maintenance, licence and upgrade cost

The repository is not archived, and the last push was on 2026-06-12, which is recent enough that the pinned dependency ranges look current rather than aspirational. There are no releases in the repository, so there is nothing to upgrade between; you track the master branch or you pin a commit yourself. The pins are narrow, with torch at >=2.11,<2.12 and requires-python at exactly 3.11, so a Python 3.12 or 3.13 environment will not install without editing pyproject.toml, and a torch 2.12 release will require a change to the constraint rather than just a re-sync. Two of the dependencies, envpool and ale-py, are compiled packages with their own platform requirements, and the README's benchmarks were produced on Apple silicon, so a Linux or Windows reader should treat the first uv sync as the real compatibility test. The licence is MIT, which permits commercial use, modification and redistribution provided the copyright notice and permission notice are retained; the repository ships a LICENSE file at the top level. That is a permissive arrangement, but it says nothing about the licences of the Atari ROMs or the gymnasium environments you might run against, which are separate questions and not addressed here.

## Conclusion

Adopt this repository if you want to read a working Q-learning or PPO implementation end to end without opening a framework's source tree, or if you are teaching a course and want one file per algorithm with a paper citation at the top. Do not adopt it as a training framework: there are no releases, no packaging beyond a local pyproject.toml, and the Atari benchmarks are single-seed runs on one machine. Before you build anything on it, confirm that uv sync resolves on your platform, since envpool and ale-py are compiled dependencies, and check whether the algorithm you need survived the pruning, because the README states that Monte Carlo, DDQN, A3C, Atari DQN variants and mountaincar were dropped.

## FAQ

### What is reinforcement learning and examples?

The repository answers this by construction rather than by definition: it holds nine runnable examples, from tabular policy iteration on a Grid World through SARSA and Q-learning to DQN, A2C and PPO on CartPole and Atari. Each file is one algorithm.

### Is reinforcement learning tough?

The repository's premise is that the code does not have to be. The README describes the examples as easy to read, one file per algorithm, and states that each file opens with a paper citation and the core update equation. The benchmarks, however, show a DQN Breakout run taking roughly 9 hours on a MacBook Pro M3.

### how to set up reinforcement learning

The README requires Python 3.11 and uv, then gives three commands: clone the repository, cd into it, and run uv sync. pyproject.toml pins requires-python to 3.11 exactly and declares torch, gymnasium[atari], ale-py, envpool and other dependencies.

### how to use reinforcement learning

Run a script directly with uv. The README's examples include cd 1-grid-world && uv run python 3-sarsa.py for a tabular algorithm and cd 2-cartpole && uv run python 1-dqn.py for a neural one, with --render to watch training and --test to replay a checkpoint.

### what is reinforcement learning in ai

This repository does not define the term; it demonstrates it. The README frames the project as code examples spanning the basics to deep reinforcement learning, with one file per algorithm, from policy iteration to PPO.

## Sources

- [Issues](https://github.com/rlcode/reinforcement-learning/issues)
- [License: MIT](https://github.com/rlcode/reinforcement-learning/blob/master/LICENSE)
- [README](https://github.com/rlcode/reinforcement-learning/blob/master/README.md)
- [rlcode/reinforcement-learning on GitHub](https://github.com/rlcode/reinforcement-learning)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/rlcode-reinforcement-learning
