Library / SDK
vwxyzjn/cleanrl avatar
vwxyzjn/cleanrl

CleanRL: Every Deep RL Algorithm in a Single File

High-quality single file implementation of Deep Reinforcement Learning algorithms with research-friendly features (PPO, DQN, C51, DDPG, TD3, SAC, PPG)

10,478 stars1,173 forksPythonNOASSERTION

At a glance

What is it?
CleanRL is a Python deep reinforcement learning library where each algorithm variant lives in one self-contained file. Researchers get readable implementations of PPO, DQN, SAC, and four more algorithms without tracing through a modular codebase.
Who is it for?
CleanRL fits researchers who need to read every implementation detail of PPO, DQN, or SAC in a single file without tracing through inheritance chains. It is the wrong tool for teams that want a stable, pip-importable library for production systems.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 163 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The Single-File Design and What It Means for Reading DRL Code

CleanRL takes a position unusual among deep reinforcement learning libraries: every implementation detail of an algorithm variant lives in a single standalone Python file. When you open ppo_atari.py, which the README notes has about 340 lines, you see the environment wrappers, network architecture, training loop, loss function, and logging calls together in one place. Nothing is hidden in a parent class or a separate utilities module. This design accepts code duplication across files in exchange for local readability. The README states this directly: the library is not meant to be imported as a dependency. Its purpose is to serve as a reference for researchers who want to understand exactly what runs during training, without tracing through inheritance chains or locating scattered configuration parameters. The trade-off is explicit: duplicate code everywhere, but one file per algorithm variant that a reader can understand from top to bottom.

Installing CleanRL and Running a First Experiment

CleanRL requires Python 3.8 or newer, with an upper bound of Python 3.10. The pyproject.toml file specifies Python <3.11. The README recommends uv for dependency management. Clone the repository and install core dependencies with these two commands.

bash
git clone https://github.com/vwxyzjn/cleanrl.git && cd cleanrl
uv pip install .

Running a PPO experiment on CartPole takes a single command. The --env-id flag selects a Gymnasium environment, and --total-timesteps controls how long training runs.

bash
python cleanrl/ppo.py --env-id CartPole-v1

After the run, open Tensorboard to inspect logged metrics from the runs/ directory.

bash
tensorboard --logdir runs

Alternatively, if pip is preferred over uv, install dependencies from the requirements files in the requirements/ directory, then run algorithm files directly with python.

Algorithms, Environments, and Optional Dependencies

The library includes seven main algorithm families: PPO, DQN, C51, DDPG, TD3, SAC, and PPG. Each algorithm has multiple file variants for different environment types. PPO alone has files for classic control, Atari with a CNN, Atari with LSTM, continuous action spaces, multi-GPU training, and JAX-backed execution. The core install covers Gymnasium environments for classic control tasks. Additional environment groups require separate optional dependencies. Atari games need the atari extras, which include ale-py and the ROM license acceptor. MuJoCo physics simulation requires the mujoco extras. Multi-agent training with PettingZoo uses the pettingzoo extras. JAX-backed algorithm variants need jax, jaxlib, flax, optax, and chex, all at pinned versions listed in pyproject.toml. These version pins exist because CleanRL's benchmarked results were obtained with specific framework versions. Reproducing those published results requires matching the exact versions.

bash
uv pip install ".[atari]"
python cleanrl/ppo_atari.py --env-id BreakoutNoFrameskip-v4

Experiment Tracking with Tensorboard and Weights and Biases

Every CleanRL file writes Tensorboard logs to a runs/ directory by default. You can launch Tensorboard during or after training. For experiment management across many runs, pass the --track flag to enable Weights and Biases integration. The first time you use W&B, authenticate with wandb login. The README provides an example showing how to attach a project name and combine tracking with seed-based reproducibility.

bash
wandb login
uv run python cleanrl/ppo.py \
    --seed 1 \
    --env-id CartPole-v0 \
    --total-timesteps 50000 \
    --track \
    --wandb-project-name cleanrltest

The library also captures video of gameplay during training and stores it alongside Tensorboard logs. This combination of fixed seeds, Tensorboard output, and video capture is what the README describes as local reproducibility. Each algorithm file handles these features independently, so you can see every line responsible for logging by reading the file directly.

Cloud Scaling with Docker and AWS Batch

CleanRL includes a Dockerfile and a cloud/ directory to support running experiments at scale. The Dockerfile builds on nvidia/cuda:11.4.2-runtime-ubuntu20.04 and installs CleanRL's core dependencies via uv. An entrypoint.sh script handles container startup. The repository layout also contains cleanrl_utils/, which holds utilities for managing experiment submissions. The README mentions that the library can scale to run thousands of experiments using AWS Batch, though the detailed cloud setup lives at docs.cleanrl.dev rather than in the README. The repository also includes a Gitpod configuration at the root that provisions a browser-based development environment. This allows a researcher without a local GPU to explore the code and run lightweight experiments without configuring a local Python environment. The Gitpod option uses the .gitpod.Dockerfile to set up the container.

Limitations: Python Version Ceiling and No Import Interface

The most immediate constraint is the Python version ceiling. CleanRL does not support Python 3.11 or newer. Teams using newer Python versions for other work need a separate virtual environment or a containerized setup. The intentional lack of a public API means CleanRL cannot be added to a project as a library dependency; you copy files rather than import classes. When an algorithm file needs a bug fix, every copy held by a research team diverges unless manually synchronized. The pinned dependency versions in pyproject.toml, such as torch==2.4.1 and gymnasium==0.29.1, can conflict with other packages in a shared environment. The README also notes that the `ppo_atari_envpool.py` speedup using EnvPool is only available on Linux, so macOS users cannot run those particular variants. If you need a stable, versioned API, CleanRL is not designed for that use case.

CleanRL versus Stable Baselines3 and LeanRL

Stable Baselines3 is a modular library for deep RL designed to be imported as a Python package. It provides a consistent API across algorithms and targets applied research and production use. CleanRL is not a replacement for this role; the README explicitly states it is not meant to be imported. The two serve different needs: Stable Baselines3 for teams that want a stable dependency, CleanRL for researchers who want to read or modify implementations. For raw training speed, the README mentions LeanRL, a project from pytorch-labs that provides optimized PyTorch implementations of CleanRL algorithms using CUDA Graphs to reduce host overhead. The trade-off is that CUDA Graph versions are harder to read and modify. For offline RL in a comparable single-file style, the CORL project from corl-team provides implementations following CleanRL's design conventions. Both related projects are listed in the README.

Editorial conclusion

CleanRL fits researchers who need to read every implementation detail of PPO, DQN, or SAC in a single file without tracing through inheritance chains. It is the wrong tool for teams that want a stable, pip-importable library for production systems. Before using it, confirm your Python version falls between 3.8 and 3.10, and install the specific optional requirements file that matches your target environment, such as requirements-atari.txt or requirements-mujoco.txt.

Frequently asked questions

How do I get started with CleanRL?

Clone the repository with git clone https://github.com/vwxyzjn/cleanrl.git, install with uv pip install ., then run any algorithm file directly, such as python cleanrl/ppo.py --env-id CartPole-v1. The README also describes a Gitpod-hosted environment for browser-based setup without a local installation.

How does CleanRL compare to Stable Baselines3?

CleanRL puts every algorithm implementation into a single, self-contained file so a reader can understand all details without tracing through a modular inheritance hierarchy. Stable Baselines3 is a modular library designed to be imported as a dependency for production projects. CleanRL is explicitly not designed to be imported.

What Python versions does CleanRL support?

The pyproject.toml file specifies Python >=3.8 and <3.11. Python 3.11 and newer are not supported by the current configuration.

Official sources

  1. Issues
  2. Project website
  3. README
  4. Releases
  5. vwxyzjn/cleanrl on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/vwxyzjn-cleanrl.svg)](https://hysenlabs.com/projects/vwxyzjn-cleanrl)