Model or dataset
DLR-RM/rl-baselines3-zoo avatar
DLR-RM/rl-baselines3-zoo

RL Baselines3 Zoo: tuned hyperparameters and pretrained agents on top of Stable Baselines3

A training framework for Stable Baselines3 reinforcement learning agents, with hyperparameter optimization and pre-trained agents included.

2,886 stars609 forksPythonMIT

At a glance

What is it?
RL Baselines3 Zoo is a training framework that wraps Stable Baselines3 with per-environment hyperparameter files, evaluation and plotting scripts, and a collection of pretrained agents. It is for engineers who want reproducible RL training runs rather than a new algorithm library.
Who is it for?
Adopt it if you are already on Stable Baselines3 and want the tuned YAML files, the train.py and enjoy.py entry points, and the pretrained agents under rl-trained-agents. Skip it if you need a library that implements new algorithms: the zoo contributes configuration and tooling, not methods, and its benchmark.md is explicitly a single-run check, not a quantitative comparison.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 8 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What RL Baselines3 Zoo actually adds to Stable Baselines3

Stable Baselines3 gives you algorithm implementations and a common API. It does not tell you which learning rate, batch size or n_steps value works for HalfCheetahBulletEnv-v0, and it does not ship a runner that evaluates every 10000 steps and writes logs you can plot later. RL Baselines3 Zoo fills that gap. The README lists its goals plainly: a simple interface to train and enjoy agents, benchmark the algorithms, and provide tuned hyperparameters for each environment and algorithm.

The intended user is someone who has already picked an algorithm and now has to make it converge. That is a configuration and experiment-management problem, and the repository treats it as one. The hyperparameters live in YAML files under hyperparams/, one per algorithm name, and the training entry point reads them by convention rather than by flags. The README states that if the environment exists in hyperparameters/algo_name.yml, then you can train with a single command.

The second audience is people who want to watch a trained policy without training anything. The repository ships a git submodule of pretrained agents and a Hugging Face page, and the README claims a collection of 200+ trained agents. Those are useful for sanity checks and demos. They are not a leaderboard, and the README says so directly: the benchmark corresponds to only one run per configuration and is meant to check maximal performance and find bugs.

How training, evaluation and tuning fit together

The data flow is file-driven. A YAML file under hyperparams/ maps an algorithm to environments and their parameter dictionaries. train.py resolves the algorithm name and environment id, loads the matching entry, constructs the Stable Baselines3 model, and runs the training loop with periodic evaluation. Evaluation frequency, episode count and the number of evaluation environments are command-line arguments, so the same script covers both a quick smoke run and a long job.

Tuning is a separate path. The dependencies include optuna, and the optional extras include optunahub for what the pyproject file labels Optuna auto. The README points to a dedicated tuning section of the documentation rather than describing the search space inline, so the exact sampler and pruning configuration is something you read in the docs, not something the README guarantees. The presence of optuna in the core dependency list means the tuning machinery is installed by default, not behind an extra.

The repository also separates concerns at the top level: train.py and enjoy.py are thin scripts, rl_zoo3/ holds the package, scripts/ holds shell helpers, and the Makefile exposes pytest, mypy, lint, format, doc and docker targets. The Docker targets build CPU and GPU images through scripts/build_docker.sh, with the GPU variant selected by USE_GPU=True. That layout matters if you plan to run training inside a container, because the build script is the supported path rather than a hand-written Dockerfile.

Installing rl_zoo3 and running a first training job

The README gives two installation routes. From a clone, install in editable mode; as a package, install rl_zoo3 from PyPI. Both are single commands, and the package installs a console entry point so you can run the same code from any directory.

bash
pip install -e .
bash
pip install rl_zoo3

The README notes that python -m rl_zoo3.train works from any folder and that rl_zoo3 train is equivalent to python train.py. If you want the extra environments, plotting dependencies and test tooling in one go, the full installation is the documented path.

bash
apt-get install swig cmake ffmpeg
pip install -e .[plots,tests,extras]

On macOS the README flags a pybullet build failure with errors from stdio.h and gives a workaround, which is worth knowing before you assume the install is broken.

bash
CFLAGS="-fno-define-target-os-macros" pip install -e .[plots,tests,extras]

With the package installed, training an environment that already has an entry in hyperparams/ takes one command. The README's example uses sac on HalfCheetahBulletEnv-v0, evaluating every 10000 steps with 10 episodes across a single evaluation environment.

bash
python train.py --algo sac --env HalfCheetahBulletEnv-v0 --eval-freq 10000 --eval-episodes 10 --n-eval-envs 1

What you should see is a training run that periodically evaluates the policy and writes logs and model checkpoints. To watch a pretrained agent instead, the README gives the enjoy.py path with a folder argument and a timestep budget.

bash
python enjoy.py --algo a2c --env BreakoutNoFrameskip-v4 --folder rl-trained-agents/ -n 5000

That command only works if the pretrained agents are present. The README is explicit that you must clone with --recursive to pull the submodule. A plain clone leaves rl-trained-agents empty and enjoy.py with nothing to load.

Where the zoo is the wrong tool

The repository does not implement reinforcement learning algorithms. Anything you get here, you get because Stable Baselines3 or sb3_contrib provides it. If your research needs an off-policy method that SB3 does not ship, the zoo will not help you; you would be adding the algorithm yourself and then, optionally, adding a YAML entry for it.

The pretrained agents are demos, not baselines for publication. The README states that the benchmark corresponds to only one run per configuration and links to an issue about that limitation. If you need statistical comparisons across seeds, the zoo's own numbers are not the evidence you want, even though the optional plots extra pulls in rliable, which is built for aggregated multi-seed reporting. That combination is a reasonable hint about intended use, but the shipped benchmark.md is not that.

Environment coverage is also uneven by design. The README describes the collection as incomplete and asks for contributors. The Atari tables mark seven games as covered and label a second table of three games as to be completed. A new or unusual environment may simply have no tuned entry, and the release notes for v2.8.0 mention adding default hyperparameters for unlisted environments, which tells you that unlisted environments were previously handled less gracefully.

Finally, the tooling assumes Gymnasium-style environments. The dependency range on gymnasium was relaxed in v2.9.1 to allow any 1.x, and the extras still pin gym==0.26.2 for compatibility work. If your stack is built on a different environment interface, expect to write an adapter before you write a training config.

Compared with running Stable Baselines3 directly

The honest alternative is Stable Baselines3 on its own. You get the same algorithms, the same callbacks and the same save and load format, and you avoid a layer of YAML resolution and script conventions. For a single experiment where you already know the hyperparameters, writing a short script against SB3 is less indirection than learning the zoo's file layout.

The difference shows up at the second and tenth experiment. The zoo's contribution is that the hyperparameters for a given environment are a checked-in artifact rather than a value buried in someone's script, and that train.py, enjoy.py and the benchmark module all read from the same place. Adding a new environment means adding a YAML block, not editing Python. The trade-off is that you inherit the zoo's conventions: the algorithm name must match a file under hyperparams/, and the environment id must match a key in it, or you fall back to the default hyperparameters path.

A second alternative for the tuning half is Optuna used directly. The zoo depends on Optuna anyway, so the question is whether you want the zoo's search-space definitions and its integration with the training script, or your own study code. Direct Optuna gives you full control over the objective and the sampler; the zoo gives you a starting point that already knows how to construct and evaluate an SB3 model. Neither is strictly better, and the pyproject extras list optunahub, which suggests the project expects users to bring additional Optuna components rather than rely only on what is bundled.

Maintenance, licence and upgrade cost

The repository is not archived, and the last push was on 2026-09-09, so it is current. The release cadence is visible in the version history: v2.8.0 on 2026-04-01 added default hyperparameters for unlisted environments, and v2.9.0 and v2.9.1 both landed on 2026-06-15, the latter relaxing the Gymnasium version range to allow any 1.x. Recent releases have been dependency and quality-of-life work rather than new algorithms, which is consistent with a project whose value is configuration and tooling.

Upgrade cost concentrates in the dependency ranges. The package requires Python 3.10 or newer and pins sb3_contrib and stable-baselines3 to the 2.x line, gymnasium to a bounded range, and huggingface_sb3 to 3.x. The plots extra pins pandas below 3 and scipy at roughly 1.10, and the pyproject file carries a comment explaining that rliable 1.2.0 required an older architecture that conflicts with pandas 3.x. If your environment already resolves a newer pandas, the plots extra is where you will feel friction.

The licence is MIT, declared both in the repository and in the pyproject metadata, with the LICENSE file listed among the top-level entries. MIT is permissive, so the practical question is not whether you may use it but what obligations you take on with the pretrained agents you download and redistribute; those are separate artifacts, and the README points to a Hugging Face page rather than stating terms for them. Check the licence attached to the specific agent files you ship, and treat the MIT grant on the repository code as covering the code.

Editorial conclusion

Adopt it if you are already on Stable Baselines3 and want the tuned YAML files, the train.py and enjoy.py entry points, and the pretrained agents under rl-trained-agents. Skip it if you need a library that implements new algorithms: the zoo contributes configuration and tooling, not methods, and its benchmark.md is explicitly a single-run check, not a quantitative comparison. Before committing, clone with --recursive so the submodule is present, confirm that your environment id appears in hyperparams/, and read the tuning guide for the Optuna path you intend to use.

Frequently asked questions

What is the purpose of reinforcement learning?

The repository does not define reinforcement learning in general; it assumes you already work in that field and provides training, evaluation, hyperparameter tuning and video recording for Stable Baselines3 agents. Its stated goals are a simple interface to train and enjoy agents, benchmark algorithms, and supply tuned hyperparameters per environment and algorithm.

What are the different types of reinforcement learning?

The zoo does not categorise RL methods itself. It organises its tuned hyperparameters by algorithm name in files under hyperparams/, and the benchmark tables cover A2C, PPO, DQN, QR-DQN and ARS across Atari and classic control environments, so the algorithm taxonomy you get is whatever Stable Baselines3 and sb3_contrib expose.

What are the algorithms used in reinforcement learning?

The README's benchmark tables list A2C, PPO, DQN and QR-DQN for the Atari games, and add ARS for the classic control environments such as CartPole-v1 and Pendulum-v1. The actual implementations come from Stable Baselines3 and sb3_contrib, which the package depends on.

Is RL the future of AI?

The repository takes no position on that. It is a training framework for Stable Baselines3 agents, and its stated goals are a simple training and demo interface, benchmarking the algorithms, and supplying tuned hyperparameters per environment and algorithm.

Official sources

  1. DLR-RM/rl-baselines3-zoo on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/dlr-rm-rl-baselines3-zoo.svg)](https://hysenlabs.com/projects/dlr-rm-rl-baselines3-zoo)