# Stable Baselines3: a PyTorch RL baseline library for people who already know RL

> Stable Baselines3 packages PPO, SAC, DQN and other algorithms behind a sklearn-like API, with Gymnasium environments and TensorBoard logging. It is a baseline toolkit, not a tutorial, and its own README says so.

**DLR-RM/stable-baselines3** — PyTorch version of Stable Baselines, reliable implementations of reinforcement learning algorithms. 

- Repository: https://github.com/DLR-RM/stable-baselines3
- Website: https://stable-baselines3.readthedocs.io
- Stars: 13,818 · Forks: 2,182
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/dlr-rm-stable-baselines3

## What Stable Baselines3 is for, and who it is not for

Stable Baselines3 (SB3) is a set of PyTorch implementations of reinforcement learning algorithms, described in its README as "reliable implementations" and as the next major version of Stable Baselines. The problem it addresses is reproducibility: instead of every lab and team writing its own PPO, SB3 provides one implementation with a common interface so results can be compared against a shared baseline.

The README is direct about the audience. It states that despite the simplicity of use, SB3 assumes some knowledge of reinforcement learning, and that you should not use the library without some practice. That sentence matters more than the feature table. If you do not know what a rollout buffer is or why a value loss curve should fall, the library will run and produce a model, and you will have no way to tell whether the model is any good. SB3 is a tool for people who can already diagnose a training run.

The intended users are researchers who need a baseline to compare a new idea against, and engineers who want an algorithm that works without reading a paper first. The README frames it as a base around which new ideas can be added.

## The common interface: Gymnasium environments, policies, and a learn loop

The architecture is deliberately flat. You construct an environment, pair it with a policy, hand both to an algorithm class, and call learn. The README notes that most of the library follows a sklearn-like syntax using Gym. The example in setup.py shows the shape: gymnasium.make returns an environment, PPO wraps it with a policy string such as "MlpPolicy", model.learn takes a total_timesteps argument, and model.get_env returns a vectorised environment you can step through manually.

That design means algorithms are interchangeable at the call site. Swapping PPO for SAC changes the class name and possibly the policy, not the surrounding script. Custom environments plug in through the Gymnasium interface, and the feature table lists custom policies, custom callbacks, Dict observation spaces and TensorBoard support as supported.

Logging is part of the mechanism rather than an add-on. The pytest configuration in pyproject.toml carries filterwarnings entries for TensorBoard deprecations, which tells you the logging path is exercised by the test suite. The README also points to the OpenRL Benchmark platform for detailed logs and reports, and mentions integrations with Weights & Biases for tracking and Hugging Face for storing and sharing trained models.

The trade-off is that the flat interface hides the training loop. When something goes wrong inside the update step, you are reading library code, not your own.

## Installing Stable Baselines3 with pip and training PPO on CartPole

The README states that SB3 requires Python 3.10 or newer and supports PyTorch >= 2.8. The v2.9.0 release notes add that pandas is now optional and that gymnasium 1.3.0 is supported. The README points Windows users to the documentation for platform-specific instructions rather than giving them inline, so treat Windows as a documented case you should read before installing.

Install from PyPI into an environment that already has a suitable PyTorch build:

```bash
pip install stable-baselines3
```

Then the quick example from setup.py, which trains PPO on CartPole and then steps the environment manually:

```python
import gymnasium

from stable_baselines3 import PPO

env = gymnasium.make("CartPole-v1", render_mode="human")

model = PPO("MlpPolicy", env, verbose=1)
model.learn(total_timesteps=10_000)
```

With verbose=1 you should see a progress bar and periodic rollout statistics printed to stdout during learn. The render_mode="human" argument opens a window, so on a headless machine drop it or switch to a different render mode. After training, model.get_env() gives you the vectorised environment used for stepping the policy yourself, as the README example continues to do.

For a containerised setup, the repository ships a Dockerfile built on mambaorg/micromamba:2.0-ubuntu24.04 with PYTHON_VERSION defaulting to 3.12 and a CPU PyTorch index at https://download.pytorch.org/whl/cpu. It installs the package editable with the extra, tests and docs groups, then swaps opencv-python for opencv-python-headless because the container has no display.

## Where Stable Baselines3 stops: maintenance mode and the contrib split

The README states that since most features from the original roadmap have been implemented, there are no major changes planned for SB3 and it is now stable, with development focused on bug fixes and maintenance such as documentation updates and user experience. That is a deliberate boundary, and it is the single most important thing to understand before adopting it.

Newer algorithms live elsewhere. The README lists Recurrent PPO (PPO LSTM), CrossQ, Truncated Quantile Critics (TQC), Quantile Regression DQN (QR-DQN) and PPO with invalid action masking (Maskable PPO) as features implemented in SB3-Contrib, a separate repository. Faster variants are developed in SBX, the Jax version, which the README says provides a minimal number of features compared to SB3 but can be much faster, citing a figure of up to 20x in a linked post. If your project needs one of those algorithms, you are installing a second package, not configuring SB3.

The practical failure mode is expecting the core library to grow. It will not, by the maintainers' own statement. A second failure mode is subtler: the README's performance claims point to per-algorithm results pages and to the OpenRL Benchmark, so the honest way to judge whether an algorithm suits your environment is to run it, not to read the feature table. The table lists capabilities, not outcomes.

## Stable Baselines3 compared with CleanRL and RLlib

The two comparisons users ask about most are CleanRL and RLlib, and they differ from SB3 in opposite directions.

CleanRL takes the single-file approach: each algorithm is a self-contained script you read top to bottom. SB3 gives you a class with a shared interface across algorithms. If your goal is to modify the update rule and understand every line, CleanRL's structure is closer to that. If your goal is to run four algorithms on one environment with the same script and compare curves, SB3's shared interface removes the glue code. The cost of SB3's abstraction is that the training loop is inside the library; the cost of CleanRL's transparency is that you maintain the loop yourself.

RLlib is a distributed training framework. SB3 is a library you import into a single Python process, with a Dockerfile in the repository for containerised runs. The README describes SB3 as a baseline toolkit and does not present distributed training as a goal. If your problem is scale across many workers, that is a different category of tool; if your problem is a trustworthy PPO on one machine, SB3 is aimed at exactly that.

The README also positions SB3 against its own relatives rather than external tools: SB3-Contrib for experimental algorithms, SBX for speed, and the RL Baselines3 Zoo as the training framework that provides scripts for training, evaluating agents, tuning hyperparameters, plotting results and recording videos, plus tuned hyperparameters for common environments.

## Licence, versioning and what an upgrade actually costs

Stable Baselines3 is MIT licensed, and the repository carries a NOTICE file alongside the LICENSE. MIT is permissive, so the practical constraint is attribution rather than copyleft obligations. This is not legal advice; read the LICENSE and NOTICE files yourself if you are redistributing.

The release cadence visible in the release notes is roughly quarterly, and the changes are dependency-shaped rather than API-shaped. v2.8.0 dropped Python 3.9 and added Python 3.13 support. v2.9.0 moved to gymnasium 1.3.0 and torch>=2.8 and made pandas optional. That pattern tells you where upgrade cost lands: on your environment, not on your training script. A pinned PyTorch or an older Python is the thing that will block an upgrade, which is why the README states the Python 3.10+ and PyTorch >= 2.8 requirements up front.

The repository layout supports this reading. pyproject.toml configures ruff with target-version py310, black at line-length 127, mypy, and pytest markers including an expensive marker for tests that can be deselected. The Makefile exposes lint, format, type, commit-checks and doc targets, and the Dockerfile pins micromamba 2.0 on Ubuntu 24.04. Upgrading SB3 means re-satisfying those pins, and the test suite is what tells you whether you did.

## Conclusion

Adopt Stable Baselines3 if you need a reference PPO, SAC or DQN implementation with a Gymnasium interface and TensorBoard logging, and you already understand RL enough to read a learning curve. Do not adopt it as a first contact with reinforcement learning, and do not expect the core repository to ship new algorithms: the README states development is focused on bug fixes and maintenance, with newer algorithms going to SB3-Contrib and faster Jax variants to SBX. Before committing, verify the Python and PyTorch versions your environment actually provides against the requirements the README states, and check whether the algorithm you need lives in the core package or in SB3-Contrib.

## FAQ

### What is Stable Baselines3 and what does it do?

It is a set of PyTorch implementations of reinforcement learning algorithms, described in the README as reliable implementations and as the next major version of Stable Baselines. It gives you a common interface over algorithms such as PPO, SAC and DQN so results can be compared against a shared baseline.

### How do I install Stable Baselines3?

The README states it requires Python 3.10 or newer and PyTorch >= 2.8, and the package installs from PyPI with pip install stable-baselines3. Windows users are pointed to the documentation for platform-specific instructions rather than inline steps.

### How do I use Stable Baselines3?

The example in setup.py makes a Gymnasium environment, wraps it with an algorithm class such as PPO using an "MlpPolicy" string, calls model.learn with a total_timesteps value, and then steps the environment returned by model.get_env. The README notes the library follows a sklearn-like syntax.

### How do I use TensorBoard with Stable Baselines3?

TensorBoard support is listed in the README feature table, and the pytest configuration in pyproject.toml carries filterwarnings entries for TensorBoard deprecations, so the logging path is covered by the test suite. The README also points to the OpenRL Benchmark platform for detailed logs and reports.

### Is there a Python library for reinforcement learning?

Stable Baselines3 is one: a Python and PyTorch library with a Gymnasium interface and a common API across algorithms. The README assumes you already have some knowledge of reinforcement learning and says you should not use it without some practice.

### What are the alternatives to Stable Baselines3?

The README points to SB3-Contrib for experimental algorithms such as Recurrent PPO and Maskable PPO, SBX for Jax-based faster variants, and the RL Baselines3 Zoo as the training framework with tuned hyperparameters. CleanRL and RLlib are different in approach: CleanRL is single-file per algorithm, RLlib is a distributed training framework.

## Sources

- [DLR-RM/stable-baselines3 on GitHub](https://github.com/DLR-RM/stable-baselines3)
- [License: MIT](https://github.com/DLR-RM/stable-baselines3/blob/master/LICENSE)
- [Project website](https://stable-baselines3.readthedocs.io)
- [README](https://github.com/DLR-RM/stable-baselines3/blob/master/README.md)
- [Releases](https://github.com/DLR-RM/stable-baselines3/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/dlr-rm-stable-baselines3
