Library / SDK
thu-ml/tianshou avatar
thu-ml/tianshou

Tianshou 2.0: a PyTorch RL library split into Algorithm and Policy

An elegant PyTorch deep reinforcement learning library.

10,997 stars1,341 forksPythonMIT

At a glance

What is it?
Tianshou 2.0.1 separates learning algorithms from policies, keeps online, offline and multi-agent RL in one codebase, and breaks compatibility with 0.x. Here is what that buys you and what it costs.
Who is it for?
Adopt Tianshou 2.x if you are starting a new project on Python 3.11, want on-policy, off-policy and offline algorithms behind one trainer interface, and can read the migration notes in CHANGELOG.md before porting anything. Do not adopt it as a drop-in upgrade for an existing 0.x codebase: the release note states version 2 is not backwards compatible, and the cost lands on every custom policy or algorithm subclass you own.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 179 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Tianshou solves, and who ends up using it

Deep reinforcement learning code has a habit of collapsing into one file: the network, the rollout loop, the replay buffer and the update rule fused together. Tianshou's answer is to keep the pieces apart. The README describes it as a reinforcement learning library based on pure PyTorch and Gymnasium, with a modular low-level interface for algorithm developers and a convenient high-level interface for people who just want to train an implemented algorithm on a custom environment.

That split maps onto two audiences. The first is the RL researcher who wants to modify a loss function or swap an advantage estimator without rewriting the training loop. The second is the practitioner who has a Gymnasium environment and wants PPO or SAC running on it. The supported algorithm list covers DQN and its variants (Double, Dueling, Branching, C51, Rainbow, QRDQN, IQN, FQF), the policy-gradient family (PG, NPG, A2C, TRPO, PPO), continuous control (DDPG, TD3, SAC, REDQ, SAC-Discrete), imitation learning (vanilla imitation, GAIL), and offline methods (BCQ, CQL, TD3+BC, BCQ-Discrete, CQL-Discrete, CRR-Discrete). Replay and exploration components such as Prioritized Experience Replay, GAE, ICM and HER are listed as separate building blocks rather than baked into one algorithm.

The scope is wider than most single-library RL projects: online on-policy, online off-policy, offline, experimental multi-agent support, and experimental model-based support. If your work sits entirely in one of those boxes, that breadth is not a selling point, it is extra surface area. If your work crosses between them, it is the reason to look here.

The Algorithm and Policy split in version 2

The version 2 release note is explicit about what changed: a clear separation between learning algorithms and policies through two abstractions, Algorithm and Policy. In practice this means a policy owns the mapping from observation to action and the networks behind it, while the algorithm owns the update rule that consumes collected data and produces gradients. A trainer then wires a policy, an algorithm, an environment and a buffer together.

The release note also states that the class hierarchy was revised to separate on-policy, off-policy and offline algorithms at the type level, and that parameter names were made more consistent, with the documentation expanded to cover algorithm and trainer parameters. The type-level separation matters more than it sounds: offline algorithms consume a fixed dataset and cannot interact with the environment during training, while on-policy algorithms discard data after each update. Encoding that in the hierarchy stops you from assembling a nonsensical combination and only discovering it at runtime.

Data collection is the other half. The README states that vectorized environments, synchronous or asynchronous, are supported for all algorithms, and that EnvPool-based vectorized environments are supported for all algorithms as well. Recurrent state representations in actor and critic networks are also supported, which is what you need for partially observable problems. The cost of the redesign is stated plainly: version 2 is not backwards compatible with previous versions, and migration information lives in CHANGELOG.md.

Installing Tianshou 2.0.1 and running a first training job

The package is distributed on PyPI under the name tianshou, and the repository builds with poetry-core. The pyproject.toml pins python = "^3.11", so a 3.11 interpreter is the baseline. The dependency list includes gymnasium, torch, numpy, numba, h5py, pandas, tensorboard, tqdm, pettingzoo and sensai-utils. Note the torch constraint in pyproject.toml: torch = "^2.0.0, !=2.0.1, !=2.1.0", with a comment in the file pointing at a PyTorch issue as the reason those two versions are excluded.

A plain pip install is the shortest path:

bash
pip install tianshou

Environment families such as Atari, box2d, classic-control and MuJoCo are optional dependencies of gymnasium, and the pyproject.toml comment explains that the project maintains its own list of dependencies for them instead of relying on extras of another package. So installing tianshou does not install MuJoCo. Check the extras section of pyproject.toml for the group matching your environment before you assume the environment will import.

The repository also ships a Dockerfile for a reproducible setup. It is built from python:3.11-slim, installs poetry through pipx, sets WORKDIR /workspaces/tianshou, copies pyproject.toml, poetry.lock and README.md, then runs poetry config virtualenvs.create false followed by poetry install --no-root --with dev. The entrypoint re-runs poetry install --with dev and trusts the notebooks under notebooks/ and docs/02_notebooks/ before executing your command. The Dockerfile comment states the code is expected to be mounted into the container; if you do not mount it, you override the entrypoint.

For a first real run, the repository keeps worked examples in examples/ grouped by setting: atari, box2d, discrete, mujoco, offline, inverse, modelbased and vizdoom. Pick the directory that matches your environment rather than adapting a tutorial from a different family, since the policy and algorithm pairing differs between discrete and continuous control. Before you port any existing 0.x code, read CHANGELOG.md, because the release note states the API is not backwards compatible.

Where Tianshou 2.x gets in your way

The migration is the first limitation, and it is not a small one. The release note says version 2 is a complete overhaul of the procedural API and is not backwards compatible. Every custom policy, custom algorithm or custom trainer subclass written against 0.x has to be rewritten against the new abstractions. The changelog is the only migration path the repository documents.

The second is the Python floor. pyproject.toml requires python = "^3.11". If your training cluster, your lab's shared environment or a dependency such as an older simulator is pinned to 3.9 or 3.10, Tianshou 2.x is simply unavailable to you. That is a real constraint, not a preference.

The third is the dependency weight. The core install pulls in numba, h5py, pandas, tensorboard, pettingzoo and sensai-utils alongside torch and gymnasium. For a small experiment that only needs PPO on a classic-control task, that is a lot of surface to keep in sync.

The fourth is maturity signalling inside the project's own metadata. pyproject.toml classifies the package as "Development Status :: 4 - Beta", the README labels multi-agent and model-based RL as experimental, and the release history shows v2.0.0b2, then v2.0.0, then v2.0.1. Treat the 2.x line as a young API, and pin your version accordingly. The last push to the repository was on 2026-04-03, and v2.0.1 was released on 2026-04-02, so the 2.x line has not had a long tail of patch releases behind it.

Tianshou against Stable-Baselines3, and when to pick the other one

Stable-Baselines3 is the obvious comparison point, and the difference is architectural rather than a matter of feature checklists. SB3 is built around a small set of well-tested on-policy and off-policy algorithms with a stable, deliberately narrow API: you construct a model, call learn, and the library owns the training loop. Tianshou exposes the loop. The Algorithm and Policy abstractions, the separate buffer, and the trainer mean you assemble the pieces yourself, which is more code up front and more control afterwards.

The second difference is scope. SB3 concentrates on online RL. Tianshou's README lists offline algorithms (BCQ, CQL, TD3+BC, CRR-Discrete), imitation learning (GAIL, vanilla imitation), experimental multi-agent support through pettingzoo, and experimental model-based support. If your project is offline RL or imitation learning, SB3 does not cover the ground and the comparison ends there.

The third is the migration posture. SB3 has kept its interface stable across releases. Tianshou 2.x did the opposite on purpose: it rewrote the procedural API to get cleaner type-level separation. That trade is defensible for a research library, but it means the version you pin matters more here than it does with SB3, and it means any code you find online written for Tianshou 0.x will not run against 2.x.

Licence and the cost of staying current

Tianshou is MIT licensed, and pyproject.toml declares license = "MIT" with the matching classifier. MIT is permissive: you can use, modify and redistribute the code, including in closed products, provided the copyright notice and permission notice are preserved. That is the general shape of the licence, not legal advice, and your own dependency obligations (torch, gymnasium and the rest) are separate questions you should check independently.

Upgrade cost is where the version history matters. The jump from 0.x to 2.x is a rewrite, not a patch, and the changelog is the documented route through it. Within 2.x, the release cadence so far is one patch after the initial release, so the practical advice is to pin tianshou to a specific version in your lockfile rather than tracking the latest, and to re-read CHANGELOG.md before each bump. If you maintain a fork with custom algorithms, budget for re-applying your changes whenever the Algorithm or Policy interfaces move, since those are the abstractions the 2.x redesign reorganised.

Editorial conclusion

Adopt Tianshou 2.x if you are starting a new project on Python 3.11, want on-policy, off-policy and offline algorithms behind one trainer interface, and can read the migration notes in CHANGELOG.md before porting anything. Do not adopt it as a drop-in upgrade for an existing 0.x codebase: the release note states version 2 is not backwards compatible, and the cost lands on every custom policy or algorithm subclass you own. Before you commit, verify three things in the repository: that the environment group you need is installed through the extras listed in pyproject.toml, that the example under examples/ matching your setting (atari, mujoco, offline, modelbased) runs, and that the torch constraint in pyproject.toml is compatible with the rest of your stack.

Frequently asked questions

What is Tianshou?

It is a deep reinforcement learning library built on pure PyTorch and Gymnasium, covering online on-policy and off-policy algorithms, offline RL, and experimental multi-agent and model-based support. Version 2 separates learning algorithms from policies through the Algorithm and Policy abstractions.

How do I install Tianshou?

It is published on PyPI as tianshou, so pip install tianshou works. The package requires Python 3.11 or later according to pyproject.toml, and environment families such as Atari, box2d, classic-control and MuJoCo are optional dependencies you install separately.

Is Tianshou 2.0 backwards compatible with earlier versions?

No. The release note states that version 2 is a complete overhaul of the procedural API and is not backwards compatible with previous versions, with migration information in CHANGELOG.md.

Which reinforcement learning algorithms does Tianshou support?

The README lists DQN and its variants, PG, NPG, A2C, TRPO, PPO, DDPG, TD3, SAC, REDQ, imitation learning including GAIL, and offline methods such as BCQ, CQL, TD3+BC and CRR-Discrete, along with components like Prioritized Experience Replay, GAE, ICM and HER.

Official sources

  1. License: MIT
  2. Project website
  3. README
  4. Releases
  5. thu-ml/tianshou on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/thu-ml-tianshou.svg)](https://hysenlabs.com/projects/thu-ml-tianshou)