Library / SDK
pytorch/rl avatar
pytorch/rl

TorchRL: a TensorDict-first PyTorch library for reinforcement learning

A modular, primitive-first, python-first PyTorch library for Reinforcement Learning.

3,584 stars492 forksPythonMIT

At a glance

What is it?
TorchRL is a PyTorch-native RL toolkit built around one data model, TensorDict, rather than a single algorithm. It suits engineers who want to swap environments, collectors, buffers and losses without rewriting the training loop, and it asks them to accept that data model first.
Who is it for?
Adopt TorchRL if your training loop already lives in PyTorch and you need one data model across vectorized, multiprocess, recurrent, multi-agent or offline workflows, because the collector, replay buffer and loss modules all read and write named TensorDict keys. Do not adopt it if you want a finished algorithm with a stable config schema and no interest in how data is structured, since the repository self-describes as Beta and the tutorial examples are not a supported API.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The glue code problem TorchRL is aimed at

Most reinforcement learning codebases rot in the same place. One environment returns tuples, another returns dicts, recurrent states live in a separate list next to the batch, masks sit outside the data they describe, and the loss function silently assumes a particular batch layout. Every new environment or algorithm means another adapter. TorchRL's answer is to make the layout explicit and shared: observations, pixels, actions, rewards, masks, recurrent states, agent groups, sampled indices and priorities all travel as named fields in a single TensorDict. The README states the library is built around three ideas, the first being that data should have names, structure, batch dimensions and devices all the way through the training loop. The audience is therefore not someone who wants a finished PPO script. It is someone building a training system, or porting research code between settings, who is tired of rewriting the same plumbing. The three ideas are stated as composable pieces, independent modules that can be swapped without rewriting the rest of the stack, and a data model that survives the move from a local prototype to vectorized, multiprocess, distributed, compiled, recurrent, multi-agent, model-based or offline workflows. That is a systems pitch, not an algorithms pitch.

How a TensorDict moves through the stack

The README gives a short mental model, and it is worth reading literally. A TensorDict enters a policy module, which writes actions and log-probs. The environment reads actions and writes next observations, rewards and done flags. A collector batches trajectories from one or many workers. A replay buffer stores, samples, prioritizes and transforms that data. A loss module reads named keys and writes differentiable losses. An optimizer then updates ordinary PyTorch parameters. Two details matter here. First, the environment is not special: it reads and writes the same container as every other component, so a transform stack is just more keys being added or rewritten. Second, the loss module is described as reading named keys, which is what keeps algorithm code from depending on a particular environment's output shape. TensorDict itself is a separate project, pytorch/tensordict, described in the README as a dictionary-like tensor container with PyTorch operations, device transfers, shared-memory support, memmaps, lazy views and nn.Module wrappers. TorchRL pins it tightly, at tensordict>=0.14.2,<0.15.0 in pyproject.toml, so the data model and the library version move together rather than independently.

Installing TorchRL and running a first rollout

The package is published on PyPI under the name torchrl, and pyproject.toml lists torch>=2.1.0, pyvers>=0.2.3, hoptorch>=0.1.4, numpy, packaging, cloudpickle and tensordict>=0.14.2,<0.15.0 as runtime dependencies, with Python 3.10 or newer required. Installation is therefore a normal pip install, and the build backend pulls in torch, ninja, pybind11[global] and cmake because TorchRL ships C++ extensions; setup.py raises an explicit error telling you to install pybind11 with pip install 'pybind11[global]' if it is missing during a build. That command is the fix the build itself names when the extension build cannot proceed.

bash
pip install torchrl

For a source checkout, the Makefile exposes a develop target that cleans, rebuilds the extensions in place and then installs the package in editable mode without build isolation.

bash
make develop

The README's quick demo shows what a local rollout looks like. It builds a Pendulum environment with a step counter, wraps a small network as a policy with explicit input and output keys, and asks the environment to roll out 32 steps. The assertions are the useful part: the returned rollout has batch size 32, and the reward under the next key has 32 as its leading dimension. If they pass, your environment, policy and data model are wired together correctly.

python
import torch
from tensordict.nn import TensorDictModule
from torch import nn

from torchrl.envs import PendulumEnv, StepCounter, TransformedEnv

env = TransformedEnv(PendulumEnv(), StepCounter(max_steps=200))

policy = TensorDictModule(
    nn.Sequential(
        nn.LazyLinear(64),
        nn.Tanh(),
        nn.Linear(64, 1),
        nn.Tanh(),
    ),
    in_keys=["observation"],
    out_keys=["action"],
)

rollout = env.rollout(max_steps=32, policy=policy)
assert rollout.batch_size == torch.Size([32])
assert rollout["next", "reward"].shape[:1] == torch.Size([32])

The README is explicit that nothing in this pattern is specific to Pendulum. The same keys-and-TensorDict interface is used by batched environments, multi-agent tasks, collectors, replay buffers, recurrent modules, transforms and losses. That is the claim to test on your own problem: write a custom environment that emits the keys your loss expects, and see whether the rest of the stack accepts it unchanged. Beyond the demo, the repository carries examples/ directories for agents, collectors, distributed runs, environments, llm, memmap, microduck, mujoco_macros, multiagent, replay-buffers, rlhf, satellite, services and video, plus a sota-implementations/ tree. Those are the realistic starting points for a real project, not the README snippet.

Where TorchRL gets in your way

The first cost is the data model itself. If your environment or simulator already returns a well-defined structure that your loss consumes directly, TensorDict is an extra layer with its own conventions, its own version pin, and its own debugging surface. The second cost is maturity signalling. pyproject.toml classifies the project as Development Status :: 4 - Beta, which is the project's own description of itself, and the release cadence is fast: v0.13.1 on 2026-06-08, v0.13.2 on 2026-06-17, v0.13.3 on 2026-07-14, with the last push on 2026-09-10. Fast releases are good for fixes and bad for anyone who needs a frozen API. The third cost is build friction. Installing from source requires a working C++ toolchain, pybind11 and cmake, and setup.py inspects CUDA_HOME and nvcc to decide whether CUDA extensions can be built, so a machine without a matching CUDA setup takes a different build path than one with it. If you only want to run an existing algorithm on a standard benchmark and never intend to restructure the data, TorchRL is the wrong tool: a smaller, single-algorithm implementation will be faster to read and to modify.

How it differs from Stable-Baselines3 and CleanRL

The difference is architectural, not a matter of which trains faster. Stable-Baselines3 and CleanRL are organized around algorithms: you pick PPO or SAC, supply an environment through a Gym-style interface, and the library owns the training loop. TorchRL is organized around components. The README says outright that it is not a single algorithm implementation or a narrow benchmark suite, but a collection of composable pieces. In practice that means you assemble the collector, replay buffer and loss yourself, and you are responsible for the keys that connect them. The payoff is that swapping a replay buffer for a prioritized one, or moving from single-agent to multi-agent, is a change to one component rather than a fork of the trainer. The cost is that there is no single trainer to fork when something goes wrong. If your work is comparing published algorithms on standard benchmarks, the algorithm-first libraries are a better fit. If your work is building a training system that has to survive several of those settings, the component split is the point.

Licence, maintenance and the upgrade tax

TorchRL is MIT licensed, which is permissive and imposes no copyleft obligation on your own code; the LICENSE file is at the repository root. Nothing here is legal advice, and if you redistribute a modified build you should read the licence text yourself. On maintenance, the last push was on 2026-09-10 and the repository is not archived, so the project is being worked on, but the Beta classifier and the three releases between June and July 2026 tell you to expect movement. The upgrade tax is concentrated in one dependency: tensordict is pinned to >=0.14.2,<0.15.0, so a tensordict minor release can leave your TorchRL version outside the supported range. Pin both in your own lockfile. The second tax is the C++ extensions. The Makefile's clean target removes build/, dist/, egg-info directories, the compiled _torchrl*.so files and torchrl/version.py, which is the documented way to recover when you switch Python or PyTorch versions and the in-place build stops loading.

Editorial conclusion

Adopt TorchRL if your training loop already lives in PyTorch and you need one data model across vectorized, multiprocess, recurrent, multi-agent or offline workflows, because the collector, replay buffer and loss modules all read and write named TensorDict keys. Do not adopt it if you want a finished algorithm with a stable config schema and no interest in how data is structured, since the repository self-describes as Beta and the tutorial examples are not a supported API. Before committing, check that the installed tensordict version falls inside the >=0.14.2,<0.15.0 range TorchRL pins, and run one of the sota-implementations configs end to end on your own environment to see how much glue code the TensorDict keys actually remove.

Frequently asked questions

Is there a Python library for reinforcement learning?

Yes. TorchRL is a PyTorch-native reinforcement learning library published on PyPI as torchrl, and it requires Python 3.10 or newer along with torch>=2.1.0. It is built around TensorDict rather than around a single algorithm.

Is ChatGPT built on PyTorch?

The repository does not cover ChatGPT or any OpenAI product. TorchRL is a PyTorch-native library for reinforcement learning and decision making, and its only stated relationship to PyTorch is that it uses PyTorch tensors and nn.Module as its programming model.

Is TensorFlow losing to PyTorch?

Nothing in the repository compares framework adoption or market share, so this cannot be answered from it. What can be said is that TorchRL is a PyTorch-native toolkit, so choosing it means your training loop already runs on PyTorch.

Official sources

  1. License: MIT
  2. Project website
  3. pytorch/rl on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/pytorch-rl.svg)](https://hysenlabs.com/projects/pytorch-rl)