RLtools: a header-only C++ deep RL library for control and TinyML
The Fastest Deep Reinforcement Learning Library
At a glance
- What is it?
- RLtools compiles deep reinforcement learning algorithms into plain C++ binaries, so training runs on a laptop CPU and the resulting policy can be deployed to a microcontroller. It is aimed at control and robotics engineers, not at Python-first research workflows.
- Who is it for?
- Adopt RLtools if you are training continuous-control policies in C++ and want the same code to run on a microcontroller, or if you need a single g++ command instead of a Python environment.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 15 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem RLtools solves: training and deployment in one language
Most deep reinforcement learning stacks are Python. You install a framework, an environment suite and a simulator, then train a policy and export it. Getting that policy onto a robot or a microcontroller means a second implementation, usually in C or C++, and a conversion step where the network weights, the observation normalization and the action scaling have to be reproduced exactly. RLtools removes that second implementation. The library is C++, and the README states that the same code paths cover training and inference, including inference on microcontrollers, which the repository lists as a topic alongside tinyrl and tinyml.
The intended user is an engineer working on continuous control: quadrotors, pendulum swing-up, acrobot, racing cars, MuJoCo Ant. The README's own examples are drawn from that world, and the projects based on RLtools are drone papers, including Learning to Fly in Seconds and a quadrotor system-identification paper. If your problem is a discrete-action game or a language-model fine-tune, this is not the library for you. The algorithm table lists SAC, TD3, PPO and Multi-Agent PPO, and everything in the table is continuous control or multi-agent control.
How RLtools is structured: headers, backends and per-environment entry points
The repository layout tells you most of the architecture. There is an include/ directory and a src/ directory, plus CMakeLists.txt and a cmake/ folder. The quick start compiles a single .cpp file with -I include, which is the signature of a header-only library: the algorithms and the environments live in headers under include/, and src/ holds example programs and environment definitions rather than a library to link against.
Environment code is organized by algorithm and by device. The README's algorithm table links to paths such as src/rl/environments/pendulum/td3/cpu/standalone.cpp, src/rl/environments/mujoco/ant/ppo/cpu/training.h and src/rl/environments/mujoco/ant/ppo/cuda/training_ppo.cu. That directory pattern, environment then algorithm then device, is the data flow in practice: you pick an environment, pick an algorithm, and pick CPU or CUDA by choosing a file. Numerical work is delegated to a backend selected at compile time. The README names two: -DRL_TOOLS_BACKEND_ENABLE_ACCELERATE on macOS and -DRL_TOOLS_BACKEND_ENABLE_OPENBLAS on Ubuntu. There is also a src/rl/zoo/ tree holding named examples such as l2f and bottleneck-v0, which is what the quick start builds.
The README does not document the internal tensor or memory model, so anyone who wants to add a new backend has to read the headers. That is the cost of the header-only design.
Installing RLtools and running the learning-to-fly example
There is no package manager step in the README. You clone the repository and compile. The quick start builds a Zoo example, the learning-to-fly SAC quadrotor, with a single g++ invocation that adds include/ to the search path.
g++ -std=c++17 -O3 -ffast-math -I include src/rl/zoo/l2f/sac.cppThe result is ./a.out. The README then says to run it with a seed number and to start the visualization tool.
./a.out 1337
./tools/serve.shAfter ./tools/serve.sh runs, open http://localhost:8000 in a browser and navigate to the ExTrack UI to watch the quadrotor flying. The README does not document how long training takes on arbitrary hardware, but it does give two platform-specific builds that enable a fast linear algebra backend. On macOS, append -framework Accelerate -DRL_TOOLS_BACKEND_ENABLE_ACCELERATE; the README quotes roughly 4s on an M3. On Ubuntu, install the backend first and then compile against it.
sudo apt install libopenblas-dev
g++ -std=c++17 -O3 -ffast-math -I include src/rl/zoo/l2f/sac.cpp -lopenblas -DRL_TOOLS_BACKEND_ENABLE_OPENBLASThe README quotes roughly 6s on Zen 5 for that build. Those are the project's own numbers for one example, not a general benchmark, and they will not transfer to a different environment or a larger network.
Where RLtools stops being the right tool
The environment story is the sharpest limitation. The README's Getting Started section says that to implement your own environment you should refer to the separate rl-tools/example repository. In other words, the library does not ship an adapter layer for the environment suites most people already have. If your task is defined in Python, wrapping it for RLtools means either reimplementing the dynamics in C++ or building a bridge, and the README does not describe a bridge.
Algorithm coverage is the second boundary. SAC, TD3, PPO and Multi-Agent PPO cover a large share of continuous-control work, but they are not the whole field. Offline RL, model-based methods, and the long tail of variants that exist as Python reference implementations are absent from the README's table. If you need to reproduce a specific paper's algorithm, you will be writing it yourself in headers.
CUDA support is partial rather than uniform. The table lists MuJoCo Ant PPO and Pendulum SAC with CUDA variants, but the other entries are CPU-only. There is no statement in the README that every algorithm has a GPU path, so plan on CPU training unless your exact combination appears in the table.
The maintenance picture is worth stating plainly. The repository is not archived, and the last push was on 2026-07-04. The most recent release listed is v2.2.0 from 2025-10-23, with v2.1.0 shortly before it and v2.0.0 back in 2024-11-19. The README documents no deprecation or migration policy, so upgrading across a major version means reading the release notes yourself.
RLtools compared with Stable-Baselines3 and Gymnasium-based stacks
The obvious alternative for most readers is Stable-Baselines3 on top of Gymnasium in Python. The difference is not speed in the abstract; it is where the boundary between training and deployment falls. Stable-Baselines3 gives you a wide set of algorithms, an environment interface that thousands of existing tasks already implement, and a Python workflow that is easy to instrument. It also gives you a policy that has to be exported and reimplemented before it runs on a robot.
RLtools inverts that. The environment interface is C++ and you write it, but the trained policy is already C++ and the README's microcontroller inference benchmarks suggest deployment is a first-class target rather than an afterthought. The trade is ecosystem for deployment path. If your deliverable is a paper or a simulation result, Python wins on convenience. If your deliverable is firmware that runs a control policy at a fixed rate on a small device, the C++ side is where you want to be, and RLtools is one of the few libraries whose README treats that as the main use case.
A second alternative is writing the algorithm yourself. For a single algorithm on a single environment, that is often less work than learning any library. RLtools becomes worth it when you want several algorithms, a backend abstraction for BLAS, and the Zoo examples as reference implementations to compare against.
Licence and the cost of keeping up
RLtools is MIT licensed. That is permissive: you can use it in commercial and closed-source products, and you can modify it, provided you keep the copyright notice and the licence text. It is not a copyleft licence, so it does not force you to publish your own environment code. The repository ships a LICENSE file at the top level. This is a description of the licence identifier, not legal advice; if the licence terms matter to your organization, read the LICENSE file and get your own review.
Upgrade cost is harder to estimate from the README alone. The version history shows a major bump from v2.0.0 to v2.1.0 and then v2.2.0 within about a month, which suggests the API is still moving. Because the library is header-only, an upgrade is a recompile rather than a relink, and any breaking change in a header surfaces at build time across every target you compile. That is the practical cost: there is no shared object to pin, so you pin the repository commit instead. The README does not describe a compatibility guarantee between minor versions.
Editorial conclusion
Adopt RLtools if you are training continuous-control policies in C++ and want the same code to run on a microcontroller, or if you need a single g++ command instead of a Python environment. Do not adopt it if your work depends on Python-native environments, Gymnasium wrappers, or a large ecosystem of third-party algorithm implementations, because the README points to rl-tools/example for writing environments and the documented algorithm set is SAC, TD3, PPO and Multi-Agent PPO. Before committing, verify the following in order: build the l2f SAC example on your own machine with the OpenBLAS or Accelerate flag for your platform, confirm the ExTrack visualization loads at http://localhost:8000 after running ./tools/serve.sh, and check whether the environment you care about already exists under src/rl/environments or has to be written from scratch.
Frequently asked questions
What is RLtools used for?
RLtools is a C++ deep reinforcement learning library aimed at continuous control and robotics. The README's examples cover pendulum swing-up, acrobot, racing cars, MuJoCo Ant and a learning-to-fly quadrotor, and the repository also targets microcontroller inference.
Which algorithms does RLtools implement?
The README's algorithm table lists TD3, PPO, Multi-Agent PPO and SAC, with example paths for each. Some entries have CUDA variants, such as MuJoCo Ant PPO and Pendulum SAC, while others are CPU only.
How do I install and build RLtools?
There is no package manager step; you clone the repository and compile. The quick start builds a Zoo example with g++ -std=c++17 -O3 -ffast-math -I include src/rl/zoo/l2f/sac.cpp, and the README gives platform flags for the Accelerate backend on macOS and the OpenBLAS backend on Ubuntu.
How do I visualize a trained RLtools policy?
The README says to run the compiled binary with a seed, then run ./tools/serve.sh, open http://localhost:8000 and navigate to the ExTrack UI. That is the flow shown for the learning-to-fly quadrotor example.
Can I add my own environment to RLtools?
The README's Getting Started note directs readers who want to implement their own environment to the separate rl-tools/example repository rather than documenting the interface inline. Expect to read the headers and that example repository.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/rl-tools-rl-tools)