facebookresearch/ReAgent: an archived PyTorch platform for applied reinforcement learning
A platform for Reasoning systems (Reinforcement Learning, Contextual Bandits, etc.)
At a glance
- What is it?
- ReAgent bundles off-policy deep RL algorithms, contextual bandits and counterfactual policy evaluation into one Python platform built on PyTorch and TorchScript. The README states it is officially archived and points elsewhere, so the question is whether the code is still worth reading or reusing.
- Who is it for?
- ReAgent is worth adopting only as a reference implementation or as a starting point you are prepared to fork and maintain yourself; the README states it is officially archived and no longer maintained, and it points to Pearl for production-ready reinforcement learning. Teams that need a supported library, security fixes or a dependency upgrade path should not standardise on it.
- Can I use it commercially?
- Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 29 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem ReAgent was built to solve
ReAgent targets a specific setting: large-scale recommendation and optimisation tasks where no simulator exists. The README is explicit that the platform is designed for cases where you cannot interact with the environment freely, so the usual online RL loop of act, observe, update does not apply. Instead you train offline on batches of logged data and release new policies slowly over time. That constraint shapes everything else in the project. Because policy updates are slow and batched, the algorithm list leans heavily on off-policy methods: DQN variants, TD3, SAC, CRR and PPO, plus recommender-specific approaches such as Seq2Slate and SlateQ. The intended user is an applied engineer or research engineer at a company with a logging policy already in production, not someone learning RL from a textbook. Two supporting tools make that audience clearer: a Domain Analysis Tool that reports state and action feature importance and, per the README, identifies whether a problem is suitable for batch RL at all, and Behaviour Cloning, which clones the logging policy to bootstrap a learning policy safely. The suitability check is the more interesting of the two, because it is an admission that many problems should not be forced into this pipeline.
How the training and serving path fits together
The stack is Python for modelling and training, PyTorch for the models, and TorchScript for serving. That split matters: TorchScript export means a trained policy can be serialised and served without carrying the training dependency graph into production, which is the usual reason a research codebase is hard to deploy. The repository layout reflects the stages. reagent/ holds the library, preprocessing/ handles data preprocessing and feature transformation, serving/ covers optimised serving, and scripts/ holds the runnable entry points. There is also a rasp_requirements.txt at the top level, which suggests a separate dependency set for at least one component rather than a single monolithic install. The README lists the workflow components as data preprocessing, feature transformation, distributed training, counterfactual policy evaluation and optimised serving. Counterfactual evaluation is the piece that distinguishes this from a generic RL library: it provides Doubly Robust estimators for bandits and for sequential decisions, plus MAGIC, so a candidate policy can be scored against logged actions before anyone deploys it. The bandit algorithms (UCB1, MetricUCB, Thompson Sampling, LinUCB) sit alongside the deep methods, which means the same platform covers both the simple exploration problem and the neural one.
Installing ReAgent and running a first workflow
The README gives two installation routes, Docker or manual, and defers the details to docs/installation.rst in the repository. It does not reproduce the commands inline, so treat the docs directory as the source of truth rather than any snippet you find elsewhere. The repository also ships a pyproject.toml whose build backend is setuptools.build_meta, with setuptools and setuptools_scm declared as build requirements, and a setup.py that calls setup() and notes that configuration lives in config.cfg. That means the package is installed as a standard Python distribution, and the version is derived by setuptools_scm from the repository state rather than hardcoded.
git clone https://github.com/facebookresearch/ReAgent.git
cd ReAgent
pip install -r rasp_requirements.txtCloning and installing the requirements file is the manual path the repository layout implies; the README does not spell out this exact sequence, so confirm it against docs/installation.rst before relying on it. Once the package is importable, the README points to docs/usage.rst for how to use it, and says the platform is designed for offline batch training rather than simulator-driven loops. Expect the first real task to be data preparation rather than model training: preprocessing and feature transformation come before any algorithm runs. If your data is already in a batch format and you have a logging policy to clone, Behaviour Cloning is the documented way to bootstrap a learning policy safely before trying an off-policy algorithm.
The archive notice is the main limitation
The first line of the README states that ReAgent is officially archived and no longer maintained, and directs readers to Pearl for production-ready reinforcement learning. The repository metadata does not mark it as archived and records a last push on 2026-09-01, but the project's own documentation is unambiguous about its status, and that documentation is what a reviewer should trust. The practical consequences are ordinary ones: no upstream fixes for dependency drift, no compatibility work as PyTorch moves forward, and no security response. A TorchScript serving path is only as good as the PyTorch version it was exported against, and nothing in the repository promises that will be kept current. The second limitation is scope. This is not a general-purpose RL toolkit for games or robotics; the README frames it around recommendation and optimisation without a simulator. If you have a simulator, the off-policy, slow-release design is a mismatch, and you would be fighting the architecture rather than using it. Third, the documentation is thin on operational detail. The README describes what the components do but does not document rollback, deployment procedures or failure recovery, so anyone running the serving path is on their own for those decisions.
ReAgent compared with Stable-Baselines3 and Pearl
Stable-Baselines3 is the obvious alternative for someone who wants to train an RL agent in Python today. The difference in approach is the target setting. Stable-Baselines3 is built around environments and the standard gym-style interaction loop, which makes it the right tool when you can step an environment and collect fresh experience. ReAgent assumes the opposite: no simulator, logged batches, off-policy learning and counterfactual evaluation before deployment. If you have an environment, Stable-Baselines3 fits without any adaptation; if you have only logs from a production policy, ReAgent's counterfactual evaluation and bandit algorithms address a problem Stable-Baselines3 does not set out to solve. Pearl, which the ReAgent README itself names as the successor for production-ready reinforcement learning, is the more direct comparison, since it comes from the same applied RL team at Meta. The README gives no migration guide from ReAgent to Pearl, so the two should be evaluated separately rather than treated as a drop-in replacement. For bandit work specifically, ReAgent's inclusion of UCB1, Thompson Sampling and LinUCB alongside the deep algorithms is broader than what a deep-RL-only library offers.
Maintenance cost and licence terms
Adopting ReAgent means adopting its maintenance. The README states the project is archived and no longer maintained, so any dependency upgrade, PyTorch compatibility fix or bug fix becomes your work. The repository carries a CircleCI configuration and a codecov configuration, but a CI pipeline that no longer receives upstream attention tells you the code once built, not that it builds now. Budget for pinning dependencies and for verifying that the Docker path in docs/installation.rst still resolves. On licensing, ReAgent is released under BSD 3-Clause, a permissive licence that allows commercial use and modification provided the copyright notice and licence text are retained. The README links to the LICENSE file for the full terms and to Meta's Terms of Use and Privacy Policy pages. This is a description of what the repository states, not legal advice; if you plan to redistribute a modified version or ship it inside a product, have your own counsel read the LICENSE file and the terms pages it links to. The citation block in the README points to the 2018 Horizon paper on arXiv, which is the reference to use if you need to describe the algorithms in academic work.
Editorial conclusion
ReAgent is worth adopting only as a reference implementation or as a starting point you are prepared to fork and maintain yourself; the README states it is officially archived and no longer maintained, and it points to Pearl for production-ready reinforcement learning. Teams that need a supported library, security fixes or a dependency upgrade path should not standardise on it. Before committing, read docs/installation.rst and docs/usage.rst in the repository, confirm which of the listed algorithms your workflow actually needs, and check whether the Docker path still builds against current PyTorch releases.
Frequently asked questions
Is facebookresearch/ReAgent still maintained?
No. The README states that ReAgent is officially archived and no longer maintained, and points to Pearl for production-ready reinforcement learning. The repository itself is not marked as archived and its last push was on 2026-09-01, but the project's own documentation is the clearer signal.
What algorithms does ReAgent support?
The README lists off-policy deep RL methods including DQN variants, C51, QR-DQN, TD3, SAC, CRR and PPO, recommender-specific methods Seq2Slate and SlateQ, counterfactual evaluation with Doubly Robust and MAGIC, and bandit algorithms UCB1, MetricUCB, Thompson Sampling and LinUCB.
How do I install ReAgent?
The README says ReAgent can be installed via Docker or manually, and that detailed instructions are in docs/installation.rst in the repository. The build is a standard setuptools distribution, with pyproject.toml declaring setuptools and setuptools_scm as build requirements.
What licence does ReAgent use?
ReAgent is released under the BSD 3-Clause licence, and the README links to the LICENSE file for the full text. That is a permissive licence, but it is not legal advice and the terms pages linked from the README should be read directly.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/facebookresearch-reagent)