alibaba/ROLL: a Ray-based RL library for LLM post-training
An Efficient and User-Friendly Scaling Library for Reinforcement Learning with Large Language Models
At a glance
- What is it?
- ROLL is Alibaba's Apache-2.0 library for reinforcement learning with large language models, built on a multi-role Ray architecture with Megatron-Core, SGLang and vLLM backends. It is aimed at teams that already have multi-GPU clusters and want RLVR, DPO, distillation or agentic training pipelines without writing the distributed plumbing themselves.
- Who is it for?
- ROLL fits teams that already run multi-GPU training and want RLVR, agentic or distillation pipelines in one framework with example configs per model size. It is the wrong first stop for a single-GPU experiment: the example set is built around Megatron or FSDP2 clusters, and the requirement files are split per torch and inference-engine combination.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 3 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem ROLL addresses: post-training RL without rebuilding the cluster glue
Training a language model with reinforcement learning is not a single script. It is at least three workloads that must run at once: an inference engine generating rollouts, a training engine consuming them, and some orchestrator moving data and weights between the two. Most published research code solves this by pinning one combination, for example vLLM for generation and FSDP for training, and leaving the reader to port it. ROLL takes the other position. It treats the generation engine, the training engine and the resource scheduler as separately configurable roles, and ships example configurations for several model sizes so a new user starts from a known-good set of values rather than from a blank file.
The intended audience is visible in the repository layout. There are example directories for Qwen2.5 at 0.5B, 3B and 7B, for Qwen3 at 8B, 30B and 235B, and for Qwen3-Next at 80B, split between Megatron and FSDP2 training backends. That is a range from a laptop-scale model to something that needs a real cluster, and the presence of a 0.5B agentic example is the clearest signal that the project expects people to try the pipeline small before scaling it. The topics listed on the repository are agentic, rlhf and rlvr, which matches the three training modes the README highlights: human preference alignment, reasoning, and multi-turn tool use.
Multi-role distributed architecture: Ray schedules, Megatron and SGLang or vLLM execute
The README describes a multi-role distributed architecture with Ray for resource allocation and heterogeneous task scheduling. In practice that means the roles are separate actors: the inference engine that produces rollouts, the training engine that updates weights, and whatever reward or environment component scores the output. Ray decides where each role lands on the cluster, which is what allows one job to mix a vLLM rollout worker and a Megatron training worker on the same GPU pool.
On the training side, the repository shows two strategies. Megatron-Core appears in the majority of example directories, and the March 2026 release notes add FSDP2 as a second strategy alongside Megatron with LoRA support. That matters because it changes what hardware and expertise the project demands: Megatron paths assume tensor and pipeline parallelism knowledge, while FSDP2 paths are closer to standard PyTorch distributed training. On the inference side, the README names both SGLang and vLLM, and the requirements files are split accordingly, with separate files for torch 2.6 and torch 2.8 against each engine. There is also a dedicated requirements file for AMD and one for Ascend NPU, which the September 2025 news entry confirms as a supported platform.
The architecture is not free. Splitting generation and training into separate roles means weights have to be synchronized between them, and the project has published work on that cost: the RollPacker paper is described as mitigating long-tail rollouts for synchronous RL post-training, and a later paper covers asynchronous RLVR and agentic training. Those papers are the honest signal here. If weight synchronization and rollout latency were trivial, there would be no paper about them.
Installing alibaba/ROLL and running a first RLVR example
The repository does not put a single install command in the README. Installation is split by torch version and inference backend through the requirements files at the top level, so the first real decision is which one applies to your environment. The naming is explicit: requirements_torch260_vllm.txt pairs torch 2.6 with vLLM, requirements_torch280_sglang.txt pairs torch 2.8 with SGLang, and there are separate files for AMD and for the vision and diffusion paths.
Start by cloning the repository and picking the requirements file that matches your stack. The setup.py declares python_requires of 3.10 or later and the package name roll, so the install itself is a standard setuptools invocation once the dependencies resolve.
git clone https://github.com/alibaba/ROLL.git
cd ROLL
pip install -r requirements_torch280_vllm.txt
pip install -e .The editable install picks up the roll package declared in setup.py. After that, the practical entry point is one of the example directories rather than a bare command line. Each example carries its own YAML configuration, for example the Qwen3 8B RLVR configuration under examples/qwen3-8B-rlvr_megatron/ or the FSDP2 variant under examples/qwen3-8B-rlvr_fsdp2/. The configuration file is where the model path, the training backend, the rollout engine and the parallel sizes are set, so the workflow is: copy an example directory, edit the YAML for your model and cluster, then launch it.
For a smaller starting point, examples/qwen2.5-0.5B-agentic/ is the lowest-parameter example in the tree and the most sensible place to confirm that Ray, the inference engine and the training backend all come up before moving to a 30B or 235B configuration. The Makefile exposes the test suite and the pre-commit hooks, which is a quick way to check that the environment is coherent.
python -m pytest -n auto --dist=loadfile -s -v ./tests/Where ROLL is the wrong tool
The dependency surface is the first constraint. There is no single requirements.txt. The project maintains separate files for torch 2.6 and torch 2.8, for vLLM and SGLang, for AMD, for Ascend and for vision workloads. That is a reasonable response to real version conflicts, but it also means an environment that does not match one of those combinations is unsupported by construction. If your cluster is pinned to a torch version outside the maintained set, or to an inference engine the project has not paired with it, you are on your own.
The second constraint is scale. The example set is built around Megatron and FSDP2, and the model sizes range up to 235B. Nothing in the repository suggests a single-GPU path, and the architecture assumes Ray is scheduling across a pool. A researcher with one GPU who wants to run GRPO on a 7B model will find more friction here than in a library designed around a single process.
The third is documentation depth. The README is a news feed with a feature list; the substantive material lives in the docs site and in the example YAML files. The README does not document rollback, does not describe failure recovery for a partially completed run, and does not state what happens when a rollout worker dies mid-generation. Those are the questions that matter once a job has been running for hours, and the top-level README is silent on them.
How ROLL differs from verl, OpenRLHF and TRL
The closest comparisons are verl, OpenRLHF and TRL, and the difference is mostly in what each one assumes about your cluster. TRL is a library of trainers that runs inside a single process or a standard accelerate launch; it is the right choice when the model fits on the hardware you have and you want the training loop to look like ordinary PyTorch. OpenRLHF packages the actor, critic and reward models into a Ray-based distributed setup, which is closer in spirit to ROLL, but it is oriented around the PPO family with vLLM for generation.
ROLL's distinguishing choice is the breadth of the training backend. Megatron-Core and FSDP2 are both first-class, and the example tree shows the same model family configured twice, once for each. That is unusual. It means a team that has already invested in Megatron parallelism can keep it, and a team that prefers PyTorch-native sharding can use FSDP2 without switching frameworks. The second difference is the agentic emphasis: the repository carries examples for multi-turn tool use and the README notes alignment with the GEM environment definition, which puts environment interaction inside the training loop rather than bolted on as a reward function. If your work is single-turn preference optimization on a modest model, that extra machinery is overhead you will pay for and not use.
Maintenance, licence and what an upgrade costs
The repository is not archived and the last push was on 2026-09-23. The release cadence visible in the release list runs v0.2.0 in February 2026, v0.2.1 in March 2026 and v0.3.0 in June 2026, with the news feed showing feature additions between releases. That is a project that is moving, and moving projects in this space tend to move their dependency pins with them.
The upgrade cost follows from the requirements layout. Because the dependency sets are split by torch and inference engine, a torch upgrade is not a version bump in one file; it may mean moving to a different requirements file entirely and revalidating the combination. The version in setup.py is 0.3.0, matching the latest release tag, so the package version and the release tag are kept in step. The practical implication is that pinning to a release tag is safer than tracking main if you need a reproducible environment, and the Makefile's pre-commit target is the cheapest way to see whether your local edits still pass the project's own lint rules.
The licence is Apache-2.0, stated in the LICENSE file and in the badge at the top of the README. Apache-2.0 is permissive and includes an explicit patent grant, which matters for a library that may end up inside a commercial training stack. It also means redistributed modifications must carry the licence and notice files. This is a description of the licence text, not legal advice; the terms that apply to a specific deployment depend on how the software is used and redistributed, and that is a question for counsel rather than for a README.
Editorial conclusion
ROLL fits teams that already run multi-GPU training and want RLVR, agentic or distillation pipelines in one framework with example configs per model size. It is the wrong first stop for a single-GPU experiment: the example set is built around Megatron or FSDP2 clusters, and the requirement files are split per torch and inference-engine combination. Before committing, verify which requirements file matches your torch version and inference backend, and check that an example directory exists for your model family, since the README only lists configurations for the models the project has already validated.
Frequently asked questions
What is alibaba/ROLL?
ROLL is a reinforcement learning library for large language models, described in its README as an efficient and user-friendly scaling library for RL with LLMs. It uses a multi-role distributed architecture built on Ray and integrates Megatron-Core, SGLang and vLLM for training and inference.
How do I install alibaba/ROLL?
There is no single install command in the README. You clone the repository, install one of the version-specific requirements files such as requirements_torch280_vllm.txt or requirements_torch280_sglang.txt, then run pip install -e . to install the roll package declared in setup.py.
Which models does alibaba/ROLL have example configurations for?
The examples directory covers Qwen2.5 at 0.5B, 3B and 7B, Qwen3 at 8B, 30B and 235B, and Qwen3-Next at 80B, with separate directories for Megatron and FSDP2 training backends. Qwen3.5 Dense and MoE configurations are also listed in the June 2026 news entry.
What licence does alibaba/ROLL use?
The repository is licensed under Apache-2.0, as stated in the LICENSE file and the licence badge at the top of the README.
Does alibaba/ROLL support hardware other than NVIDIA GPUs?
The requirements files include a dedicated AMD variant and an Ascend NPU variant, and the September 2025 news entry announces Ascend NPU support with a linked usage guide. The README does not describe the performance characteristics of either path.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/alibaba-roll)