# rLLM: Agentic RL for Language Agents, From Rollout to Weight Update

> rLLM is an Apache-2.0 Python framework that runs your existing agent code as a rollout, captures token IDs and logprobs through a model gateway, and feeds the traces into GRPO or REINFORCE training on verl, tinker or Fireworks. It is aimed at teams who already have an agent and want to train it, not at people looking for a finished model.

**rllm-org/rllm** — Democratizing Reinforcement Learning for LLMs

- Repository: https://github.com/rllm-org/rllm
- Website: https://docs.rllm-project.com
- Stars: 5,826 · Forks: 617
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/rllm-org-rllm

## The problem rLLM targets: your agent and your trainer speak different languages

Most agent code is written for inference. It calls an OpenAI-compatible endpoint, gets text back, and moves on. Reinforcement learning needs more than text: it needs the token IDs and logprobs that produced each completion, grouped by the task that generated them, so an advantage can be computed per trajectory. Teams that want to train an agent usually end up maintaining two versions of it, one for evaluation and one instrumented for training, and the two drift.

rLLM's answer is a decorator and a gateway. The README states that the same agent code drives both eval and training, and that during training `config.base_url` points to a gateway that transparently captures token IDs and logprobs. The intended user is a research engineer or a small lab team with a working agent and a reward they can express as a function. It is not a product for someone who wants to download a trained model. The repository's topics list places it alongside verl, SWE-agent and search-agent work, which matches that audience.

## How rLLM works: Episodes, Trajectories, Steps and a URL-routed gateway

The README describes the pipeline as run your agent, collect traces, compute rewards, update the model. The data model is three levels deep. An Episode is one task. Inside it sit Trajectories, one per agent run. Inside those sit Steps, one per LLM call. The gateway captures token IDs and logprobs by URL-routed sessions and structures them into that shape, so the reward function receives an Episode with an artifacts dictionary rather than raw strings.

Under the hood the README names four components. A Workflow Engine runs N parallel agent instances to collect rollouts. The Model Gateway routes requests and captures token IDs and logprobs. A Transform Pipeline groups trajectories for advantage computation. A Training Backend, one of verl, tinker or fireworks, performs the policy update. The separation matters because the sandbox and the training backend are independent choices: Docker, Daytona, Modal or local on one side, distributed multi-GPU or single-machine on the other.

The README also mentions snapshot and warm-pool acceleration for keeping rollouts cheap at training scale, though it does not quantify the saving. Treat that as an unverified design claim rather than a measured result.

## Installing rLLM and running a first eval plus a first training job

rLLM requires Python 3.11 or newer according to the README, although the pyproject.toml sets `requires-python = ">=3.10"` and the tinker dependency is marked `python_version >= '3.11'`. The README is the stricter and more accurate guide here.

The default install pulls the tinker backend, which the README describes as single-machine via the Tinker API.

```bash
uv pip install "rllm @ git+https://github.com/rllm-org/rllm.git"
```

For distributed multi-GPU training with verl and vLLM or SGLang, the README gives a separate extra. This one drags in pinned versions of verl, vLLM, torch and flash-attn, so install it in a dedicated environment.

```bash
uv pip install "rllm[verl] @ git+https://github.com/rllm-org/rllm.git"
```

The CLI path needs no Python file at all. The README's quickstart is three commands: configure a model provider, evaluate on a benchmark, then train on the same benchmark. `rllm eval` auto-pulls the benchmark by name, so gsm8k does not need to be downloaded by hand.

```bash
rllm model setup
rllm eval gsm8k
rllm train gsm8k
```

If you would rather write the agent yourself, the README's Python API expects a rollout function decorated with `@rllm.rollout` that returns an Episode, and a reward function decorated with `@rllm.evaluator` that returns an EvalOutput carrying a reward and a list of signals.

```python
import rllm
from rllm.types import AgentConfig, Episode, Task, Trajectory

@rllm.rollout
def solve(task: Task, config: AgentConfig) -> Episode:
    answer = call_your_agent(task.instruction, config.base_url, config.model)
    return Episode(
        trajectories=[Trajectory(name="solver", steps=[])],
        artifacts={"answer": answer},
    )
```

The trainer then takes the rollout and the evaluator together. The README's example passes `backend="tinker"`, `agent_flow=solve`, `evaluator=score`, a config and a train dataset, then calls `trainer.train()`. The cookbooks directory is the place the README points to for complete working examples, including a single-turn VLM solver and a multi-agent solver-judge.

## Where rLLM gets in the way: version pinning, pre-releases and thin documentation

The distributed extra is the sharpest constraint. It pins `verl==0.8.0`, `vllm==0.22.1`, `flash-attn==2.8.3`, `torch>=2.10.0` and `transformers>=5.5.3`. If your cluster already runs a different vLLM or flash-attn build, the install will fight it, and the Dockerfile tells the same story from the other direction: it clones volcengine/verl, checks out v0.6.1 and installs it editable before installing rLLM. Two different verl versions appear in the repository, one in the extra and one in the Dockerfile, so you should not assume the container and the pip extra are interchangeable.

The release history is another signal. The most recent release listed is `v0.3.0-pre` from 2026-04-30, and the pyproject version is `0.3.0.pre`. A pre-release at the top of the tree is normal for a research framework, but it means the API you read about today is not frozen.

There is also a licence inconsistency worth noticing before you file anything internally: the repository description says Apache-2.0 and the LICENSE file is what the build reads, while the pyproject classifier still declares `License :: OSI Approved :: MIT License`. That classifier is a leftover. The README does not document rollback, does not describe checkpoint compatibility between the tinker and verl backends, and does not say what happens to in-flight rollouts when a gateway session is interrupted. If any of those matter to your workflow, the documentation is silent and you will be reading the source.

## rLLM versus verl alone, and versus a closed training platform

The closest comparison is running verl directly, since rLLM uses it as one of three backends. The difference is where the agent lives. With verl alone you write the rollout inside verl's own abstractions and your agent becomes part of the training code. With rLLM the agent stays a Python function that returns an Episode, and the gateway does the instrumentation by routing requests through a URL. That is the whole trade: you keep one agent implementation for eval and training, and in exchange you accept rLLM's gateway in the request path and its pinned dependency set.

Against a closed training platform such as the Fireworks path that rLLM itself supports, the difference is control over the environment. Fireworks is one of the three backends the README lists, so this is not either-or at the project level; the choice is whether rollouts run in your Docker, Daytona, Modal or local sandbox, or on someone else's infrastructure. The README's claim that you can switch backends with one flag is the part to test first, because the extras install very different dependency trees and the flag does not change what is on disk.

A third option is to skip RL entirely and do supervised fine-tuning on traces you collected by hand. That is cheaper and needs no gateway, but it cannot use a reward function, which is the reason to be here.

## Maintenance, upgrades and what the licence actually commits you to

The repository is not archived and the last push was on 2026-09-06, so the codebase is being touched. That is the only maintenance claim the repository supports. There is no published support window, no deprecation policy and no migration guide between 0.2.x and 0.3.0-pre, so an upgrade is a re-read of the release notes plus a re-run of your own eval, not a routine version bump.

The upgrade cost concentrates in the extras. Moving the verl extra forward means moving vLLM, flash-attn and torch with it, and those three are the usual source of build failures on a shared cluster. The tinker extra is far lighter: `tinker>=0.22.2` and `tinker-cookbook>=0.4.1`, both gated on Python 3.11 or newer. If you can develop against tinker and only move to verl for scale, you avoid most of the churn.

On licensing, the repository description states Apache-2.0 and the build reads the LICENSE file, while the pyproject classifier still says MIT. Apache-2.0 includes an explicit patent grant and requires attribution and notice retention, which matters if you redistribute a modified rLLM inside a product. This is a description of what the files say, not legal advice; if the classifier discrepancy affects a compliance decision, resolve it with the project rather than assuming either label.

## Conclusion

Adopt rLLM if you already have an agent loop written in Python and want to train it without rewriting that loop for training, and if you can accept that the distributed path pins exact versions of verl, vLLM and flash-attn. Skip it if you need a single stable release: the current line is 0.3.0-pre, the repository's own classifiers still say MIT while the LICENSE file is Apache-2.0, and the README does not document rollback or checkpoint compatibility between backends. Before committing, run the CLI quickstart end to end on gsm8k and confirm that your gateway URL, your sandbox choice and your backend all survive one full training run.

## FAQ

### What Python version does rLLM need?

The README states that rLLM requires Python 3.11 or newer. The pyproject.toml sets requires-python to >=3.10, but the tinker dependency is only installed on Python 3.11 and above, so 3.11 is the safer floor.

### Which training backends can rLLM use?

The README lists three: verl for distributed multi-GPU training, tinker for single-machine training through the Tinker API, and fireworks for the Fireworks platform. Each has its own install extra, and the README says you switch between them with one flag.

### How does rLLM capture token IDs and logprobs without changing my agent code?

During training, config.base_url points at a model gateway that routes requests by URL and captures token IDs and logprobs transparently, structuring them into Episodes, Trajectories and Steps. The README states that the same agent code therefore works for both eval and training.

### Does rLLM support Docker or local sandboxes for agent rollouts?

The README lists Docker, Daytona, Modal and local as sandbox options, with snapshot and warm-pool acceleration mentioned for rollouts at training scale. The repository also ships a Dockerfile that builds on the verl vLLM image and installs rLLM editable.

## Sources

- [License: Apache-2.0](https://github.com/rllm-org/rllm/blob/main/LICENSE)
- [Project website](https://docs.rllm-project.com)
- [README](https://github.com/rllm-org/rllm/blob/main/README.md)
- [Releases](https://github.com/rllm-org/rllm/releases)
- [rllm-org/rllm on GitHub](https://github.com/rllm-org/rllm)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/rllm-org-rllm
