# RLinf: a Ray-based RL training stack for embodied and agentic models

> RLinf is an Apache-2.0 Python framework that wires simulators, rollout workers and trainers into one Ray cluster. It is aimed at teams doing VLA fine-tuning, not at single-GPU experimentation.

**RLinf/RLinf** — RLinf: Reinforcement Learning Infrastructure for Embodied and Agentic AI.

- Repository: https://github.com/RLinf/RLinf
- Website: https://rlinf.readthedocs.io/en/latest/
- Stars: 5,374 · Forks: 751
- Language: Python
- License: Apache-2.0
- Published: 2026-08-08 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/rlinf-rlinf

## What RLinf actually coordinates

Most reinforcement learning code for language and vision-language models assumes one loop: generate samples, score them, update weights. Embodied training breaks that assumption. The sample source is a simulator or a physical robot, the policy may be a vision-language-action model with its own serving stack, and the update step may need a different parallel layout from the rollout step. RLinf exists to hold those pieces in one process graph instead of one script.

The README describes it as an open-source RL infrastructure for Embodied and Agentic AI, and the 'inf' is explicitly read as both Infrastructure and Infinite. The audience is narrow and identifiable: teams already training VLA policies or reasoning agents who have outgrown a single training script. The examples directory confirms this, splitting into agent/, diffusion/, embodiment/, offline_rl/, reasoning/, reward/ and sft/. A reader who only wants to fine-tune a small language model with GRPO has simpler options.

## Ray as the scheduler, Hydra as the config surface

The dependency list in pyproject.toml is the clearest statement of architecture. ray[default]>=2.47.0 and torch>=2.5.0 sit under a comment labelled Core System, described as the dependencies of the core scheduler. Ray is therefore the substrate: rollout workers, environment workers and trainers are Ray actors, and the topology of a run is a placement decision rather than a hardcoded process launch.

Configuration is Hydra. The pin is hydra-core<1.4.0.dev8, with a comment explaining that a later dev build rejects version_base="1.1" used elsewhere in the tree. That is a real constraint worth knowing before you upgrade Hydra for another reason.

Observation transfer between environment and rollout is compressed. The dependency block names lz4 and zstandard as codecs for env.obs_compression, which tells you the framework treats image observations crossing process boundaries as a bandwidth problem. Logging is pluggable across wandb, swanlab and tensorboard. The repository also carries ray_utils/ at the top level, which is where cluster and placement helpers live.

## Installing RLinf and running a first example

The project publishes under the name rlinf, and pyproject.toml requires Python 3.10 or newer. The README gives no pip install line; it points readers at the documentation site at rlinf.readthedocs.io, and the repository ships a docker/ directory and a .dockerignore alongside pyproject.toml. Setup therefore goes through either the Docker path or the documentation's installation page.

Once the package is installed, runs are launched through the example scripts rather than a single CLI entry point. The README points each feature at its own documentation page, for instance the Evo-1 example under the embodied docs, so the pattern is to pick the script matching your model and simulator. The repository layout puts those under examples/embodied/.

Configuration overrides follow Hydra conventions, so values are passed as key=value arguments on the command line. The environment variable env.obs_compression controls observation compression between environment and rollout, with lz4 and zstandard listed as the codecs. The README does not name the individual example script filenames; the documentation pages linked from each bullet in the What's NEW section do.

## Compression, parallelism and the BEHAVIOR numbers

The README reports a 25x end-to-end speedup for the BEHAVIOR simulator, with rollout latency falling from 1028.7 ms/step to 41.2 ms/step. Those figures come from a project blog post, not from an independent measurement, and the README attributes the gain to three named techniques: slimming, on-demand observation and hybrid pipeline parallelism.

The mechanism is consistent with the dependency list. On-demand observation means the environment does not materialise every camera frame the policy might need; slimming means unused state is dropped before it crosses a process boundary; hybrid pipeline parallelism means the rollout and update stages overlap rather than alternating. The lz4 and zstandard dependencies are the transport half of the same idea.

This is the part of RLinf that is genuinely differentiated. Simulator throughput, not GPU FLOPs, is usually the binding constraint in embodied RL, and a framework that treats observation serialisation as a first-class concern is making a defensible bet. Treat the 25x figure as a claim tied to one simulator and one configuration until you reproduce it on yours.

## Where RLinf is the wrong tool

The classifiers in pyproject.toml declare Development Status 2 - Pre-Alpha and a single environment target, NVIDIA CUDA 12.8. The README announces support for Moore Threads MUSA, Huawei Ascend CANN and AMD ROCm as of August 2026, but pyproject.toml still lists only the CUDA classifier. If you are on ROCm or Ascend, the packaging metadata will not tell you that; you have to read the platform guides.

The version number is also a moving target. pyproject.toml declares version 0.4.0 while the most recent tagged release in the repository is v0.3 from 2026-07-15. Anyone pinning to a release tag is pinning to something older than the source tree.

The sharper limitation is scope. RLinf assumes a cluster. Ray is a hard dependency, not an optional backend, and the whole design leans on distributing environment workers away from trainers. On one GPU with one environment, that machinery costs more than it returns. For plain LLM RLHF or GRPO on text, a single-process trainer with vLLM is less to install and less to debug. RLinf earns its complexity when the sample source is a simulator or a robot, and not before.

## How RLinf differs from verl and TRL

verl and TRL both target language-model post-training, and both are reasonable defaults for that job. The difference is what sits on the left side of the loop. In verl and TRL the sample source is a dataset of prompts and a generation engine, so the data path is tokens in, tokens out, and throughput work concentrates on the inference engine.

RLinf puts a simulator or a physical robot in that position. Isaac Lab v3.0.0 adopting RLinf as its RL training infrastructure, per the README, is the clearest illustration: the environment is a physics engine producing images and proprioception, and the policy is a VLA model rather than a text generator. That changes the bottleneck from KV-cache management to observation transport and step latency.

The overlap is real for agentic workloads. RLinf's examples/agent/ and examples/reasoning/ directories cover territory verl also covers, including a GRPO recipe for Moonlight-16B-A3B with DeepSeek-V3 MLA and MoE. If your work is purely text, the embodied half of RLinf is dead weight. If your work involves a robot, verl has no equivalent.

## Licence, releases and the upgrade bill

RLinf is Apache-2.0, with the licence file at the repository root and referenced from pyproject.toml. That permits commercial use and modification with the usual notice and patent terms. It is a permissive licence, so the practical question is not legal but operational: how much churn you absorb between releases.

Three releases are visible: v0.1 on 2025-12-17, v0.2 on 2026-03-26 and v0.3 on 2026-07-15. The last push to the default branch was 2026-07-15, the same timestamp as the v0.3 tag. The cadence is roughly quarterly, and the v0.3 notes describe upgrades across the real-world RL pipeline, additional simulators and system-level optimisations. A release that touches the pipeline end to end is a release that can move configuration keys.

Upgrade cost therefore concentrates in your Hydra configs and in the pinned dependency set. The Hydra pin is a live example: an upper bound on hydra-core exists specifically because a dev build broke the version_base used in the tree. Budget for reading the release notes before bumping anything.

## Conclusion

Adopt RLinf if you already have a multi-GPU cluster and want one codebase that spans LIBERO or Isaac Lab simulation, SFT, GRPO-style RL and real-robot deployment, and if you are willing to read the ReadTheDocs tree rather than the README. Do not adopt it for a single-GPU proof of concept or if you need a stable API: pyproject.toml still declares Development Status 2 - Pre-Alpha. Before committing, verify which accelerator path you need (the CUDA 12.8 classifier is the only one declared in pyproject.toml), confirm that the simulator you intend to use has an entry under examples/embodied/, and check the release notes for v0.3 to see whether the pipeline stage you need is already covered.

## FAQ

### What is RLinf?

RLinf is an open-source reinforcement learning infrastructure for embodied and agentic AI, written in Python and licensed under Apache-2.0. The README describes it as a scalable backbone for training, with the 'inf' standing for both Infrastructure and Infinite.

### What is RL Infra?

The question is close to asking what RLinf is: an RL training infrastructure rather than an algorithm library. Its core scheduler is built on Ray, with Torch as the computation backend and Hydra for configuration.

### Is RL a dead end?

The README does not address this, but it does document where RL is being applied: RLinf's examples cover embodied control, agentic reasoning, diffusion and video generation models, and reward modelling. Two of its associated papers were accepted to OSDI 2026 and two to RSS 2026.

### What are the four elements of reinforcement learning?

RLinf is an infrastructure project, and its documentation is organised around simulators, models and training recipes rather than reinforcement learning theory. The README does not cover textbook concepts.

### Is ChatGPT using reinforcement learning?

The README does not discuss ChatGPT. RLinf's agentic examples do include GRPO training for Moonlight-16B-A3B, which is a reinforcement learning method applied to a language model.

## Sources

- [Official documentation](https://rlinf.readthedocs.io/en/latest/)
- [Official README](https://github.com/RLinf/RLinf#readme)
- [Project repository](https://github.com/RLinf/RLinf)
- [Release notes](https://github.com/RLinf/RLinf/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/rlinf-rlinf
