OpenRLHF: a Ray and vLLM RLHF framework for PPO, GRPO and agentic RL
An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)
At a glance
- What is it?
- OpenRLHF splits RLHF training across a Ray controller, vLLM rollout engines and DeepSpeed policy workers, and now exposes the same pipeline for multi-turn agents. Here is what the repository documents, what it leaves open, and who should stay away.
- Who is it for?
- Adopt OpenRLHF if you already run multi-GPU or multi-node training and need PPO, GRPO or REINFORCE++ over a vLLM rollout engine, or if you are training a multi-turn agent whose reward comes from an environment rather than a static dataset. Do not adopt it for single-GPU experimentation or if you cannot pin deepspeed==0.19.6, transformers==5.15.0 and ray[default]==2.55.0 inside a clean environment.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 13 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap OpenRLHF fills: RLHF that outgrows one GPU
The target user is a team that has already fine-tuned a language model and now wants to align it with reinforcement learning, but whose model or rollout budget no longer fits on a single device. The README frames the project as a production-ready RLHF framework combining a Ray plus vLLM distributed architecture with an agent-based execution design. In practice that means the generation step, which is usually the slowest part of an RLHF loop, runs on separate vLLM workers while policy updates run on DeepSpeed workers, all coordinated by Ray. The repository ships examples/scripts for training and examples/python for agent definitions, so the intended workflow is script-driven rather than notebook-driven. If your alignment work is a LoRA experiment on one 24 GB card, this is more machinery than you need.
How Ray, vLLM and DeepSpeed split the work
The architecture is a single-controller design. Ray holds the placement group and the controller process, vLLM serves rollouts, and DeepSpeed holds the trainable policy and reference weights. The README states that the source was refactored around a single controller and unified packing of samples, which is the part that matters operationally: samples from different prompts and different agent turns are packed together for the forward and backward passes instead of being padded per sequence. The algorithm surface is exposed through config keys rather than separate entry points. FlashREINFORCE, added in September 2026, is described as a composition of configs: --algo.advantage.estimator flash_reinforce, a binary-KL trust region on the vLLM logprobs, and sample-mean aggregation. That is a useful signal about the design philosophy. New algorithms arrive as config combinations, not forks.
Installing OpenRLHF and running a first PPO job
The repository does not carry a long install section in the README excerpt; the documentation site at openrlhf.readthedocs.io is where the project points readers. What the repository does pin is the dependency set. requirements.txt fixes deepspeed==0.19.6, transformers==5.15.0, ray[default]==2.55.0 and flash-attn==2.8.3, and setup.py reads version.txt for the package version. Install into a fresh environment so those pins are not silently upgraded.
pip install -r requirements.txt
pip install -e .The first command resolves the pinned stack; the second installs the local openrlhf package in editable mode so edits under openrlhf/ take effect without reinstalling. Note that setup.py marks the wheel as platform-specific (root_is_pure = False) and tags Linux builds as manylinux1_x86_64, so the package is not a pure-Python wheel.
For a first real run, the examples/scripts directory is the entry point. The README links train_prorlv2_math_hybrid_engine.sh as the script used for the ProRL V2 math training described in the news section, and train_reinforce_baseline_ray_agent_async.sh as a runnable async agent example. Open one of those scripts and read the argument list before launching it, because the algorithm, the rollout engine and the data paths are all set there rather than in a config file.
bash examples/scripts/train_prorlv2_math_hybrid_engine.shExpect Ray to start a cluster, vLLM workers to load the model for generation, and DeepSpeed workers to load the policy for training. On a first run the failure you are most likely to hit is a model or dataset path that the script assumes and your machine does not have.
Agentic RL and multi-turn environments
The agent layer is what separates this project from a plain PPO trainer. The README documents --train.agent_func_path for async agent RLHF and --train.async_enable for async training, both introduced in version 0.8.0, with train_reinforce_baseline_ray_agent_async.sh as a runnable example. Version 0.10 added multi-turn VLM RL, where images appear both in the prompt and in environment feedback such as screenshots, with examples/python/vlm_multiturn_agent.py as the reference. The single-turn path is simpler: a custom reward function scores completions without an environment loop. The trade-off is that the agent path couples your training job to the reliability of whatever environment you wire in, and the README does not document a retry or timeout policy for environment calls. If your environment is flaky, that flakiness becomes training instability.
Where OpenRLHF is the wrong tool
The dependency pins are the sharpest limitation. deepspeed==0.19.6, transformers==5.15.0 and ray[default]==2.55.0 are exact, not minimums, and flash-attn==2.8.3 requires a matching CUDA toolchain at build time. If your cluster image already ships a different DeepSpeed or Transformers version for another workload, you are choosing between a separate environment and a version conflict. Second, the project assumes a Ray cluster. There is no documented single-process fallback for debugging a reward function, and the README does not describe one. Third, the README does not document rollback or checkpoint-resume semantics for a crashed async run, so a long agentic job interrupted mid-flight is an open question you should test before trusting it. Finally, if your goal is offline preference optimization rather than online RL, the machinery here is a poor fit; you would be paying for vLLM rollout workers you never use.
OpenRLHF vs verl and TRL: different centres of gravity
The most common comparison in search data is OpenRLHF vs verl. Both are Ray-based RLHF stacks with vLLM rollouts, so the difference is not the topology but the algorithm surface and the agent story. OpenRLHF's identity is bound to REINFORCE++ and REINFORCE++-baseline, which the README credits as the method behind ProRL V2 and which Logic-RL and PRIME are cited as showing to be more stable than GRPO and faster than PPO. Its agent support is expressed through --train.agent_func_path and async flags. TRL is the other end of the spectrum: a Hugging Face library that keeps training in-process and does not require Ray or a separate rollout server, which makes it far easier to start and far harder to scale past a single node. The honest summary is that OpenRLHF trades setup complexity for distributed scale and an agent execution model; TRL trades scale for a shorter path from pip install to a running trainer. Pick based on which of those two you are actually short of.
Maintenance, licensing and the cost of upgrading
The repository is not archived, and the last push was on 2026-09-17, four days before this writing. Releases are frequent: v0.11.0 on 2026-08-13, v0.11.1 on 2026-09-10, v0.11.2 on 2026-09-14. That cadence is the upgrade cost. With exact pins on deepspeed, transformers and ray, a minor release can move any of those three, and the README's news entries show features landing as config-level additions rather than deprecations, so old scripts may keep working while the dependency set underneath them shifts. Budget a re-run of your training script on every bump. A nightly channel exists: setup.py publishes openrlhf-nightly when OPENRLHF_BUILD_MODE=nightly, appending a .devYYYYMMDD suffix to the version from version.txt. The licence is Apache-2.0, which permits commercial use and modification and requires that you preserve the licence and notice files and state significant changes; the repository does not include a NOTICE file at the top level, so if you redistribute, check the licence text in LICENSE rather than assuming a notice file exists.
Editorial conclusion
Adopt OpenRLHF if you already run multi-GPU or multi-node training and need PPO, GRPO or REINFORCE++ over a vLLM rollout engine, or if you are training a multi-turn agent whose reward comes from an environment rather than a static dataset. Do not adopt it for single-GPU experimentation or if you cannot pin deepspeed==0.19.6, transformers==5.15.0 and ray[default]==2.55.0 inside a clean environment. Before committing, verify your GPU and CUDA combination against the flash-attn==2.8.3 build in requirements.txt, then run one of the examples/scripts training scripts end to end on your own cluster.
Frequently asked questions
What is OpenRLHF and what algorithms does it support?
It is a Ray-based RLHF framework that pairs a vLLM rollout engine with DeepSpeed policy training. The README lists PPO, REINFORCE++, GRPO and RLOO, and the September 2026 news entry adds FlashREINFORCE as a config composition.
How does OpenRLHF compare with TRL?
TRL is a Hugging Face library that runs training in-process without Ray or a separate rollout server, while OpenRLHF requires a Ray cluster with vLLM rollout workers and DeepSpeed policy workers. The trade is a longer setup path for distributed scale and an agent execution model.
How do I install OpenRLHF?
The repository pins its stack in requirements.txt, including deepspeed==0.19.6, transformers==5.15.0, ray[default]==2.55.0 and flash-attn==2.8.3. Install those into a fresh environment and then install the local package in editable mode; the README points to openrlhf.readthedocs.io for the full documentation.
Does OpenRLHF support multi-turn agent training?
Yes. The README documents --train.agent_func_path for async agent RLHF and --train.async_enable for async training, both from version 0.8.0, with train_reinforce_baseline_ray_agent_async.sh as a runnable example. Version 0.10 added multi-turn VLM RL with images in prompts and environment feedback.
What does --algo.advantage.estimator flash_reinforce do?
The September 2026 news entry describes FlashREINFORCE as critic-free, single-rollout asynchronous RL for agentic models, configured as a composition of --algo.advantage.estimator flash_reinforce, a binary-KL trust region on the vLLM logprobs, and sample-mean aggregation. The training script is train_flash_reinforce_ray_agent_async.sh.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/openrlhf-openrlhf)