Library / SDK
verl-project/verl avatar
verl-project/verl

verl: ByteDance's Reinforcement Learning Framework for Large Language Model Post-Training

verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework.

23,609 stars4,589 forksPythonApache-2.0

At a glance

What is it?
verl is an open-source RL training library initiated by ByteDance Seed and maintained by the broader verl community, designed to run post-training algorithms like PPO and GRPO on large language models at scale. It integrates with FSDP, Megatron-LM, vLLM, and SGLang and requires Python 3.10 to 3.12.
Who is it for?
verl suits ML teams that need to run RL post-training on large language models, particularly those who want to combine FSDP or Megatron-LM for training with vLLM or SGLang for rollout. The 3D-HybridEngine for actor model resharding is the main performance mechanism, and it requires GPU infrastructure; the recipe directory and the examples show concrete configurations for PPO, GRPO, and a range of other algorithms.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 4 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What verl Does and Who It Is For

Reinforcement learning post-training for large language models requires coordinating a training engine, a rollout engine, and a reward model across a distributed GPU cluster. verl addresses this by providing a flexible framework that decouples computation from data dependencies, allowing teams to plug in their preferred LLM infrastructure.

The README describes verl as a flexible, efficient, and production-ready RL training library for large language models. It is the open-source implementation of the HybridFlow paper (arXiv:2409.19256v2), which introduced the hybrid-controller programming model for managing complex post-training dataflows. The project was initiated by the ByteDance Seed team and is now maintained by the broader verl community. It is written in Python and released under the Apache-2.0 licence.

The Hybrid-Controller Model and 3D-HybridEngine

The core architectural concept in verl is the hybrid-controller programming model. The README describes it as enabling flexible representation and efficient execution of complex post-training dataflows, with RL algorithms like GRPO and PPO expressible in a few lines of code.

The 3D-HybridEngine is the performance-critical mechanism for managing the actor model during RL training. The README states it eliminates memory redundancy and significantly reduces communication overhead during transitions between training and generation phases. This is relevant to large-scale training where the actor model must be resharded between the training step and the generation (rollout) step.

The framework decouples computation and data dependencies to support integration with existing LLM frameworks including FSDP, Megatron-LM, vLLM, and SGLang. The setup.py lists Ray as a required dependency (ray[default]>=2.41.0), reflecting the distributed execution model.

Installing verl and Syncing Dependencies

verl uses uv for dependency management with a single universal uv.lock file. The pyproject.toml documents the approach: each backend is a PEP 621 extra and mutually exclusive ones are managed with uv conflicts. The recommended sync command for a FSDP and vLLM combination:

bash
uv sync --extra fsdp --extra vllm

Or using the included management script:

bash
python manage_envs.py sync fsdp vllm

The requirements.txt lists the core runtime dependencies: accelerate, codetiming, datasets, dill, hydra-core, liger-kernel, numpy, pandas, peft, pyarrow, pybind11, ray, tensordict, transformers, wandb, and tensorboard. The transformers version constraint explicitly excludes 5.6.0 due to a known flash-attention bug, as noted in the requirements.txt comment. The setup.py serves as a fallback when pyproject.toml does not work.

Supported Algorithms and the Examples Directory

The examples directory contains training script folders for a range of RL algorithms. The README and the examples directory listing show configurations for PPO, GRPO, RLOO, REINFORCE++, DAPO, GSPO, ReMax, CISPO, DPPO, and others. Each example directory is named after its algorithm: examples/ppo_trainer/, examples/grpo_trainer/, examples/rloo_trainer/, examples/gspo_trainer/, examples/cispo_trainer/, examples/dppo_trainer/, examples/reinforce_plus_plus_trainer/, and so on.

The recipe directory, migrated to a separate verl-recipe repository and added as a git submodule, holds community-contributed recipes including VeRL-Tinker for SFT, RL, and distillation workflows. The README notes that after cloning, running git submodule update --init --recursive recipe is required to fetch the submodule. Experimental features including transfer_queue, fully_async_policy, one_step_off_policy, and vla are kept under verl/experimental/ and are planned for eventual merge into the main library.

The examples also include reward model configurations (examples/rewards/), data preprocessing scripts (examples/data_preprocess/), generation utilities (examples/generation/), and router replay configurations (examples/router_replay/). These auxiliary examples are needed before or alongside a training job and are not covered by the core library alone.

Ecosystem Extensions and Related Projects

The verl organization hosts several extension projects described in the README news section. verl-vla (v0.1.0) is a unified post-training framework for vision-language-action (VLA) models supporting simulators, real robots, and distributed cloud-edge resources. VeRL-Omni (v0.2.0) is an RL stack for diffusion and omni-modal model post-training, covering DPO and GSPO training for Qwen3-Omni multimodal models. VeRL-Tinker integrates the Tinker Cookbook with verl-managed GPU workers for SFT, RL, and distillation workflows. uni-agent is a unified agent framework for building, running, and training LLM agents at scale.

RL-Insight provides online observability for RL training, connecting training-side metrics, RL state traces, and service dashboards across distributed rollout and optimization workloads. verl-SpeCo is a co-training framework for speculative decoding that keeps draft models aligned during training and reusable for accelerated serving.

All of these are separate repositories under the verl-project GitHub organization. The main verl repository focuses on the core training library and references them only in the news section of the README; they are not bundled into the verl Python package.

Limitations and When verl Is the Wrong Choice

verl requires GPU infrastructure. There is no CPU-only training path documented in the README or requirements. The dependency tree is large: Ray, vLLM or SGLang, PyTorch, and multiple optional LLM frameworks must all be compatible versions, and the uv.lock file is the single source of truth for keeping them consistent. The pyproject.toml comment notes that all backends use torch 2.13.0, but torchaudio stays at 2.11.0 because that is its last release, introducing a version mismatch that is managed by the lock file.

Python 3.13 is not supported; the pyproject.toml specifies requires-python = '>=3.10,<3.13'. The transformers package version window is constrained: 5.6.0 is explicitly excluded due to a broken flash-attention path (the requirements.txt comment references huggingface/transformers#45588), and the maximum is 5.13. These constraints narrow the compatible environment significantly.

OpenRLHF is an alternative RL post-training framework for LLMs that also integrates with Ray and vLLM. The architectural difference is that OpenRLHF uses a simpler pipeline structure, while verl's hybrid-controller model aims for more flexible dataflow composition at the cost of additional configuration. Teams running algorithms beyond PPO and GRPO will find verl's example directory useful for starting points, but each new algorithm requires understanding the controller abstraction.

Maintenance, Versions, and the Apache-2.0 Licence

The last push to the main branch was on 2026-09-25. The most recent releases are v0.8.0 (2026-06-01), v0.9.0 (2026-08-14), and v0.9.1 (2026-09-20), showing active release activity with roughly six-week intervals between minor versions. The project was presented at NVIDIA GTC26, PyTorch Conference EU 2026, and earlier at PyTorch Conference 2025 and an Expert Exchange Webinar in August 2025, reflecting community engagement.

The news section of the README documents multiple external integrations: verl-SpeCo for speculative decoding co-training, vexact for zero-mismatch HuggingFace rollout with batch-invariant kernels, and RL-Insight for online observability. These are all separate repositories under verl-project and are not bundled into the main library.

The README mentions that Mind Lab used verl and Megatron-Bridge to train GRPO LoRA for a trillion-parameter model on 64 H800 GPUs, and that DAPO achieved 50 points on AIME 2024, per the README's news section.

verl is released under the Apache-2.0 licence. The Notice.txt file is present in the repository root, as required by Apache-2.0 when distributing the source. Documentation is hosted on ReadTheDocs at verl.readthedocs.io. The project uses a .readthedocs.yaml configuration for documentation builds.

Editorial conclusion

verl suits ML teams that need to run RL post-training on large language models, particularly those who want to combine FSDP or Megatron-LM for training with vLLM or SGLang for rollout. The 3D-HybridEngine for actor model resharding is the main performance mechanism, and it requires GPU infrastructure; the recipe directory and the examples show concrete configurations for PPO, GRPO, and a range of other algorithms. Verify that your target Python version is between 3.10 and 3.12 before starting, as 3.13 is not yet supported.

Frequently asked questions

What does verl do?

verl is a reinforcement learning training library for large language models. It provides a hybrid-controller programming model for implementing RL algorithms like PPO and GRPO, integrates with FSDP, Megatron-LM, vLLM, and SGLang, and includes the 3D-HybridEngine for efficient actor model resharding during training.

Is verl open source?

Yes. verl is open source and released under the Apache-2.0 licence. It is the open-source implementation of the HybridFlow paper and is hosted on GitHub under the verl-project organization.

Does verl use vLLM?

Yes. vLLM is one of the supported rollout backends in verl. The pyproject.toml lists it as an optional extra, and the recommended install for a vLLM-based setup is uv sync --extra fsdp --extra vllm. SGLang is also a supported alternative rollout backend.

How do I install veRL?

verl uses uv for dependency management. Run uv sync --extra fsdp --extra vllm to install the core library with FSDP training and vLLM rollout support. A requirements.txt is also provided as a reference. Python 3.10 to 3.12 is required.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/verl-project-verl.svg)](https://hysenlabs.com/projects/verl-project-verl)