EasyR1: A veRL Fork for Vision-Language GRPO Training
EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL
At a glance
- What is it?
- EasyR1 is a clean fork of veRL that adds vision-language model support and a catalogue of RL algorithms. It is worth adopting if you already have multi-GPU hardware and want working GRPO examples, not if you are looking for a CPU-friendly entry point.
- Who is it for?
- Adopt EasyR1 if you have the GPUs the README's table lists for your model size and you want a runnable GRPO or DAPO script for a Qwen-VL or Qwen3-VL checkpoint. Do not adopt it for CPU-only experimentation or for text-only work where plain veRL already covers you.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 30 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap EasyR1 fills: RL training for models that see images
veRL is a reinforcement learning framework for language models. EasyR1 is a clean fork of it, and the fork exists for one reason: vision language models. The README states the project "is a clean fork of the original veRL project to support vision language models." Everything else in the repository follows from that decision.
The target user is an engineer who has a Qwen2-VL, Qwen2.5-VL or Qwen3-VL checkpoint and a reward signal, and wants to run GRPO or DAPO over image-text pairs without rebuilding the distributed training plumbing. The supported model list also covers Llama3, Qwen2, Qwen2.5, Qwen3 language models and DeepSeek-R1 distill models, so text-only runs are possible, but they are not the reason this fork exists.
The repository ships example scripts for Geometry3K, a geometry dataset, plus two R1-V baselines: CLEVR-70k-Counting and GeoQA-8k. Those baselines are the clearest statement of scope. This is a project for people reproducing or extending multimodal reasoning results, not a general-purpose fine-tuning library.
HybridEngine, vLLM SPMD, and where the rollouts come from
The README attributes the framework's efficiency to the HybridEngine design, citing arXiv 2409.19256, and to vLLM's SPMD mode. That is the architecture in one line: generation and training share the same workers rather than living in separate processes, and vLLM handles the rollout side.
The algorithm list is where the project diverges most from a minimal fork. GRPO, DAPO, Reinforce++, ReMax, RLOO, GSPO and CISPO are all listed as supported. The examples directory confirms this with separate scripts: qwen2_5_vl_7b_geo3k_grpo.sh, qwen2_5_vl_7b_geo3k_dapo.sh, qwen2_5_vl_7b_geo3k_gspo.sh, qwen2_5_vl_7b_geo3k_cispo.sh, qwen2_5_vl_7b_geo3k_reinforce.sh, qwen2_5_vl_7b_geo3k_sapo.sh. Each is a single bash file, so comparing two algorithms means diffing two scripts rather than reading two papers first.
Training tricks are listed as padding-free training, LoRA training, resuming from the latest or best checkpoint, and tracking through Wandb, SwanLab, Mlflow or Tensorboard. Datasets are described as "any text, vision-text dataset in a specific format," with four Hugging Face example datasets linked: math12k for text, geometry3k for image-text, journeybench-multi-image-vqa for multi-image-text, and rl-mixed-dataset for text-image mixed data. That set is the practical specification of the data format.
Installing EasyR1 and running a first GRPO job
The README recommends the pre-built Docker image over a bare install. That is the fastest path and it avoids resolving flash-attn against your CUDA version yourself.
docker pull hiyouga/verl:ngc-th2.8.0-cu12.9-vllm0.11.0
docker run -it --ipc=host --gpus=all hiyouga/verl:ngc-th2.8.0-cu12.9-vllm0.11.0The --ipc=host flag matters for shared-memory use across worker processes, and --gpus=all exposes every card. If Docker is not available, the README offers Apptainer as an alternative:
apptainer pull easyr1.sif docker://hiyouga/verl:ngc-th2.8.0-cu12.9-vllm0.11.0
apptainer shell --nv --cleanenv --bind /mnt/your_dir:/mnt/your_dir easyr1.sifFor a source install, clone and install in editable mode:
git clone https://github.com/hiyouga/EasyR1.git
cd EasyR1
pip install -e .The README lists Python 3.9+, transformers>=4.54.0, flash-attn>=2.4.3 and vllm>=0.8.3 as software requirements. Note that requirements.txt pins transformers>=4.54.0,<5.0.0 and vllm>=0.8.0, so the README and the requirements file do not state identical lower bounds for vLLM. The Docker image resolves this by shipping vLLM 0.11.0.
With the environment ready, the first real run is the Geometry3K GRPO example:
bash examples/qwen2_5_vl_7b_geo3k_grpo.shThe README presents this as the tutorial path: Qwen2.5-VL GRPO on Geometry3K "in just 3 steps." Expect the script to download the model and dataset, then write checkpoints under a path beginning with checkpoints/easy_r1/. When the run finishes, merge the actor checkpoint into Hugging Face format:
python3 scripts/model_merger.py --local_dir checkpoints/easy_r1/exp_name/global_step_1/actorTwo environment variables appear in the README and are worth knowing before you start. Set USE_MODELSCOPE_HUB=1 to download models from the ModelScope hub, and set HF_ENDPOINT=https://hf-mirror.com if Hugging Face connectivity is the problem.
The hardware table is the real admission bar
The README's hardware table is marked as estimated, and it is the most honest part of the documentation. GRPO full fine-tuning at BF16 for a 7B model is listed as 4*40GB. At AMP it is 8*40GB. A 72B model at BF16 is listed as 16*80GB. LoRA changes the picture considerably: 7B LoRA at AMP is 2*32GB, and 32B LoRA at AMP is 2*80GB.
Read that as a boundary rather than a suggestion. A single 24GB card appears in the table only for 1.5B AMP full fine-tuning and 3B BF16, and for LoRA at 1.5B and 3B. If your hardware is one consumer GPU, this is the wrong tool and no amount of configuration will change the table.
The table also implies a second constraint that the README does not spell out: the numbers are per-method and per-precision, so the precision flags are not cosmetic. The README notes that bf16 training requires worker.actor.fsdp.torch_dtype=bf16 and worker.actor.optim.strategy=adamw_bf16. Choosing BF16 halves the card count in several rows, which means those two config keys are load-bearing for anyone planning capacity.
The multi-node path goes through Ray. The README's sequence is to start a head node with ray start --head --port=6379 --dashboard-host=0.0.0.0, connect workers with ray start --address=<head_node_ip>:6379, verify with ray status, and then run the training script on the head node only. It defers to veRL's official multi-node documentation for debugger details. That deferral is a fair signal: EasyR1 inherits veRL's operational model, including its failure modes, and the fork does not re-document them.
Where EasyR1 stops being the right choice
The project is a fork, and forks carry a maintenance tax. When veRL changes its internals, EasyR1 either absorbs the change or drifts. The README does not describe a merge or rebase policy, and it does not document a rollback procedure if an upgrade breaks a training run. Anyone pinning EasyR1 in a production pipeline should treat the upgrade path as undocumented until they verify it themselves.
The documentation is also thin in specific places. The custom dataset section says only to "refer to the example datasets" and links four Hugging Face datasets; the format itself is not written out in the README. There is no documented reward function interface beyond the examples/reward_function/ directory in the repository layout. The README does not document checkpoint compatibility across versions, so a checkpoint written by one release and resumed by another is an open question rather than a supported path.
Finally, the project's own README carries a promotion for an unrelated project, PenguinHarness, in a centered block near the top, and a Twitter follow badge. That is a cosmetic annoyance rather than a technical problem, but it tells you the README is partly a marketing surface, and you should read the hardware table and requirements.txt rather than the feature list when you are deciding.
One more boundary worth stating plainly: EasyR1 is a training framework, not a serving stack. The README's only mention of inference infrastructure is vLLM as the rollout engine during training. If your goal is to deploy a model, this repository is upstream of your problem.
EasyR1 against veRL, TRL and OpenRLHF
The most direct alternative is veRL itself. EasyR1 is a clean fork of it, so the difference is not architecture but surface area: EasyR1 adds vision-language support, the algorithm scripts (DAPO, GSPO, CISPO, SAPO), the LoRA examples, and the R1-V baselines. If you are training text-only models and veRL already works for you, the fork adds maintenance overhead without adding capability you need. If you are training Qwen-VL checkpoints, the fork is the shorter path.
TRL is the other reference point, and the README points to Hugging Face's GRPO trainer blog as the place to learn the algorithm. The difference in approach is integration depth. TRL's GRPO trainer is a library component you call from Python; EasyR1 is a distributed training project with Ray, FSDP and vLLM wired together and driven by bash scripts. TRL is easier to embed in an existing training loop. EasyR1 is easier to scale to the multi-node configurations in its hardware table.
OpenRLHF is a third option in the same space. The README does not compare EasyR1 to it, and the repository gives no benchmark against it, so any claim about relative throughput would be invented. What can be said from the documentation is structural: EasyR1's distinguishing features are the vision-language model support and the breadth of algorithm example scripts, and those are the criteria to compare on rather than raw speed.
Licence, maintenance and what an upgrade costs
EasyR1 is Apache-2.0. The setup.py carries the Apache 2.0 header and the LICENSE file is at the repository root. The package name in pyproject.toml and setup.py is verl, not easyr1, which is a fork artifact worth knowing: pip install -e . installs a distribution named verl. That matters if you have the upstream veRL package installed in the same environment, because the two would collide on the same distribution name. The setup.py url field points at github.com/volcengine/verl, and the author email list includes addresses associated with both projects. None of this is a licence problem, but it is a namespace problem you should resolve before you install.
On maintenance: the repository is not archived, and the last push was on 2026-08-31. The most recent release is v0.3.2, tagged "RL Baselines," dated 2025-09-18. v0.3.1 was "Multi-modal DAPO" and v0.3.0 was the initial release. The gap between the last release tag and the last push means the main branch carries work that is not in a tagged release, so pinning to v0.3.2 and tracking main are two different commitments.
Upgrade cost is dominated by the dependency pins, not by the EasyR1 code itself. requirements.txt pins transformers>=4.54.0,<5.0.0 and vllm>=0.8.0, and also pulls flash-attn, liger-kernel, ray[default], tensordict and qwen-vl-utils. A transformers major-version bump is excluded by the upper pin, so you are insulated from that particular break, but a vLLM minor release can still change rollout behaviour. The Dockerfile exists precisely so you do not have to negotiate that graph by hand, and the README's recommendation to use the pre-built image is the practical answer to upgrade cost.
Editorial conclusion
Adopt EasyR1 if you have the GPUs the README's table lists for your model size and you want a runnable GRPO or DAPO script for a Qwen-VL or Qwen3-VL checkpoint. Do not adopt it for CPU-only experimentation or for text-only work where plain veRL already covers you. Before committing, verify that your flash-attn and vLLM versions satisfy the requirements.txt pins, and check whether the example script you plan to copy matches your model size, since the repository ships separate scripts per size and algorithm.
Frequently asked questions
What is EasyR1 and how does it relate to veRL?
EasyR1 is a clean fork of the veRL project created to support vision language models. It keeps veRL's HybridEngine design and vLLM SPMD rollout mode, and adds vision-language model support plus example scripts for GRPO, DAPO, GSPO, CISPO and other algorithms.
How do I install EasyR1?
The README recommends the pre-built Docker image hiyouga/verl:ngc-th2.8.0-cu12.9-vllm0.11.0. For a source install, clone the repository and run pip install -e . from the project root, which requires Python 3.9+, transformers>=4.54.0, flash-attn>=2.4.3 and vllm>=0.8.3.
How much GPU memory does EasyR1 GRPO training need?
The README's hardware table, marked as estimated, lists GRPO full fine-tuning at BF16 as 4*40GB for a 7B model and 16*80GB for a 72B model. LoRA reduces this: 7B LoRA at AMP is listed as 2*32GB and 32B LoRA at AMP as 2*80GB.
Can EasyR1 train on multiple nodes?
The README documents a Ray-based multi-node setup: start a head node with ray start --head --port=6379 --dashboard-host=0.0.0.0, connect workers with ray start --address=<head_node_ip>:6379, check with ray status, then run the training script on the head node only. It points to veRL's official documentation for multi-node debugging details.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/hiyouga-easyr1)