Model or dataset
om-ai-lab/VLM-R1 avatar
om-ai-lab/VLM-R1

VLM-R1: training Qwen2.5-VL and InternVL with GRPO instead of SFT

Solve Visual Understanding with Reinforced VLMs

6,028 stars383 forksPythonApache-2.0

At a glance

What is it?
VLM-R1 is a research codebase for reinforcement-learning post-training of vision-language models. It is built around GRPO, a verifiable reward, and JSONL task data, and its own experiments show where RL beats supervised fine-tuning and where it does not.
Who is it for?
Adopt VLM-R1 if you already fine-tune Qwen2.5-VL or InternVL and your task has a checkable answer, such as a box or a number, and you want to compare GRPO against SFT on your own data. Do not adopt it if you need a supported product with a stable API, a single-GPU path, or a documented rollback story; the README does not describe one.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 86 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What VLM-R1 trains, and who it is aimed at

VLM-R1 is a training repository, not an inference product. The README describes it as an R1-style large vision-language model project: it takes an existing VLM and post-trains it with reinforcement learning so that the model reasons about an image before answering. The two tasks the README covers in detail are Referring Expression Comprehension, where the model must localize an object described in text, and Open-Vocabulary Detection, where it must produce boxes for a set of categories. A third variant, VLM-R1 Math, is described as reaching the top of the Open-Compass Math Leaderboard under 4B parameters.

The intended user is someone who already runs fine-tuning jobs on multimodal models and wants to know whether RL helps. The README's central comparison is exactly that: for REC, it reports that the SFT model barely moves on in-domain data between 100 and 600 training steps, while the R1 model improves steadily, and that on out-of-domain data the SFT model degrades slightly as steps increase while the RL model generalizes. That is a claim about their runs, not a guarantee for yours, but it is the reason the project exists. If you are looking for a pretrained model to call from an application, the checkpoints on Hugging Face are the relevant artifact, not this repository.

GRPO with a verifiable reward, and where the reward code lives

The training loop is GRPO, implemented inside a vendored copy of open-r1-multimodal under src/open-r1-multimodal. The entry point for the JSONL-based tasks is src/open-r1-multimodal/src/open_r1/grpo_jsonl.py, which the April 2025 update says now handles REC alongside the other tasks for consistency. GRPO needs a reward it can compute without a separate reward model. For REC that reward is geometric: the predicted box is compared against the ground-truth box. The June 2025 update adds a post-resize step for the bounding box in both training and evaluation, which the release notes say improved results slightly.

Reward logic can live in two places. By default it sits with the task code, but the parameter is_reward_customized_from_vlm_module, introduced in the April 2025 update, moves it into the model module instead: QwenVL2Module or InternVLModule depending on which backbone you selected. That split matters when you add a task, because you have to decide whether the reward belongs to the model wrapper or to the dataset code. For the OVD task the README points to three additional reward terms, odLength, weighted_sum and cosine, implemented in grpo_jsonl.py, with the rationale in the project's blog posts. The README does not spell out the formula for each term, so reading the source is the only way to know what you are optimizing.

Installing VLM-R1 and running the REC training script

The repository ships a Dockerfile, which is the clearest install path documented in the repository. It starts from pytorch/pytorch:2.5.1-cuda12.4-cudnn9-devel, installs flash-attn from a prebuilt wheel with a fallback build, then installs wandb 0.18.3, tensorboardx, qwen_vl_utils, torchvision and transformers from the Hugging Face git repository. It copies src/open-r1-multimodal into the image and installs it in editable mode with the dev extra, then installs vllm 0.7.2. There is also a setup.sh at the top level, but the README does not document what it does, so the Dockerfile is the path to follow.

bash
docker build -t vlm-r1 .
docker run --gpus all -it --rm -v $PWD:/workspace vlm-r1

Inside the container, the full fine-tuning entry point for REC is a shell script rather than a Python command. The README lists run_scripts/run_grpo_rec.sh for full GRPO fine-tuning, run_scripts/run_grpo_rec_lora.sh for the LoRA variant, run_scripts/run_grpo_gui.sh for multi-image input, and run_scripts/multinode_training_demo.sh for multi-node runs.

bash
cd /workspace/src/open-r1-multimodal
bash run_scripts/run_grpo_rec.sh

Expect the script to pull a Qwen2.5-VL checkpoint, read a JSONL dataset, and log to wandb or TensorBoard. The README does not state the default batch size, GPU count or expected wall-clock time, and it does not give a sample of the console output, so the first run is a matter of reading the script and adjusting it. To freeze the vision tower instead of training it, the README says to set freeze_vision_modules to true in the script. For your own data, the README points to a For your own data section and, for multi-image records, to the format used by the GUI script.

The constraints the README does not hide

The first constraint is hardware. Every training path in the README is a shell script written for multi-GPU or multi-node execution, and the only single-machine inference documentation is for Huawei Ascend hardware: ascend_inference/910B/vllm_ascend/README.md, ascend_inference/300IDuo/README.md, and ascend_inference/910B/xllm/README.md. There is no documented single-GPU training recipe and no CPU path. If you have one consumer card, the LoRA script is the only plausible starting point, and the README does not say what memory it needs.

The second constraint is that rewards are task-specific code. GRPO only works when the answer can be checked automatically. REC has a box to compare, Math has a numeric answer, OVD has category and box structure. A task whose quality is judged by a human, such as captioning or open-ended visual question answering, has no reward here, and the README offers no reward model or preference-based alternative. The third constraint is that the codebase is a research tree with a vendored copy of open-r1-multimodal inside it. The Dockerfile installs transformers from a git URL rather than a pinned release, so a rebuild months apart can pull different code. The repository also contains a test.py at the top level whose purpose the README does not explain.

How VLM-R1 differs from SFT-only toolchains

The obvious alternative is to skip reinforcement learning entirely and run supervised fine-tuning on the same data. VLM-R1's own comparison is the argument against that for this task: the README states that on in-domain REC data the SFT model changes little between 100 and 600 steps while the RL model improves, and that on out-of-domain data SFT degrades slightly as steps grow while RL generalizes. The mechanism behind the difference is the reward signal, which scores sampled completions instead of matching a fixed target string, so the model is not penalized for phrasing a correct answer differently.

A second comparison is with the upstream open-r1 project, which VLM-R1 vendors and extends. Upstream targets text-only reasoning models; VLM-R1 adds the vision modules, the bounding-box reward and the QwenVL and InternVL adapters under src/open-r1-multimodal/src/open_r1/vlm_modules/. The README's How to add a new model document describes extending that adapter layer, and it states that QwenVL and InternVL are supported today. A third point of comparison is the inference side: rather than serving from this repository, the intended deployment path appears to be the published checkpoints, with the August 2025 updates documenting vllm-ascend and xllm on Ascend hardware. The README claims TTFT reduced by 50% and throughput increased by 127% versus vllm-ascend when using xllm, which are the project's own numbers and not independently reproduced here.

Maintenance, licence and what an upgrade costs

The repository is not archived, and the last push was on 2026-07-07, which is recent enough that the tree is still moving. The published releases are older: v0.2.1 on 2025-04-15, v0.2.0 on 2025-03-24 and v0.1.0 on 2025-03-17. There is a gap between the release tags and the repository activity, and the README's update log continues past v0.2.1 with entries through August 2025 that are not reflected in a newer tag. Practically, that means pinning to a tag gives you the April 2025 state, and tracking main gives you the later changes without a version number to cite.

The licence is Apache-2.0, stated at the top level of the repository. That permits commercial use and modification, and it includes a patent grant, but it also carries notice and attribution obligations that you should read in LICENSE rather than take from a summary. This is not legal advice. On upgrade cost: the Dockerfile installs transformers from git, so the dependency graph is not fully pinned; the vendored open-r1-multimodal tree means upstream fixes have to be merged by hand; and the reward code lives in task files, so a change to the box post-resize logic in qwen_module.py affects both training and the evaluation path in src/eval/test_rec_r1.py. Budget time for re-reading those two files after any pull.

Editorial conclusion

Adopt VLM-R1 if you already fine-tune Qwen2.5-VL or InternVL and your task has a checkable answer, such as a box or a number, and you want to compare GRPO against SFT on your own data. Do not adopt it if you need a supported product with a stable API, a single-GPU path, or a documented rollback story; the README does not describe one. Verify first that your JSONL fields match the REC example, that your GPU memory fits the chosen script, and that the checkpoint you pick is the REC, OVD, or Math variant rather than the base model.

Frequently asked questions

Is VLM better than OCR?

VLM-R1 does not address OCR or document text extraction. It is a training repository for vision-language models on referring expression comprehension, open-vocabulary detection and math, so the question falls outside what the project covers.

Is VLM better than LLM?

VLM-R1 does not make this comparison. It post-trains vision-language models such as Qwen2.5-VL and InternVL with GRPO, so the comparison the README draws is between reinforcement learning and supervised fine-tuning on the same VLM.

Which are the vision language models?

For VLM-R1 specifically, the README states that QwenVL and InternVL are supported, and the training and checkpoint work is built on Qwen2.5-VL, including the 3B REC, OVD and Math variants published on Hugging Face.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. om-ai-lab/VLM-R1 on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/om-ai-lab-vlm-r1.svg)](https://hysenlabs.com/projects/om-ai-lab-vlm-r1)