VAGEN: multi-turn RL for VLM agents with world-model rewards
World model reinforcement learning for multi-turn VLM agents. RL for vision framework (NeurIPS 2025).
At a glance
- What is it?
- VAGEN is a Python reinforcement learning framework that trains vision-language model agents across multiple turns and can reward state estimation and transition prediction separately from task success. It is built on VERL and ships environment harnesses for Sokoban, FrozenLake, Navigation, ManiSkill and SVG.
- Who is it for?
- Adopt VAGEN if you are training a vision-language model agent over multi-turn episodes and you want the world-modeling signal to be a separate, swappable reward rather than something baked into the training loop.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 26 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What VAGEN solves, and who it is for
Most reinforcement learning code for language models assumes a single response graded once. A VLM agent that plays Sokoban or navigates a room does not work that way: it sees a frame, acts, sees a new frame, and only at the end does it learn whether it won. VAGEN frames that setting as a partially observable Markov decision process, which the README presents as Figure 2, and trains the policy across the whole episode rather than per turn.
The second problem it addresses is credit assignment inside the episode. A failed Sokoban run tells the model nothing about whether it misread the board or mispredicted the consequence of a push. VAGEN adds two auxiliary rewards on top of task success: StateEstimation, which asks what the current state is, and TransitionModeling, which asks what comes next. Both are scored by an LLM-as-Judge, so the world-model signal is produced by a separate grader rather than by the environment's own reward function.
The audience is narrow and specific. You need a VLM you can run rollouts against, a machine with a CUDA 13 capable GPU stack, and enough familiarity with VERL to read its agent-loop configuration. The repository ships example directories under examples/train and examples/evaluate, and the README names five environments (FrozenLake, Navigation, Sokoban, ManiSkill, SVG) with paired screenshots. If you are doing supervised fine-tuning on static image-question pairs, none of this applies to you.
The harness, environment and rollout boundaries
The design decision worth paying attention to is that VAGEN keeps the environment, model, rollout and evaluation boundaries separate from the training algorithm. That is what makes the world-modeling rewards optional rather than structural: you can run the same environment harness with the auxiliary signals turned off and get a plain multi-turn RL run.
The 2026/08 news entry describes this as a major update that added support for Verl 0.9.0 and decoupled the environment and harness layer. The README also notes that the last commit before those changes is preserved at a tag, which tells you the decoupling was disruptive enough that the maintainers expected people to need the old layout. Read that as a real migration cost, not a footnote.
Environments are resolved through a registry. The setup.py comments explain that vagen/envs/registry.py resolves configs/env_registry.yaml relative to __file__ and raises FileNotFoundError when the file is absent. That comment exists because the packaging was wrong: find_packages() alone ships no data files, and vagen/configs/ holds no .py files, so nothing could reach it. The fix was an explicit package_data entry covering configs/*.yaml, configs/*.flags, envs/navigation/assets/*.json and the per-environment requirements.txt files. This is a useful signal about the project's maturity: the bug never appeared under pip install -e . because the editable install maps back to the source tree, so it only surfaced on a non-editable install.
Installing VAGEN and running a first training job
The README's installation section is short and assumes conda plus a GPU machine. It clones a specific branch from a contributor's fork rather than the main branch, which is worth noticing before you script it into CI.
conda create -n vagen python=3.12 -y
conda activate vagen
git clone --branch dev/bi_level https://github.com/JamesKrW/VAGEN.git
cd VAGEN
bash scripts/install.sh # SGLang 0.5.13 (default)The script initializes the pinned verl submodule, which is why verl appears as a top-level directory in the repository, and installs the CUDA 13 stack: Torch 2.11.0, SGLang 0.5.13 and Transformers 5.8.1. According to the README the script is safe to re-run, and setting SKIP_ENGINE=1 tells it to verify and keep an existing rollout engine instead of reinstalling one. That flag is the one to reach for when you already have a working SGLang build and do not want the installer to touch it.
After installation the entry points are the example directories. The repository lists examples/train and examples/evaluate, so a first run means picking a config under examples/train and pointing it at an environment. The README reports that the stack completed 10-step Sokoban runs for Qwen3-VL and both Qwen3.5 modes, and that GLM-4.6V and InternVL3.5 also completed actor and critic runs. Treat those as the configurations the maintainers exercised, not as a guarantee about your model. If you want a CPU-only smoke test before touching a GPU, the setup.py comments mention that per-environment servers and smoke tests are run as python -m vagen.envs.<env>.<env>_env, but the README does not document a CPU path.
Where VAGEN breaks or wastes your time
The dependency list is the first real constraint. setup.py pins uvicorn<0.41 and adds ninja as a direct requirement with a comment explaining that vLLM compiles kernels through ninja at engine startup even with --enforce-eager, and that neither engine extra pulls it in. Without it, the comment says, the server dies with a bare FileNotFoundError several frames below anything that mentions vLLM. That is a debugging session you avoid only because someone else already had it.
The second limitation is churn. The main branch was migrated to VAGEN-Lite in 2026/02, a lightweight reimplementation built on the VERL agent-loop, and the previous full-featured release lives on the vagen-legacy branch. Then in 2026/08 the environment and harness layers were decoupled for Verl 0.9.0. Two structural rewrites in six months means any tutorial or blog post older than the 2026/08 entry may describe code that no longer exists on main. The README itself warns to check out the tag if you need the previous codebase.
The third is the LLM-as-Judge reward. Using a language model to score StateEstimation and TransitionModeling introduces a second model into your training loop with its own latency and its own failure modes. The README does not document how the judge is configured, what happens when it returns malformed output, or how to audit its scores. If your task has a cheap ground-truth state, a hand-written reward will be faster and more predictable. VAGEN is the wrong tool when your episodes are single-turn, when you cannot afford a judge model in the loop, or when you need a frozen API surface.
How VAGEN differs from VERL and from single-turn RLHF stacks
VERL is the closest reference point because VAGEN is built on it, pins it as a submodule, and migrated main onto the VERL agent-loop. The difference in approach is where the multi-turn logic lives. A plain VERL setup gives you the distributed training machinery and a rollout abstraction; you supply the environment stepping and the reward. VAGEN supplies the environment protocol, the registry, the per-environment servers and the world-modeling reward on top, and keeps those pieces independent of the training algorithm so you can swap the algorithm without rewriting the environment.
Against a single-turn RLHF or GRPO pipeline, the difference is the unit of reward. Single-turn stacks score one completion against one preference or verifier signal. VAGEN scores an episode, and then optionally scores intermediate reasoning about state and transitions within that episode. That extra signal is the entire research claim: the README describes World Modeling RL as supervising the world-model reasoning process explicitly to improve multi-turn performance.
The practical consequence is that you cannot evaluate VAGEN the way you evaluate a single-turn trainer. A run that improves task success but degrades StateEstimation is a meaningful outcome here, and the framework is built so those two numbers move independently. If you only ever look at final success rate, you are not using the part of VAGEN that distinguishes it.
Maintenance, licensing and upgrade cost
The repository is not archived, and the last push was on 2026-09-05, roughly two weeks before this writing. The release list shows one entry, v25.12.30, published on 2026-02-10, which is more than six months before the last push. That gap between a tagged release and ongoing commits is the practical upgrade problem: if you pin to v25.12.30 you are pinning to code from before the VAGEN-Lite migration and before the Verl 0.9.0 decoupling. If you track main you get the current layout but no release boundary to diff against.
Licensing is MIT, per both the repository and setup.py, and the repository carries a NOTICE file alongside LICENSE. MIT is permissive and places few obligations on how you redistribute, but the project vendors or depends on upstream components with their own terms, including VERL as a submodule and the model weights you choose to train. Read the NOTICE file and the upstream licences before shipping anything. This is not legal advice; it is a pointer to the files that matter.
The upgrade cost is dominated by the environment layer. Because the 2026/08 update decoupled environments from the harness, any custom environment you wrote against the earlier protocol needs review against the current one. The README links a guide at docs/custom-environment.md describing the new environment protocol, and a separate document at vagen/envs/_common/remote/README.md for hosting environments in a separate process. Start there rather than reading the diff.
Editorial conclusion
Adopt VAGEN if you are training a vision-language model agent over multi-turn episodes and you want the world-modeling signal to be a separate, swappable reward rather than something baked into the training loop. Do not adopt it if you need a CPU-only setup, a stable tagged API, or a single-turn fine-tuning recipe: the install script targets a CUDA 13 stack with Torch 2.11.0, SGLang 0.5.13 and Transformers 5.8.1, and the README points anyone needing the previous codebase at the vagen-legacy branch or the vagen-lite-verl-v0.6.1-final tag. Before committing, verify that your environment can resolve the pinned verl submodule and that the specific environment you care about still has its entry in vagen/envs/registry.py, because the 2026/08 update decoupled the environment and harness layers and the pre-update commit is kept only as a tag.
Frequently asked questions
How do I install VAGEN?
The README gives a conda recipe: create an environment with Python 3.12, clone the dev/bi_level branch, and run scripts/install.sh. The script initializes the pinned verl submodule and installs the CUDA 13 stack with Torch 2.11.0, SGLang 0.5.13 and Transformers 5.8.1, and it is safe to re-run.
What is the difference between VAGEN and VAGEN-Lite?
VAGEN-Lite is a lightweight reimplementation built on the VERL agent-loop, and the README states that the main branch was migrated to it in 2026/02. The previous full-featured release is kept on the vagen-legacy branch.
What are the StateEstimation and TransitionModeling rewards in VAGEN?
They are two auxiliary world-modeling signals that VAGEN can reinforce alongside task success: StateEstimation asks what the current state is, and TransitionModeling asks what comes next. Both are scored by an LLM-as-Judge, according to the README.
Which environments does VAGEN support?
The README names FrozenLake, Navigation, Sokoban, ManiSkill and SVG, each with paired screenshots. Environments are resolved through a registry that reads configs/env_registry.yaml.
What licence does VAGEN use?
MIT, according to both the repository and setup.py. The repository also includes a NOTICE file, and VAGEN depends on upstream components such as VERL as a submodule that carry their own terms.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/mll-lab-nu-vagen)