# verl-agent: step-independent multi-turn RL training for LLM agents

> verl-agent extends veRL with a step-wise rollout mechanism for long-horizon agent training and ships the reference implementation of GiGPO. It is a research training stack, not a serving framework, and it inherits veRL's dependency constraints.

**langfengQ/verl-agent** — verl-agent is an extension of veRL, designed for training LLM/VLM agents via RL. verl-agent is also the official code for paper "Group-in-Group Policy Optimization for LLM Agent Training"

- Repository: https://github.com/langfengQ/verl-agent
- Website: https://huggingface.co/papers/2505.10978
- Stars: 2,353 · Forks: 227
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/langfengq-verl-agent

## What verl-agent solves that plain veRL does not

Standard RL post-training for language models treats a rollout as one prompt and one completion. Agent tasks do not work that way. A task in ALFWorld can require up to 50 steps, and each step produces an observation that has to go back into the model's context. The common shortcut is to concatenate the full interaction history into a single growing prompt. That works until the context window fills, and it removes any ability to decide what the model should see at step 12 versus step 40.

verl-agent's answer is a step-independent multi-turn rollout mechanism. The README describes it as allowing fully customizable per-step input structures, history management, and memory modules. In practice this means the training loop treats each environment step as its own model input, so you can drop old observations, summarize them, or route them through a memory component without rewriting the rollout code. The repository has an agent_system/memory directory for the modular memory manager added in July 2025, and a recipe/ directory holding algorithm variants such as hgpo and GraphGPO.

The audience is narrow and specific: RL researchers and post-training engineers who already work with veRL or a comparable Ray-based stack and who need agent trajectories rather than single-turn preference data. If your task fits in one turn, this project adds machinery you will not use.

## The rollout loop and where GiGPO fits

The architecture follows veRL's split between a trainer and a rollout worker, with the environment layer inserted between them. Environments run in parallel, and the README lists group environment support specifically for group-based RL. That detail matters because GiGPO, the project's own algorithm, is a group-in-group method: it needs more than one rollout per prompt to compute its advantage estimates.

GiGPO is the paper this repository is the official code for, accepted at NeurIPS 2025. The README also records that GiGPO was adopted by Alibaba's ROLL and that the verl-agent style training pipeline was picked up by OpenManus-RL, which is a reasonable signal that the step-wise rollout abstraction is portable rather than tied to one codebase.

Beyond GiGPO, the repository ships GRPO, PPO, DAPO, GSPO, RLOO and REINFORCE++, plus dynamic sampling and clip-higher. Each has its own example directory under examples/, for example examples/gigpo_trainer/, examples/grpo_trainer/ and examples/ppo_trainer/. The examples are shell scripts that launch training through the veRL entry points, so the algorithms differ in configuration rather than in a separate Python API you would import.

## Installing veRL plus an environment, then running one training job

The repository does not publish a single install command. The README's installation section is split into installing veRL and then installing each supported environment separately, with subsections for ALFWorld, WebShop, Search, Sokoban, Gym Cards and AppWorld. AppWorld is marked experimental. The pyproject.toml declares the project name as verl, not verl-agent, with requires-python >=3.8, and setup.py is described in its own header comment as the fallback installation script when pyproject.toml does not work.

Dependency pinning is the first thing to check. setup.py constrains transformers to <=4.57.3, ray[default] to >=2.41.0,<=2.50.0, and tensordict to >=0.8.0,<=0.10.0,!=0.9.0. The vLLM extra pins vllm>=0.8.5,<=0.11.0, and the SGLang extra pins sglang[srt,openai]==0.5.5 together with torch==2.8.0. Note that requirements.txt at the repository root carries different pins, including transformers==4.51.1 and tensordict<=0.6.2, so the two files disagree and you should treat setup.py as the authoritative one for installation.

A typical install from a clone looks like this. The pyproject.toml sets the build backend to setuptools.build_meta, so the editable install path is the standard one.

```bash
git clone https://github.com/langfengQ/verl-agent.git
cd verl-agent
pip install -e .
```

After that you install whichever environment you intend to train against, following the matching README subsection. Then you launch an algorithm from its example directory. The README points at the GiGPO trainer examples, and the Qwen3-VL support added in December 2025 has its own script:

```bash
bash examples/gigpo_trainer/run_sokoban_qwen3vl.sh
```

There is also a docker/ directory in the repository. The README does not document what the image contains or how to build it, so treat it as a starting point to read rather than a supported distribution channel. What you should expect to see when a job starts is a Ray cluster spun up locally, the environment workers initialised, and training metrics written to Weights & Biases, since wandb is an unconditional dependency in setup.py.

## Where verl-agent is the wrong tool

The dependency surface is the first real cost. An unconditional dependency on flash-attn and liger-kernel means you need a CUDA GPU environment before anything installs cleanly, and the pinned ranges for ray, tensordict and transformers will conflict with whatever else is in your environment. This is not a library you add to an existing application. It is a training stack you build a container around.

The second limitation is that the environment installs are separate and uneven. AppWorld is labelled experimental in the README. The other environments each have their own setup subsection, which implies their dependencies are not unified, and the README does not describe a shared environment interface that would let you swap one for another without touching the training configuration. If your task is not one of the six listed environments, you are writing an adapter, and the README's FAQ entry on adding new environments is where that work starts.

Third, the release history is thin. There is one release, v0.1.0, dated 2025-12-11. The last push to the repository was on 2026-06-09, so the code has moved since that tag, but there is no published upgrade path between versions and no changelog beyond the README's News section. The README does not document rollback, and it does not state a versioning policy. If you need to pin a known-good revision, you are pinning a commit hash and reading the News entries yourself.

Finally, this is not an inference server. There is no serving path here, no API, and no deployment story. If you want to run a trained agent in production, you export the weights and use something else.

## How it compares to OpenRLHF and TRL

The honest alternative depends on what you are giving up. OpenRLHF and TRL both target RL fine-tuning of language models, and TRL in particular is a pip-installable library with a stable public API. The difference is the unit of training. TRL's trainers operate on prompt-completion pairs; you supply a dataset and a reward function, and the trainer handles the rest. There is no environment loop inside the trainer, so multi-turn agent interaction has to be implemented outside it.

verl-agent inverts that. The environment loop is the core abstraction, and the single-turn case is the degenerate one. The cost of that inversion is the dependency weight and the fact that you configure training through shell scripts and Hydra configs rather than a Python class you subclass. setup.py even lists trl<=0.9.6 as an optional extra, which is a reminder that the two can coexist: TRL for the supervised or preference stage, verl-agent for the agentic RL stage.

Within the veRL family the comparison is simpler. Plain veRL gives you the trainer, the parallelism strategies and the algorithm implementations, but not the step-wise agent rollout. verl-agent is a fork that merges upstream veRL features (the June 2025 News entry describes merging all features from the latest veRL, bringing Qwen3, LoRA and REINFORCE++) and adds the agent layer on top. Choosing between them is choosing whether you need that layer.

## Conclusion

Adopt verl-agent if you already run veRL or another Ray-based post-training stack and need per-step control over what an agent sees at each turn, or if you want to reproduce GiGPO. Do not adopt it if you need a stable pip-installable library with semantic versioning and documented upgrade paths, or if you have no GPU cluster to train on. Before committing, verify three things against your own setup: that the pinned transformers and tensordict ranges resolve against your existing environment, that the environment you care about (ALFWorld, WebShop, Search, Sokoban, Gym Cards, AppWorld) has a working install path, and that you can reproduce the example training script for your chosen algorithm from examples/gigpo_trainer/.

## FAQ

### What does verl-agent do?

It is an extension of veRL for training LLM and VLM agents with reinforcement learning. It adds a step-independent multi-turn rollout mechanism so each environment step can have its own input structure, history handling and memory module.

### How do I install veRL and verl-agent?

The README splits installation into installing veRL and then installing each supported environment separately, with subsections for ALFWorld, WebShop, Search, Sokoban, Gym Cards and AppWorld. From a clone, pip install -e . uses the setuptools build backend declared in pyproject.toml, and setup.py is described as the fallback script when pyproject.toml does not work.

### Can you give me an example of a reinforcement learning agent trained with verl-agent?

The README lists ALFWorld, WebShop, Search via tool calling, Sokoban, Gym Cards and AppWorld as supported environments, and notes that ALFWorld tasks can require up to 50 steps. Example launch scripts live under examples/, including examples/gigpo_trainer/run_sokoban_qwen3vl.sh for the Qwen3-VL Sokoban case.

## Sources

- [langfengQ/verl-agent on GitHub](https://github.com/langfengQ/verl-agent)
- [License: Apache-2.0](https://github.com/langfengQ/verl-agent/blob/master/LICENSE)
- [Project website](https://huggingface.co/papers/2505.10978)
- [README](https://github.com/langfengQ/verl-agent/blob/master/README.md)
- [Releases](https://github.com/langfengQ/verl-agent/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/langfengq-verl-agent
