# VerlTool: tool-agent RL training built on verl

> VerlTool is a tool-agent training framework built on verl. It decouples rollout from environment interaction, exposes a tool server, and ships training recipes for search, SQL and browser tasks.

**TIGER-AI-Lab/verl-tool** — A version of verl to support diverse tool use [TMLR 2026]

- Repository: https://github.com/TIGER-AI-Lab/verl-tool
- Website: https://arxiv.org/pdf/2509.01055
- Stars: 1,045 · Forks: 89
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/tiger-ai-lab-verl-tool

## The problem VerlTool exists to solve

Reinforcement learning for language models usually assumes a single response per prompt. Tool-calling agents break that assumption. A rollout becomes a loop: the model emits a call, an environment executes it, the result goes back into the context, and the model continues until it stops or hits a turn limit. Reward depends on the final answer, but the trajectory that produced it spans several model calls and several external processes.

VerlTool targets that gap. It is a version of verl extended so that tool interaction is part of the rollout rather than something bolted on afterwards. The README describes it as a "unified and easy-to-extend tool-agent training framework based on verl", and the repository keeps verl as a submodule so upstream changes can be pulled in. The intended user is someone doing RL research or post-training on agent behaviour: reproducing Search-R1, training NL2SQL policies, or building a new tool environment and needing the training loop to handle it.

The paper behind it was accepted at TMLR 2026, and the news entries note a Best Paper Award at ICLR 2026 SPOT. That is context for the design, not evidence that the code will fit your cluster.

## How the rollout, tool server and environment state fit together

The architecture separates two things that are usually tangled: generating tokens and executing tools. The README calls this "complete decoupling of actor rollout and environment interaction". Tool calling goes through a unified API, and adding a tool is described as adding a Python file that can be tested on its own.

The second piece is the "tool-as-environment" paradigm. Each tool interaction can change environment state, and VerlTool stores and reloads that state per trajectory. This matters because a tool is rarely pure: a filesystem, a database session or a browser page carries state across turns, so a rollout cannot be replayed correctly unless that state travels with the trajectory.

Third, the framework supports multi-turn interactive loops natively rather than treating each turn as an independent sample. The docs list separate design notes for synchronous rollout (assets/docs/sync_design.md) and asynchronous rollout (assets/docs/asyncRL.md). The 2025-06-18 news entry states that trajectory-level asynchronous support speeds up rollout generation with tool calling by at least 2x; that is the project's own claim, not an independent measurement.

An evaluation suite sits on the other end. The README says you can launch a trained model behind an OpenAI API alongside the tool server, send questions, and get final outputs with all interactions handled internally. That is a useful property: the same tool definitions used in training are used at inference, so you are not maintaining two codepaths.

## Installing VerlTool and running a first training recipe

The README points to assets/docs/install.md as the Quick Start, and the repository ships pyproject.toml and requirements.txt. Python 3.10 or newer is required, as declared in pyproject.toml. The project metadata declares a base dependency set and several optional extras, so the install path depends on which tool family you need: vllm, tool_browser, acecoder, torl, search_tool, sql_tool, mcp_tool, python_code_dep.

The extras are declared in pyproject.toml under [project.optional-dependencies]. The vllm extra is the entry that carries the pinned inference dependency, bounded at vllm<=0.11.0. The repository does not spell out a pip command for it in the README, so the extra name is what you install against.

The environment variables the tool integrations read are listed in .env.example at the top level. That file is the source of truth for key names, and it includes MCP_GATEWAY_ADDRESS, REDIS_HOST, REDIS_PORT, KAFKA_HOST, KAFKA_PORT, KAFKA_TOPIC, SERP_API_KEY and several model provider keys:

```bash
cp .env.example .env
```

Training entry points live under examples/train, with data preprocessing under examples/data_preprocess. The README links separate recipe READMEs for Search-R1, NL2SQL (examples/train/skysql) and DAPO. Those per-recipe files are where the actual launch commands live; the top-level README does not inline them, so read the recipe you intend to run rather than guessing flags.

## Where VerlTool gets in your way

The dependency surface is the first constraint. requirements.txt pins exact versions across a large stack, including accelerate==1.6.0, datasets==3.5.0, vllm-related packages and flash-attn==2.7.4.post1. That is normal for training frameworks, but it means VerlTool is not something you drop into an existing environment casually. The vllm extra is declared as vllm<=0.11.0, an upper bound, so a newer vllm in your image is outside what the project declares support for.

Second, the codebase moves. The 2025-11-10 news entry states that VerlTool reorganised its codebase for modularity and maintainability and moved to verl 0.6.0 and vllm 0.11.0, with a dedicated upgrade note. An earlier entry (2025-06-16) describes updating the verl submodule and modifying code to adapt. If you vendor VerlTool into a long-lived pipeline, expect to redo that adaptation when the submodule moves.

Third, VerlTool is the wrong tool if you do not need multi-turn tool interaction. Single-turn RLHF or preference optimisation has no environment state to store and reload, so the tool server and the trajectory state handling are overhead with no payoff. It is also a poor fit if you need a framework with a frozen API surface for a production training service; the repository is organised around research recipes.

Finally, the README does not document rollback or downgrade procedures between the reorganised versions. If you need to pin to an older layout, the upgrade note is the only material the project points at, and it is written as an upgrade path, not a compatibility matrix.

## VerlTool versus plain verl and versus a hand-rolled loop

The most direct alternative is verl itself, which VerlTool uses as a submodule. Plain verl gives you the actor, the rollout workers and the RL algorithms; what it does not give you, according to VerlTool's framing, is the tool-agent layer: the unified tool API, the tool server, per-trajectory environment state, and the evaluation service that exposes a trained model with tools attached. Choosing plain verl means writing that layer yourself and keeping it in sync with your training loop.

A second alternative is a hand-rolled rollout loop: call the model, parse the tool call, execute it, append the result, repeat, then feed the finished trajectories into whatever RL trainer you already use. This is genuinely simpler for one tool and one task, and it avoids the pinned dependency stack entirely. The cost appears when you need trajectory-level asynchrony, state reload for non-pure tools, and a shared tool definition between training and evaluation. VerlTool's answer to those three is the async rollout design, the tool-as-environment state handling, and the evaluation suite that runs the tool server next to an OpenAI-compatible endpoint.

A third point of comparison is the lineage the README credits: Search-R1, RAGEN and ToRL. Those are earlier explorations of tool-agent RL training, and VerlTool ships recipes that reproduce them (the 2025-06-30 entry describes reproducing Search-R1 on the same benchmarks, and there is a ToRL example). If you only need one of those specific setups, starting from the original project is a smaller surface; VerlTool's value is having several of them under one rollout and tool abstraction.

## Maintenance, licence and what an upgrade costs

The repository is not archived, and the last push was on 2026-07-15. Releases are tagged: v0.1.0 on 2025-11-10 and v0.2.0 on 2025-12-24. The gap between the last release and the last push means the main branch carries work that is not in a tagged release, so if you need reproducibility, pin to a tag or a commit rather than tracking main.

The licence is MIT, stated in the repository metadata and present as a LICENSE file at the top level. MIT is permissive: it allows modification and redistribution with the licence and copyright notice retained. That said, VerlTool depends on components with their own terms, including vllm, and some extras install directly from Git repositories (mini_webarena, AceCoder) rather than from a package index. Those transitive terms are not covered by VerlTool's MIT licence, and this is a description of the file layout, not legal advice.

Upgrade cost is dominated by the verl submodule and the vllm ceiling. The repository has a dedicated doc for updating the submodule (assets/docs/update_verl.md), which tells you the maintainers expect users to do this. Budget for re-reading the upgrade note and re-testing your tool integration after each move, because the 2025-11-10 reorganisation is the kind of change that can touch import paths.

## Conclusion

Adopt VerlTool if you already train with verl and need multi-turn tool interaction inside the rollout loop, or if you want to reproduce the search, NL2SQL and ToRL recipes without writing the environment plumbing yourself. Do not adopt it if you want a stable, slow-moving dependency: it tracks verl and vllm closely, requires Python 3.10 or newer, and the README's upgrade notes describe codebase reorganisations that move code around. Before committing, verify that the vllm extra (vllm<=0.11.0) matches the GPU stack you have, check which tool extras your task needs, and read assets/docs/tool_server.md plus assets/docs/asyncRL.md to confirm the tool interface matches what you plan to add.

## FAQ

### What is VerlTool and how does it relate to verl?

VerlTool is a tool-agent training framework built on top of verl, which it keeps as a submodule. It adds a unified tool API, a tool server, per-trajectory environment state and an evaluation suite on top of verl's RL machinery.

### How do I install VerlTool?

The README links assets/docs/install.md as the Quick Start. The project requires Python 3.10 or newer and declares optional extras such as vllm, search_tool, sql_tool and mcp_tool, so you install the base package plus the extra for the tool family you need.

### Does VerlTool use vLLM?

Yes. vllm is declared as an optional extra pinned to vllm<=0.11.0, and the 2025-11-10 news entry states the codebase was reorganised to support vllm 0.11.0. The README also credits vLLM and SGLang for inference support.

### What is the VerlTool tool server used for?

The tool server executes tool calls during rollouts and at evaluation time. The README links assets/docs/tool_server.md as the design document, and describes the evaluation path as launching a trained model with an OpenAI API alongside the tool server.

### What licence does VerlTool use?

The repository is MIT licensed, with a LICENSE file at the top level. Some optional extras install dependencies from Git repositories, which carry their own terms separately from VerlTool's licence.

## Sources

- [License: MIT](https://github.com/TIGER-AI-Lab/verl-tool/blob/main/LICENSE)
- [Project website](https://arxiv.org/pdf/2509.01055)
- [README](https://github.com/TIGER-AI-Lab/verl-tool/blob/main/README.md)
- [Releases](https://github.com/TIGER-AI-Lab/verl-tool/releases)
- [TIGER-AI-Lab/verl-tool on GitHub](https://github.com/TIGER-AI-Lab/verl-tool)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/tiger-ai-lab-verl-tool
