Agent-R1: Training LLM Agents with Step-Level Reinforcement Learning
Agent-R1: Training Powerful LLM Agents with End-to-End Reinforcement Learning
At a glance
- What is it?
- Agent-R1 is a Python framework for training multi-step LLM agents through reinforcement learning, treating each agent turn as a discrete MDP transition rather than as part of a growing prompt sequence. It sits on top of the verl distributed training system and provides layered abstractions for task environments, tools, and policy objectives.
- Who is it for?
- Agent-R1 is the right framework for ML researchers who need a modular RL training substrate that treats agent steps as first-class objects, supports custom tool environments, and separates training algorithms from task logic.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 3 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Agent-R1 Solves and Who It Is For
Standard single-turn RL training for LLMs treats a multi-step agent interaction as one growing prompt-response pair. The problem is that tool calls, environment feedback, and reward signals happen at intermediate steps, and a single-turn view cannot assign credit accurately to individual actions. Agent-R1 addresses this by modeling each agent turn as a step-level MDP transition.
The framework is for ML researchers and applied AI engineers who want to train LLM agents to use tools and interact with environments over multiple turns. It is not an inference-time agent runtime. The output of training with Agent-R1 is a policy-optimized model, not a deployed application.
The Step-Level MDP: How Agent-R1 Models Training
In most multi-turn agent training setups, the model sees all prior turns concatenated into a single context and produces its next output. Agent-R1 treats each turn as a structured step: the step records what the model observed, what action it produced, what feedback and reward the environment returned, and what observation should be exposed next.
This representation keeps four things aligned with real agent decisions: rollout, replay, context construction, and credit assignment. The model still performs token-level policy losses inside each generated action, but the credit assignment problem is handled at step granularity, not at token granularity across the whole episode. That distinction matters for tasks where an intermediate tool call is the critical decision and the surrounding context is noise.
The architecture's main loop is: load a sample containing a prompt, agent name, reward model, and optional environment arguments; create the configured agent flow and environment; generate an action from the current observation; parse the action, execute tools or update the environment, and return feedback; record the step; continue until done or max steps is reached; convert the structured trace into rewards, advantages, masks, and policy updates.
Layered Abstractions: Choosing the Right Entry Point
Agent-R1 provides five layers of abstraction so new tasks can reuse the same trainer without rewriting the full RL stack.
AgentFlowBase gives full control over prompt construction, model calls, branching, context management, and step assembly. Use it for complex custom agents that do not fit a standard environment loop.
AgentEnvLoop is a generic loop connecting model generation with an environment's reset() and step() interface. It handles tasks modeled as environment interaction, including traditional RL-style environments.
AgentEnv is the task environment interface that returns observations, rewards, termination signals, and metadata. Implement it when writing full environment logic for AgentEnvLoop.
ToolEnv is a built-in environment for standard multi-turn tool calling. Use it when you only need to define the tools and let the built-in loop handle the rest.
BaseTool is the standard interface for registering executable tools such as calculators, search tools, APIs, or task-specific checkers. Each layer builds on the one below, so a new task that fits the ToolEnv pattern needs only to define tools, while a task with unusual context management or branching logic uses AgentFlowBase directly.
Setting Up Agent-R1 and Running a Recipe
Agent-R1 uses the same environment as verl, and the current version requires verl 0.7.0. There is no separate Agent-R1 installation step: clone the repository and use it directly on top of the verl environment.
The recommended setup path starts with reading the verl installation documentation, then cloning this repository:
git clone https://github.com/AgentR1/Agent-R1.gitThe repository ships recipes for five tasks in the examples/ directory:
ls examples/
# alfworld/ gsm8k/ hotpotqa/ paper_search/ webshop/Each recipe is a self-contained configuration for a specific benchmark environment. The README directs readers to the Getting Started documentation at agentr1.github.io/Agent-R1/getting-started/ for the full walkthrough. Processed datasets for HotpotQA, ALFWorld, WebShop, and academic paper search are available on ModelScope at the path documented in the 2026-05-29 release notes.
Recent Additions: Online Policy Distillation and StepPO
Two significant extensions landed after the v0.1.0 architecture refactor.
StepPO (Step-level Policy Optimization) integration arrived on 2026-05-29. StepPO applies preference optimization at the step level rather than the episode level, which the framework's step-based trajectory representation supports directly. The recipe expansions on the same date added HotpotQA, ALFWorld, WebShop, and academic paper search support.
Online Policy Distillation (OPD) arrived on 2026-07-21. Generic OPD training is available on the opd branch, not on main. This allows training an agent policy by distilling from a stronger online teacher policy, which can be useful when a target benchmark environment provides limited reward signal.
The Claw-R1 extension, released on 2026-03-04, extends agentic RL to general agents through a middleware-style design and is maintained in a separate repository at AgentR1/Claw-R1.
Limitations and When to Use a Different Tool
Agent-R1 is a training framework that adds no inference-time capabilities. Once a model is trained with Agent-R1, running it as a live agent requires a separate serving setup. The framework does not ship a deployment layer.
The dependency on verl 0.7.0 is a hard constraint. The README states the current version requires exactly that release. Incompatible verl versions will prevent the training loop from functioning. GPU infrastructure capable of running distributed policy optimization is a prerequisite; there is no documented CPU-only training path.
For teams who want to run a multi-step agent at inference time without training a custom policy, frameworks like LangGraph or AutoGen are more appropriate. Those tools manage agent execution graphs and tool calls at serving time. Agent-R1 solves a different problem: it is for teams who want to train a model to become a better agent through RL, not for teams who want to deploy an existing model as an agent.
Context management is flexible but places responsibility on the implementer. The framework allows the environment to decide what the model sees next, including truncating, summarizing, or rewriting history. That flexibility requires explicit design choices; there are no automatic context management policies.
Maintenance and License
The last push to Agent-R1 was on 2026-09-21, and the repository has no formal GitHub releases. The project is licensed under MIT. The repository includes an mkdocs configuration for documentation at agentr1.github.io, a pre-commit configuration for code quality, and a pyproject.toml with ruff and mypy settings.
The arXiv technical report at 2511.14460 was substantially revised on 2026-05-30 to cover the step-level trajectory representation and layered abstractions introduced in v0.1.0. Earlier experiments predating that refactor are on the legacy branch.
Editorial conclusion
Agent-R1 is the right framework for ML researchers who need a modular RL training substrate that treats agent steps as first-class objects, supports custom tool environments, and separates training algorithms from task logic. It is a poor fit for teams looking for an inference-time agent framework or a ready-to-use deployment system: Agent-R1 is a training framework, and running it requires a verl 0.7.0 environment with a GPU stack capable of handling distributed policy optimization. Before starting, verify that your hardware and environment match the verl installation requirements, and check whether the recipe for your target task (HotpotQA, ALFWorld, WebShop, or paper search) is already included in the recipes/ directory.
Frequently asked questions
What is Agent-R1's step-level MDP and why does it matter for training?
Agent-R1 models each agent turn as a discrete MDP transition that records the observation, action, environment feedback, reward, and next observation. This structure keeps credit assignment aligned with actual agent decisions rather than spreading it across all tokens in a growing context, which makes policy optimization more accurate for multi-step tool-use tasks.
Does Agent-R1 require a specific version of verl?
Yes. The README states that the current version requires verl 0.7.0. There is no separate Agent-R1 installation step; the repository is cloned and used on top of the verl environment. Incompatible verl versions are not supported.
What benchmark tasks does Agent-R1 ship recipes for?
The examples/ directory contains recipes for five tasks: GSM8K, HotpotQA, ALFWorld, WebShop, and academic paper search. Processed datasets for the last four are available on ModelScope, as documented in the 2026-05-29 release notes.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/agentr1-agent-r1)