# AgentEvolver: a self-evolving training framework for tool-using LLM agents

> AgentEvolver is an Apache-2.0 Python framework from ModelScope that trains agents through three self-evolving loops: generating tasks, reusing experience, and attributing credit across long trajectories. It is a training stack, not a drop-in agent runtime.

**modelscope/AgentEvolver** — AgentEvolver: Towards Efficient Self-Evolving Agent System

- Repository: https://github.com/modelscope/AgentEvolver
- Website: https://modelscope.github.io/AgentEvolver/
- Stars: 1,576 · Forks: 173
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/modelscope-agentevolver

## What AgentEvolver is for, and who it is aimed at

Most agent projects assume you already have good tasks. AgentEvolver starts one step earlier. The README describes it as an "end-to-end, self-evolving training framework" that unifies self-questioning, self-navigating and self-attributing, with the stated aim of letting agents improve their own capabilities without hand-built datasets. That places it in the training-tooling category, not the agent-runtime category.

The intended user is an engineer or researcher who has a tool-using agent, an environment it can act in, and GPUs to train on. The framework assumes you can supply a model endpoint, a conda environment and an environment sandbox. It is not aimed at someone who wants to call an agent from a web app. The repository layout confirms this: agentevolver/, env_service/, config/, examples/ and external/ sit alongside launcher.py, which is a training and orchestration entry point rather than a server.

The three mechanisms map to three distinct problems. Task generation addresses dataset construction cost. Experience-guided exploration addresses rollout quality, the fact that most sampled trajectories are wasted. Attribution-based credit assignment addresses the credit problem in long trajectories, where a single final reward has to be spread across dozens of intermediate steps. Each one is a separate research contribution folded into one codebase, which is also why the configuration surface is large.

## The three self-evolving loops and how data moves between them

The README describes a service-oriented dataflow architecture in which environment sandboxes, LLMs and experience management are separate modular services. That is the clearest statement of the design: the environment is not embedded in the trainer, it is reached over an interface, and the environment service lives in env_service/ with per-environment setup scripts.

Self-questioning explores the environment and creates tasks. Self-navigating summarizes cross-task experience and reuses it to steer rollouts, which is where the optional ReMe component fits; the README calls ReMe the experience management layer and installs it from external/reme/install_reme.sh. Self-attributing processes long trajectories to estimate the causal contribution of intermediate steps, which feeds policy optimization.

The benchmark table in the README is the clearest evidence that these are separable and additive. On a Qwen2.5-7B base, AppWorld avg@8 goes from 1.8 to 23.2 with questioning alone, then to 26.3 with questioning and navigating, and to 32.4 for the full system. The same pattern holds for the 14B model. The ordering matters more than the absolute numbers: adding the mechanisms one at a time moves the metric, which suggests the components are doing distinct work rather than duplicating each other. Those figures are the project's own reported results, not independently reproduced here.

One structural consequence is worth naming. Because experience management is a separate service, the full configuration has more moving parts than the basic one, and the README presents them as two explicit paths: minimal without ReMe, full with it.

## Installing AgentEvolver and running a first training job

The README states three prerequisites: conda, the CUDA toolkit, and Python 3.11 or newer, the last taken from the badge at the top of the page. Installation is script-driven rather than package-driven. There is no pip install agentevolver step documented; requirements.txt is pinned and installed through the project's own script.

The first command sets up the training environment. The README says to run it from the repository root.

```bash
bash install.sh
```

After that, environments are set up individually. AppWorld is the worked example, and its setup script lives under env_service.

```bash
cd env_service/environments/appworld && bash setup.sh
```

Experience management is optional and comes from a separate repository, installed by its own script.

```bash
bash external/reme/install_reme.sh
```

Before training, the README says to copy example.env to .env and edit it, supplying at minimum an API key and the conda path. Then the launcher starts the environment, the log dashboard and the training process together. The minimal example uses built-in datasets from the environments and skips ReMe.

```bash
conda activate agentevolver
python launcher.py --conf examples/basic.yaml --with-appworld
```

The full example adds the ReMe-backed experience layer, and the README marks it as the configuration covering all three mechanisms.

```bash
python launcher.py --conf examples/overall.yaml --with-appworld --with-reme
```

The repository also ships shell wrappers, examples/run_basic.sh and examples/run_overall.sh, for manual execution. What you should expect to see is a launched environment service, a dashboard, and training output; the README does not describe the console output in detail, so treat the first run as a setup diagnostic rather than a benchmark run.

## Where AgentEvolver will not fit

The clearest limitation is scope. AgentEvolver trains agents; it does not serve them. If your requirement is an agent that answers requests in production, the launcher and the environment service are the wrong layer, and you would be adopting a research training stack to do a serving job.

The dependency footprint is the second constraint. requirements.txt pins a large, GPU-oriented set including cupy-cuda12x, numba, onnxruntime, chromadb, vLLM-adjacent packages such as compressed-tensors and lm-format-enforcer, and a pinned modelscope. That is not a library you add to an existing service; it is an environment you build. The README's own instruction to have the CUDA toolkit installed before running install.sh is consistent with that.

The third gap is operational. The README documents how to start training, not how to stop, resume, roll back or upgrade one. There is no release history in the repository, and no migration notes. If your team needs a documented upgrade path before adopting a training framework, this is a reason to wait rather than a reason to reject it outright.

Finally, the benchmark numbers come from the project's own report and its technical report on arXiv. They are presented as evidence of the method's effect, not as a guarantee that the same deltas transfer to your environment. The environment matters: the reported results are on AppWorld and BFCL-v3, and an environment whose tool APIs differ substantially may not respond the same way to self-questioning.

## AgentEvolver compared with Agent0 and EvoAgentX

The nearest comparisons people search for are Agent0 and EvoAgentX, and the difference is where the evolution happens. Agent0 is described in search results as unleashing self-evolving agents from zero data via tool-integrated reasoning. That is a data-free bootstrapping approach centered on the agent's own reasoning with tools. AgentEvolver instead treats task generation as one of three loops and pairs it with an explicit experience store and an explicit credit-assignment stage. The practical difference is that AgentEvolver expects you to run an environment service and, for the full configuration, an experience management service, where a zero-data approach has fewer external components to stand up.

EvoAgentX appears in the same search space as a framework for evolving agent workflows. The distinction is the unit of evolution. Workflow-oriented frameworks optimize the arrangement of agents and steps; AgentEvolver optimizes a policy over trajectories, with attribution across intermediate steps as a first-class concern. If your problem is "which agent should do which step," a workflow-evolution framework is closer. If your problem is "the final reward is too sparse to tell which of forty actions helped," AgentEvolver's self-attributing loop is aimed directly at that.

The honest summary is that these are not interchangeable. Choosing between them turns on whether you have an environment with a reward signal and a GPU budget, which is the precondition AgentEvolver's full configuration assumes.

## Maintenance, licence and the cost of keeping up

The repository is not archived, and the last push was on 2026-04-01. Given how much time has passed since then, it is accurate to describe the project by that date rather than as actively developed. The README news section shows a steady cadence through early 2026: v1 in November 2025, the technical report the same month, the Game Arena and a CuES preprint in December 2025, and a SeeUPO branch release in March 2026. That is a project that has been shipping, but the most recent activity recorded here is five months old.

The maintenance cost for an adopter is mostly environmental. Pinned dependencies plus CUDA plus conda means upgrades are not a version bump; they are a rebuild, and the README offers no upgrade procedure. Budget for re-running install.sh and the per-environment setup scripts when you move.

The licence is Apache-2.0, stated in the badge and in the LICENSE file at the repository root. That is a permissive licence with an explicit patent grant, which is generally straightforward for commercial use, but the repository also pulls in external components, notably ReMe, from a separate project. Those components carry their own licences, and the README points to the ReMe repository for installation details rather than restating terms. Check the licence of each external directory you enable before shipping anything derived from it. This is a description of what the repository states, not legal advice.

## Conclusion

Adopt AgentEvolver if you already run an agent evaluation loop and want to turn it into a training loop on your own hardware, and if you are comfortable reading Hydra configs and environment setup scripts rather than a packaged CLI. Do not adopt it if you need a production agent runtime, a managed service, or a framework whose rollback and upgrade path is documented; the README documents neither. Before committing, verify that install.sh runs end to end under your CUDA and conda setup, that the AppWorld environment setup script completes, and that examples/basic.yaml points at a model and API key your infrastructure can actually reach.

## FAQ

### What is AgentEvolver and who is it for?

It is an end-to-end self-evolving training framework that unifies self-questioning, self-navigating and self-attributing. It is aimed at engineers and researchers training tool-using agents who have an environment sandbox and GPUs, not at teams that need a production agent runtime.

### How do I install AgentEvolver?

The README requires conda, the CUDA toolkit and Python 3.11 or newer, then asks you to run bash install.sh from the repository root. Environments are set up separately, for example cd env_service/environments/appworld && bash setup.sh.

### Do I need ReMe to run AgentEvolver?

No. The README labels ReMe setup as optional and provides two launcher paths: a minimal example using built-in environment datasets without ReMe, and a full example that adds the experience management layer with --with-reme.

### What benchmarks does AgentEvolver report results on?

The README reports AppWorld and BFCL v3 results with avg@8 and best@8 columns for Qwen2.5-7B and Qwen2.5-14B bases, comparing the base model against configurations that add questioning, navigating and attributing. These are the project's own reported figures.

### What licence does AgentEvolver use?

The repository states Apache-2.0 in its badge and ships a LICENSE file at the root. Note that external components such as ReMe come from a separate project and carry their own terms.

## Sources

- [Issues](https://github.com/modelscope/AgentEvolver/issues)
- [License: Apache-2.0](https://github.com/modelscope/AgentEvolver/blob/main/LICENSE)
- [modelscope/AgentEvolver on GitHub](https://github.com/modelscope/AgentEvolver)
- [Project website](https://modelscope.github.io/AgentEvolver/)
- [README](https://github.com/modelscope/AgentEvolver/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/modelscope-agentevolver
