# Three lines of Python, two required dependencies, and eleven benchmark rows

> A-Evolve presents itself as infrastructure for evolving any agent on any benchmark with no human in the loop, and its pitch is a three-line Python call. The repository behind it is version 0.1.0, declares MIT in three places while shipping no LICENSE file, and keeps a research news feed that runs five months past its own benchmark table.

**A-EVO-Lab/a-evolve** — The official repository of "Position: Agentic Evolution is the Path to Evolving LLMs". 

- Repository: https://github.com/A-EVO-Lab/a-evolve
- Stars: 805 · Forks: 95
- Language: Python
- License: not declared
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/a-evo-lab-a-evolve

## The three-line example needs extras before it can run

The pitch is three lines:

```python
import agent_evolve as ae

evolver = ae.Evolver(agent="./my_agent", benchmark="swe-verified")
results = evolver.run(cycles=10)
```

A path to an agent, a benchmark name and a cycle count. What the manifest puts behind that is worth reading first. The base package depends on exactly two libraries, matplotlib and pyyaml, and requires Python 3.11 or newer. Every model provider and every benchmark harness is an optional extra: anthropic, openai, bedrock and litellm for the LLM side, then swe, mcp, skillbench, osworld and gepa for the harnesses, plus an all extra that lists all nine and a dev extra with pytest, ruff and hypothesis. So the three lines are honest about the call and quiet about the install. The Makefile's own answer is `pip install -e ".[all,dev]"`, which pulls every provider at once whether or not you use them.

## The headline says four benchmarks, the table has eleven rows

The Benchmark Highlights section is introduced as pushing agents into top-tier performance across four diverse benchmarks, and then prints eleven entries. MCP-Atlas goes to 79.4 percent, +3.4 points, ranked number 1. SWE-bench Verified reaches 76.8 percent, +2.6, at roughly number 5. Terminal-Bench 2.0 reaches 76.5 percent, +13.0, at roughly number 7. SkillsBench reaches 34.9 percent, +15.2, at number 2. ARC-AGI goes to 12.3 percent, +2.2, at number 2 on the community leaderboard. OSWorld reaches 69.6 percent, +3.9 with no rank given at all. The remaining four rows are marked Evolved rather than ranked: SWE-bench Lite 63.7 to 67.0, tau-bench 72.7 to 77.0, CL-Bench 29.5 to 34.0 and WebArena-Infinity 72.5 to 76.3. Only those last four print a starting number, so for the seven ranked rows the baseline has to be taken on trust.

## The biggest point gains land where the scores are lowest

Reading the uplifts against the final scores gives a consistent shape. The two largest gains are SkillsBench at +15.2 points and Terminal-Bench 2.0 at +13.0, and SkillsBench ends at 34.9 percent, roughly half of what the leaders in that row achieve by the ranking it claims. The smallest gains sit at the top: MCP-Atlas adds 3.4 points to reach 79.4 percent and SWE-bench Verified adds 2.6 to reach 76.8. ARC-AGI adds 2.2 to reach 12.3 percent, and a separate news item reports the same benchmark family improving from 10 to 12 percent. That pattern is consistent with an optimisation pass finding more headroom where a base model is weak, and it also means the headline numbers that look most impressive are the ones closest to the ceiling. Every figure in the table comes from one Claude Opus-4.6 base model and the project's own sample algorithms.

## The table was checked in March, the news feed runs to August

The footnote under the benchmark table says data checked March 2026. The news section above it is dated later: an algorithm drop in April, an integration into a skills library in April, benchmark results in May, three papers between late May and mid August. So the table and the feed are describing different moments, and the reader has to keep them apart. The news items are specific. Evo-Harness, posted on 15 August, is an EMNLP 2026 paper on compiling noisy single-shot executions into reusable harnesses, evaluated across five benchmarks, with code on a release branch. A June tech report covers autonomous post-training of a 30B Nemotron across four rounds. An earlier EMNLP paper tested seven evolver models against six solver agents on three benchmarks. The repository itself mirrors that structure: examples/ holds ten directories, one per benchmark family plus configs and a harness-disentangling folder.

## The largest claim is a 30B model post-trained with nobody watching

The strongest statement in the repository is in a June tech report rather than in the code. It describes a system running the loop with no human in the loop, post-training a 30B Nemotron across four rounds over several weeks, reaching a held-out score of 0.86 against the top human submission's 0.87 on the public NVIDIA Nemotron-Reasoning Challenge leaderboard, placing 8th of roughly 4000 entries at the time of writing. The comparison it draws is that prior public demonstrations of autonomous machine learning research sit at GPT-2-class budgets of about 124 million parameters. The same system is said to post-train the 120B and 550B Nemotron models as well. None of that code is in this repository, which is the Python package for evolving agents rather than a training pipeline.

## One headline number comes from a build the post calls leaked

The 7 April news item reads: A-Evolve added recently leaked public ClawCode, described in the same line as Claude Code, took the evolution harness and skills learned on Terminal-Bench 2.0 and transplanted them onto it. The reported result is a baseline of 67.8 percent rising to 72.9 percent, a 5.1 point uplift, linked to a post rather than to a table in the repository. That is the only item in the feed where the measured system is a proprietary product rather than a published model or benchmark, and it is also the item with the least documentation behind it. Anyone citing the 72.9 percent figure should note that the artifact it came from is described in the project's own words as recently leaked, and that no reproduction path is given.

## The Makefile lints one directory and installs everything

There are four targets, all marked phony. `install` runs `pip install -e ".[all,dev]"`, which as noted brings in every provider extra plus pytest, ruff and hypothesis. `test` runs `pytest tests/ -v`. `lint` and `fmt` both target the package directory specifically, `ruff check agent_evolve/` and `ruff format agent_evolve/`, so nothing outside the Python package is covered by either. The rest of the tree is not Python package code: DESIGN.md, QUICKSTART.md and CLAUDE.md sit at the root beside artifacts/, figs/, docs/, seed_workspaces/ and tests/. The name of the import, agent_evolve, does not match the distribution name a-evolve, so the package is installed one way and imported another, which is normal and still worth knowing before you grep for it.

## MIT is declared three times and the LICENSE file is not there

The licensing picture needs reading rather than assuming. The license metadata on the repository comes back empty. The manifest declares `license = "MIT"` and carries a classifier for the MIT License. The README's badge row links to opensource.org/licenses/MIT. And the repository tree contains no LICENSE file, no LICENSE.md and no COPYING, so the only text anyone can actually read is the two words in the manifest and the badge URL. For a project whose selling point is reproducibility, that gap is worth closing, and it costs one file. Version state is similarly thin: the manifest is at 0.1.0, the repository has no GitHub releases, and the last push landed on 22 August 2026, so there is no tag to pin and no changelog to diff.

## Conclusion

Judgment: A-Evolve is best read as a research group's public notebook rather than a product. The API surface is genuinely small, the extras are cleanly separated so a reader can install only the provider they need, and the benchmark table is specific enough to check. What sits between the marketing line and the evidence is the calendar: the table says data checked March 2026 while the news feed runs to August, the strongest claim in the feed is a 30B post-training run reported in a tech report rather than in this repository, and one headline number comes from benchmarking a build the post itself calls recently leaked. Anyone using this to evolve their own agent should start from the per-benchmark seed directories under examples/ and treat the ranking table as a claim to reproduce, not a result to cite.

## FAQ

### What is A-Evolve?

An open source Python infrastructure for evolving agents, described as the PyTorch for agentic AI. You hand it a base agent and a benchmark and it returns an evolved agent, with the call presented as three lines using an Evolver object and a run method taking a cycle count.

### How do I install A-Evolve?

The Makefile install target runs pip install -e ".[all,dev]", which pulls every provider extra. The base package itself needs only matplotlib and pyyaml on Python 3.11 or newer, with anthropic, openai, bedrock, litellm, swe, mcp, skillbench, osworld and gepa kept as optional extras so you can install only what you use.

### Which benchmarks does A-Evolve report results on?

The highlights table has eleven rows: MCP-Atlas, SWE-bench Verified, Terminal-Bench 2.0, SkillsBench, ARC-AGI, OSWorld, SWE-bench Lite, tau-bench, CL-Bench and WebArena-Infinity. All figures come from a single Claude Opus-4.6 base model evolved with the project's sample algorithms, and the table is annotated as data checked March 2026.

### Does A-Evolve have a released version?

No. The manifest is at version 0.1.0 and there are no GitHub releases, so there is no tag to pin. The last push to the repository landed on 22 August 2026.

### What license is A-Evolve released under?

The manifest declares MIT and carries an MIT classifier, and the README links an MIT badge. The repository's license metadata is empty, however, and there is no LICENSE file in the tree, so the license text itself is not present.

## Sources

- [A-EVO-Lab/a-evolve on GitHub](https://github.com/A-EVO-Lab/a-evolve)
- [Issues](https://github.com/A-EVO-Lab/a-evolve/issues)
- [README](https://github.com/A-EVO-Lab/a-evolve/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/a-evo-lab-a-evolve
