evo: an autoresearch loop for your codebase
turns your codebase into an autoresearch loop — discovers what to measure, instruments the benchmark, then runs tree search with parallel subagents.
At a glance
- What is it?
- A plugin for agentic coding tools that reads a repository, invents a benchmark, gates correctness, and then runs parallel tree search to make the code faster.
- Who is it for?
- evo is a well specified experiment harness rather than an autonomous optimizer that will rewrite your project on its own, and that distinction matters for how you should try it. The repository is unusually clear about the parts that decide whether a run is useful: gates must exist before the loop starts, a benchmark that measures the thing you care about has to come first, and the frontier strategy is something you choose rather than inherit.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 8, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The idea borrowed from Karpathy, then given structure
evo is a plugin for agentic coding tools that optimizes code by running experiments. You hand it a repository, it works out what to measure, builds an evaluation around it, and then loops: propose a change, run the benchmark, keep the edit if the score improved, throw it away if it did not. That core loop is explicitly modelled on Karpathy's autoresearch project, where an LLM runs training experiments to beat its own best score. The README is upfront that the original is a pure hill climb, try a change, keep or revert, repeat on a single branch. evo's contribution is what it layers on top, and the additions are structural rather than cosmetic.
The first addition is tree search rather than a single line of descent, so several directions can fork from any node that has already been committed and exploration does not collapse into one path. The second is parallel semi autonomous agents, each subagent running in its own git worktree so experiments cannot collide. The third is shared state: failure traces, annotations and rejected hypotheses are visible to every agent before it chooses what to try, so one agent's dead end informs the others. On top of that sit pass/fail gates and a dashboard. The repository has 1459 stars, is Apache licensed, and its last recorded push is 17 July 2026, so it is live but young.
Two commands and a seeded variant
The whole interface is two slash commands. Discovery happens once, optimization runs in a loop:
/evo:discover # one-time code discovery: figures out benchmarks and creates gates against unintended changes
/evo:optimize # run the loopThe discover step asks three things: what to optimize, the benchmark command, and which direction counts as better. If you would rather not answer them interactively you can seed the answer in the command itself, and the README gives exactly this form:
/evo:discover make the JSON parser at src/parser.py fasterAfter that, running the loop is the single command `evo optimize` invoked through the host's own syntax. How you type the slash command depends on the host: it is a slash command on Claude Code, a dollar prefixed `$evo` on Codex, a slash skill menu entry on Cursor, and plain natural language on Hermes, Opencode, OpenClaw and Pi. That host specific syntax is the main practical gotcha for anyone coming from another tool, because the command name is constant but the invocation is not. By default the loop runs unattended and pushes edits through parallel subagents. The README says you can instead ask in plain language for it to pause after each round or to run one experiment at a time, which is the sane setting for a first look.
Installing the CLI and the host plugin
There are two pieces to install, the evo command line tool and a plugin wired into whichever agent host you use. The tool is a Python package installed with uv:
uv tool install evo-hq-cliThen you register the plugin and its hooks with your host. The host argument is a fixed set of names, and the trailing comment in the README lists them all:
evo install <host> # claude-code | codex | cursor | hermes | kimi | opencode | openclaw | piThe install step writes the plugin into the host's plugin directory and registers the skills, slash commands and hooks so the next session picks them up. On Codex there is one extra detail worth knowing. `evo install codex` trusts evo's hooks for you, which is convenient but means the hooks run without you reading them first. If you would rather inspect them, pass `--no-trust-hooks` at install time and then approve them yourself through the `/hooks` command inside codex. That is the sort of default that a security conscious reader should flip.
Remote compute backends are opt-in extras rather than bundled. The worktree, pool and ssh backends are included in the base install; Modal, E2B, Daytona, AWS and Azure each need the matching extra, and there is an `[all]` extra for the full set. The pattern is `uv tool install 'evo-hq-cli[modal]'` for one provider. If you are only experimenting locally you can skip all of that, because the default worktree backend needs nothing extra.
Gates are the part that keeps it honest
A loop that only chases a number will find ways to cheat the number, and the README names the failure modes directly: without gates the search will discover how to return a constant, skip work, or trade correctness for speed. evo's answer is gates, pass/fail checks that run on every experiment. An experiment that fails a gate is discarded even if its score beats the current best. Any command that exits zero on pass and non-zero on fail can serve as a gate, which in practice means a test suite, a standalone invariant script, or a score floor on a held-out slice of the benchmark.
Gates are inherited down the experiment tree. A gate registered at the root runs on every descendant, and narrower gates can be attached to specific branches. The automatic behaviour depends on starting point: when discover builds a benchmark from scratch it attaches a held-out-slice score-floor gate automatically, but when a benchmark already exists in the repository, gates are opt-in and you have to wire them yourself. That is the detail to keep in mind. If you already have a benchmark and skip the gate step, the loop will optimise hard against a metric with nothing holding correctness in place, which is exactly the situation the design is warning about.
Where the search runs and how it decides what to extend
Each subagent works in its own isolated workspace, reads the shared state at startup, forms a hypothesis, edits, and runs the benchmark. A subagent that still has iteration budget can continue within its branch in the same round when its previous edit warrants a follow-up. Between rounds, a separate kind of scan agent reads batches of traces in parallel to surface compound failure patterns, things like gate failures that intersect or a root cause shared across otherwise separate traces, and writes those findings back into shared state for the next round. The README credits this design to a recent RLM inspired paper and the Pareto strategy to GEPA.
After each round the orchestrator chooses which committed branch to extend, and this choice is the one knob worth thinking about. The available strategies are argmax, which always extends the highest scoring branch; top_k, which round-robins among the best K; epsilon_greedy, which takes the best most of the time and tries something random occasionally; softmax, which samples weighted by score; and pareto_per_task, which keeps specialists that an aggregate score would hide. You set these in the dashboard's Frontier tab, where each strategy's parameters are listed. The compute backends are similarly plural, from a local git worktree per experiment by default, through a pool of reused local workspaces or your own SSH host, to cloud sandboxes on Modal, E2B, Daytona, AWS or Azure.
Hosts, releases, and honest expectations
evo runs inside the coding agent you already use, and the release notes show the host list growing quickly. The v0.8.0 release added Kimi Code, dropping the plugin into Kimi's plugin directory and registering it so a session picks up the skills and slash commands, with the same two command flow that already worked on Claude Code, Codex and Cursor. The v0.7.1 alpha before it was a batch of fixes against the real kimi-code plugin contract, recognising PascalCase hook events and fixing tool path resolution for the installed layout. v0.7.0 was a larger change: it removed evo's assumptions that it could create `.git` directories, write to your home folder, and bind a local port for the dashboard, so it can run inside locked-down sandboxes. That release added a gitdir backend that relocates git's metadata and shares the base repo's object store, keeping each experiment isolated without ever creating a `.git` path.
The honest summary is that this is a serious harness with real design thought behind the search strategy and the gate system, but it delegates the hard judgement to you. It will only make your code faster relative to whatever benchmark discover settled on, so a vague or self-serving benchmark produces confident nonsense. It also runs on your machine, inside your agent, editing your repository. The docs are clear about the mechanics, the host list and the backends; they are less clear about how to tell a genuinely faster parser from one that got faster by cheating a measurement, which is why the gate advice matters more than any of the search strategy options.
Editorial conclusion
evo is a well specified experiment harness rather than an autonomous optimizer that will rewrite your project on its own, and that distinction matters for how you should try it. The repository is unusually clear about the parts that decide whether a run is useful: gates must exist before the loop starts, a benchmark that measures the thing you care about has to come first, and the frontier strategy is something you choose rather than inherit. It supports eight agent hosts and eight compute backends, so the interesting question is not whether it runs but whether your benchmark is honest. The best first run is discover on a single function with a real test suite wired in as a gate, one round, watched, before letting it fork widely in a worktree per experiment.
Frequently asked questions
What is evo and what does it do to a codebase?
evo is a plugin for agentic coding tools that optimizes code by running experiments in a loop. You give it a repository, it discovers what to measure, sets up a benchmark and gates, then proposes edits, runs the benchmark, keeps improvements and discards the rest.
How do I install evo and get it running?
Install the CLI with uv tool install evo-hq-cli, then run evo install with your host name, one of claude-code, codex, cursor, hermes, kimi, opencode, openclaw or pi. On codex you can pass --no-trust-hooks to review the hooks before approving them.
What are gates in evo and why do they matter?
Gates are pass/fail commands that run on every experiment, and an experiment that fails one is discarded even if its score improves. They prevent the search from cheating by returning constants, skipping work or trading correctness for speed. Any command that exits zero on pass qualifies, such as a test suite or a score floor.
Can evo run experiments in the cloud?
Yes, through optional provider extras. The local worktree, pool and ssh backends are included by default, while Modal, E2B, Daytona, AWS and Azure each need the matching extra, installed as uv tool install with the bracketed provider name.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/evo-hq-evo)