# Hermes Agent Self-Evolution: Automated Skill Optimization via DSPy and GEPA

> Hermes Agent Self-Evolution applies the GEPA algorithm on top of DSPy to automatically mutate and retest Hermes Agent skill files using execution traces, with no GPU training involved. Only Phase 1, covering skill files, is currently implemented; four additional phases covering tool descriptions, system prompts, and code remain planned.

**NousResearch/hermes-agent-self-evolution** — ⚒ Evolutionary self-improvement for Hermes Agent — optimize skills, prompts, and code using DSPy + GEPA

- Repository: https://github.com/NousResearch/hermes-agent-self-evolution
- Stars: 5,425 · Forks: 645
- Language: Python
- License: not declared
- Published: 2026-09-22 · Updated: 2026-09-22 · Language: en
- Canonical page: https://hysenlabs.com/projects/nousresearch-hermes-agent-self-evolution

## What Problem Hermes Agent Self-Evolution Addresses

Improving an AI agent's behavior typically requires engineers to manually rewrite skill files, tool descriptions, or prompts and then run evaluations by hand. That loop is slow and does not systematically explore the space of possible improvements. Hermes Agent Self-Evolution automates that loop by treating each skill file as an organism subject to mutation and selection.

The project targets engineers who operate Hermes Agent deployments and want a way to measure and improve specific skills, such as a GitHub code review skill, without retraining any model weights. The core idea is that better prompts and skill files produce better behavior, and finding them should be a search problem rather than a manual editing task. The README describes the intended user as someone who already has Hermes Agent running and wants to make it measurably better at a defined task.

## How GEPA Reads Execution Traces to Propose Targeted Mutations

The GEPA algorithm (Genetic-Pareto Prompt Evolution), published as an ICLR 2026 Oral paper under an MIT license, operates differently from simple prompt templates that generate random variations. GEPA reads execution traces to understand why a skill failed on a given input, not merely that it failed. It then proposes mutations aimed at the specific failure mode observed in the trace.

The data flow the README documents is: read the current skill file, generate an evaluation dataset (either synthetic or from real session history), pass execution traces to the GEPA optimizer, generate candidate variants, evaluate those variants, and select the best performers. The selection uses a Pareto frontier approach that balances multiple objectives simultaneously rather than optimizing a single metric. Every candidate must pass a full constraint gate before being considered: the test suite must pass at 100 percent, skill files must stay under 15 KB, tool descriptions must stay under 500 characters, the variant must not break caching compatibility, and it must preserve the semantic intent of the original skill. All accepted changes go through human pull request review rather than being committed directly.

The entire process runs over API calls, not GPU inference. The README estimates each optimization run costs between $2 and $10 depending on the number of iterations and the size of the evaluation dataset. DSPy provides the optimization framework that GEPA plugs into.

## Installing and Running a First Skill Evolution

The repository targets Python 3.10 or higher. The core dependencies are DSPy 3.0 or later, the OpenAI client, PyYAML, Click, and Rich. Install from source with the development extras:

```bash
git clone https://github.com/NousResearch/hermes-agent-self-evolution.git
cd hermes-agent-self-evolution
pip install -e ".[dev]"
```

Before running, point the tool at your local Hermes Agent repository:

```bash
export HERMES_AGENT_REPO=~/.hermes/hermes-agent
```

To evolve a skill using synthetic evaluation data, pass the skill name and the number of iterations:

```bash
python -m evolution.skills.evolve_skill \
    --skill github-code-review \
    --iterations 10 \
    --eval-source synthetic
```

The `--eval-source synthetic` flag tells the tool to generate its own evaluation dataset rather than requiring real session history. For production use, the README recommends the `sessiondb` source, which pulls from actual session history recorded by Claude Code, GitHub Copilot, or Hermes itself:

```bash
python -m evolution.skills.evolve_skill \
    --skill github-code-review \
    --iterations 10 \
    --eval-source sessiondb
```

The `sessiondb` path produces a more grounded evaluation because the mutations are tested against real usage patterns rather than synthetic prompts. The output is a candidate SKILL.md that passed all guardrails, ready for human review before merging.

## What Phase 1 Actually Covers and What Remains Planned

The README is explicit about the current state of implementation. Phase 1, which targets SKILL.md files using DSPy and GEPA, is marked as implemented. The remaining four phases are all marked as planned:

Phase 2 will target tool descriptions (the strings that explain each tool to the agent), also using DSPy and GEPA. Phase 3 will target system prompt sections. Phase 4 will target tool implementation code itself, using a different engine called the Darwinian Evolver rather than DSPy. Phase 5 is described as a continuous improvement loop with an automated pipeline.

The Darwinian Evolver, an external CLI tool published under AGPL v3, treats code files as organisms stored in Git branches. Because it is AGPL-licensed, integrating it into a project means understanding what obligations that license creates for derivative works; the README designates it as an external CLI only, which limits how tightly it can be coupled to the rest of the system.

For anyone evaluating this project, the honest picture is that only skill file optimization is available today. The roadmap is ambitious and the architecture is coherent, but three of the five phases exist only as design documents.

## Where This Tool Fails and When to Avoid It

The guardrail system contains a meaningful constraint that limits applicability. Every evolved variant must pass the full test suite at 100 percent. If a skill does not have a test suite, or if the test suite is incomplete, the optimizer cannot verify that a proposed mutation is safe. A skill with poor test coverage is a poor candidate for automated evolution because any mutation that slips past a gap in the tests could introduce a subtle behavior change that only surfaces in production.

The tool is also wrong for optimizing skills that cannot be scored automatically. GEPA needs a signal to drive selection, and that signal must come from evaluation results. If evaluating whether a skill performed well requires human judgment that cannot be captured in an automated benchmark, the optimizer has nothing to select against.

The $2 to $10 per run estimate is reasonable for development iteration, but a team running many iterations across many skills will find costs accumulate. The tool is not designed for continuous automated evolution in production; Phase 5 describes such a pipeline, but it is not implemented. Running frequent optimization cycles against a production skill set without a clear budget and governance process could result in unexpected API costs.

The project has no GitHub releases and was last pushed on 2026-06-17. The development pace is active but the project is still in early stages.

## DSPy Without GEPA as the Nearest Alternative

DSPy on its own, without GEPA, is the nearest functional alternative for automated prompt optimization. DSPy provides its own optimizers, such as BootstrapFewShot and MIPROv2, which can improve prompts by searching over examples and instructions. The difference is in how failures are diagnosed. DSPy's built-in optimizers work from input-output examples and do not read execution traces to understand the mechanism of a failure. GEPA's trace-reading capability is meant to produce more targeted mutations because it acts on the reasons for failure rather than just the presence of failure.

For teams not specifically tied to the Hermes Agent skill format, DSPy alone offers a broader and more documented interface for prompt optimization. It works with any model and does not require a specific agent framework. Hermes Agent Self-Evolution is purpose-built for the Hermes Agent ecosystem and assumes the SKILL.md format, so it cannot be applied directly to arbitrary prompt optimization tasks outside that ecosystem.

## License and Maintenance Context

The project is MIT licensed, which permits use, modification, and distribution without restrictions on commercial use. The pyproject.toml shows dependencies on DSPy 3.0 or later and OpenAI 1.0 or later, which means the upgrade surface follows those projects' release cadences. Neither library has historically been stable across major versions without some changes to calling code, so teams adopting this project should treat dependency updates as a source of maintenance work.

The optional Darwinian Evolver dependency, brought in via the `darwinian` extras group in pyproject.toml, carries AGPL v3 obligations. Using it as an external CLI limits the scope of those obligations compared to linking it as a library, but legal review is appropriate before distribution.

The PLAN.md file referenced in the README documents the full architecture, evaluation strategy, constraints, and phased timeline. Reading it before committing to adopt the project is practical, since it describes what the system is designed to become rather than just what it is today.

## Conclusion

Teams already running Hermes Agent who want to systematically improve SKILL.md files should evaluate this project after Phase 1 stabilizes further. Anyone needing tool description or code evolution will need to wait for Phases 2 through 4. Verify that the target skill has a reliable test suite first, since the guardrail system requires 100 percent test passage before accepting any evolved variant. The MIT license permits commercial use without restrictions on redistribution.

## FAQ

### How is Hermes agent self-improving?

Hermes Agent Self-Evolution runs execution traces through the GEPA optimizer, which reads the traces to understand why specific inputs failed, generates candidate mutations to the skill file, evaluates those candidates against a benchmark, and selects variants that pass all guardrails including the full test suite. Accepted changes go through human pull request review before being merged.

### What does Phase 1 of Hermes Agent Self-Evolution optimize?

Phase 1 targets SKILL.md files, which define Hermes Agent's individual skills. It uses DSPy and GEPA to mutate and evaluate these files, then selects variants that improve benchmark performance while staying within size limits and passing the full test suite.

### What is the approximate cost to run a skill evolution with Hermes Agent Self-Evolution?

The README estimates each optimization run costs roughly $2 to $10, depending on the number of iterations and the size of the evaluation dataset. All computation runs through API calls with no GPU training involved.

## Sources

- [Issues](https://github.com/NousResearch/hermes-agent-self-evolution/issues)
- [NousResearch/hermes-agent-self-evolution on GitHub](https://github.com/NousResearch/hermes-agent-self-evolution)
- [README](https://github.com/NousResearch/hermes-agent-self-evolution/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/nousresearch-hermes-agent-self-evolution
