# Aegis's headline benchmark cannot confirm which model served the runs

> A method pack that makes coding agents plan against your real codebase before they edit, prove completion with fresh evidence, and leave trivial tasks alone. It ships as nine per-host plugin directories plus a patch aimed at someone else's runtime, and it installs by handing one prompt to your agent. The benchmark section is the most valuable thing on the page, because it states its own limitations precisely, including one that most such pages would leave out.

**GanyuanRan/Aegis** — Make AI coding agents architecture-aware: baseline-first, evidence-verified, drift-checked, and safe across long tasks.

- Repository: https://github.com/GanyuanRan/Aegis
- Website: https://github.com/GanyuanRan/Aegis
- Stars: 1,315 · Forks: 69
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/ganyuanran-aegis

## The benchmark could not confirm which model served the requests

The methodology paragraph discloses three limitations in three sentences, and the third is the one that changes how to read everything above it. The page says the evidence is bounded and advisory, that the review was an arm-hidden technical review rather than an independent human review, and that host events did not return the observed model identity. Read that last clause carefully: the harness that ran the arms did not report back which model answered. So the A/B comparison was configured to request one model at one reasoning setting, and the project cannot prove the requests were served by it. Everything else in the section is carefully bounded too, with a 95 percent case-cluster interval given as a range around the point estimate rather than a single number.

> The numbers above are bounded advisory evidence from the frozen benchmark
> below, not a universal-quality or completion-authority claim.

A project willing to print that caveat next to its own headline figure is doing something rare, and it also means the figure is not a measurement of the model it claims to be.

## The headline number is two minor releases old, and the newer snapshot has one observation per case

There are two benchmark snapshots on the page and they are not interchangeable. The one carrying the headline result was run against a specific earlier version on a specific date, across twenty cases. The newer snapshot is for a later version, covers twenty-two cases, and the page says it uses a different run profile and does not measure the changes in this release. One observation per arm per case means there is no within-case variance to average, so a single unlucky run moves the result by a full case. Neither snapshot measures the version currently published. So the evidence supports a claim about a version from the middle of August, the current release has no measured result at all, and the newest available data is too thin to replace it. The project says all of this in three lines rather than burying it.

## The page lists six things the project is not, starting with the platform it is named after

The section before the install instructions is a list of disclaimers, and it is unusually long for a project README. It says the pack is currently a runtime-ready method pack and then enumerates what that is not: not the full platform, not a daemon, not a background runner, not a runtime core, not an authoritative gate decision, not an authoritative policy snapshot, and not final completion authority. It adds the ordering rule that matters most in practice: your instructions and your target project's rules outrank the pack's guidance. That last sentence is the one to hold on to, because a method pack that can outrank your instructions is a policy engine, and this one explicitly is not. It is advice delivered to an agent, ranked below you.

## Installation is one prompt handed to an agent, and the prompt text is cut off mid-sentence

There is no package manager step. The page tells you to give your agent a prompt, and the prompt begins by telling the agent to read a URL, identify which host you are using, and install the pack globally through the right host guide. The visible text then encodes an authority condition, and this is as much of it as the page shows.

```text
Read https://github.com/GanyuanRan/Aegis, identify my current AI coding host, and install Aegis globally using the correct host guide.
```

So the install path is an LLM reading a URL and following instructions whose tail you cannot read on this page, with a nested instruction about when it is allowed to take a fallback. Compare that to a version manager before you compare anything else about the install.

## Part of the install is hand-copied into global rules, and the updater never touches it

There is a global routing prefix, and the page is candid about its status. It is optional, it is a manually copied projection into the host or profile, and it does not install anything or prove that the host can discover the skills. If your host already bootstraps and routes reliably you do not need it; if you do need it, you add it at the very beginning of your existing global user rules without changing anything else. Then the sentence that matters: this copied prefix is not managed by the updater. So a meaningful part of the installation lives in a file the project will never rewrite or repair. The page also tells you that if you previously copied now-retired profiles, you should replace only those old blocks. That is manual surgery in a shared config file, performed by hand, with no tooling.

## Nine per-host plugin directories, and a patch aimed at someone else's runtime

The top level of the tree is where the multi-host story is told. There are plugin directories for four distinct naming conventions, host configuration directories for three, and a plugin manifest for another, alongside the shared skills and extensions. In total nine host-specific integration points, for at least eight different agent hosts, plus an extensions tree with per-host subdirectories for three more. The manifest makes the preferred host explicit. It declares three peer dependencies on one vendor's agent runtime, each with a minimum at a release-candidate version, and each marked optional so the pack does not fail to load on hosts where they are absent. It also declares a bundle patch pointing at a file in the extensions tree, so on that host the integration is a patch applied to the host's plugin bundle rather than a plugin installed alongside it.

## Twenty cases and six runs each is the arithmetic behind two percentages and a zero

The headline result is stated precisely enough that you can reconstruct it. The run covered one hundred and twenty valid runs across twenty cases, with a single variable differing between the two arms. That is six runs per case per arm. The contract pass rate moves from a bit under two thirds to a bit over nine tenths, and unsafe outcomes go from a bit over one in eight to none. Both are presented to two decimal places, which reads like more resolution than twenty cases supports. The confidence interval the page gives for the contract change is a range thirty-five points wide, which is the honest number to quote. The zero is the most striking figure on the page and also the one most sensitive to sample size: it is zero observed events, not a demonstrated impossibility.

## Three README files, and the English link points at the wrong one

The language row at the top offers an English version, a Chinese version, and two fast-track playbook documents in both languages. The tree holds three README files at the top level: the default one, one with an English suffix, and one with a Chinese suffix. The English link in that row points at the default file rather than at the file with the English suffix. Small, but it tells you something about how the localisation was assembled: there is a duplicate rather than a substitution, which is the shape that produces drift. Elsewhere the tree is disciplined in the opposite direction. The benchmark directory commits a sanitised results file, an English table, a Chinese table and a methodology document, so the evidence behind the claim is in the repository rather than behind a link. There is also a version bump manifest and release notes at the root, which is more release hygiene than most projects of this kind manage.

## Conclusion

Aegis is worth reading for two things: the discipline it asks of an agent, and the benchmark write-up, which is more forthcoming about its own weaknesses than most. Treat the headline numbers as advisory rather than settled, for three reasons the page itself supplies. The run could not confirm which model actually served the requests. The measurement was taken two minor releases before this one, and the newer snapshot has a single observation per case. And the result sits on twenty cases. Before installing, know that the install is performed by an agent reading a URL rather than by a package manager, and that part of it is text you copy by hand into your host's global rules and which the updater will never touch.

## FAQ

### What does the Aegis method pack do for a coding agent?

It asks the agent to align with the project's real baseline, including owners, contracts and boundaries, before it edits code, to attach fresh verification evidence to completion claims, and to track or remove retired fallbacks rather than leaving them silently. Trivial requests stay on a fast path so ceremony only appears when a task needs it.

### How good is the evidence behind the Aegis benchmark numbers?

The page bounds it itself. The run could not confirm which model served the requests because host events did not return the observed model identity, review was arm-hidden technical review rather than independent human review, and the author calls the result bounded advisory evidence rather than a universal-quality claim. A 95 percent case-cluster interval is given for the contract change.

### Does the Aegis benchmark measure the current release?

No. The headline result was measured against an earlier version in August across twenty cases. The newer snapshot covers a later version and twenty-two cases but uses a different run profile with one observation per arm per case, and the page says it does not measure the changes in the current release.

### How do I install the Aegis method pack?

By handing one prompt to your coding agent, which reads a repository URL, identifies your host and installs through the matching host guide. There is no package manager step. The page also documents a separate optional global routing prefix that you copy into your host's global user rules by hand and which the updater does not manage.

### Does Aegis update itself automatically?

No. The page states that Aegis does not run background automatic updates by default. Updates go through the agent via a natural-language request or an explicit skill request, and updating every registered host requires an explicit all-arguments request.

## Sources

- [GanyuanRan/Aegis on GitHub](https://github.com/GanyuanRan/Aegis)
- [License: MIT](https://github.com/GanyuanRan/Aegis/blob/main/LICENSE)
- [Project website](https://github.com/GanyuanRan/Aegis)
- [README](https://github.com/GanyuanRan/Aegis/blob/main/README.md)
- [Releases](https://github.com/GanyuanRan/Aegis/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ganyuanran-aegis
