old-coder: making coding agents prove their work instead of asking you to read it
An old coder's strategy for the agent era: don't read the code — make it run the gauntlet. Evidence-first development skill for coding agents, inspired by Uncle Bob.
At a glance
- What is it?
- AmazingAng/old-coder is a markdown skill that wraps a coding agent in a SPEC, a gauntlet of checks and an EVIDENCE report, so you review two documents instead of every diff. It is MIT-licensed, Python-based in its demo, and last pushed on 2026-08-18.
- Who is it for?
- Adopt old-coder if you already run a coding agent on work where a wrong answer costs more than a review cycle, and if you are willing to sign off on a SPEC before any code exists. Skip it if your tasks are one-line edits, or if you have no test runner for the gauntlet to invoke, because the skill produces evidence from checks that must already exist in your project.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 43 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem old-coder addresses: review throughput, not code generation
Coding agents can produce more code than a human can read. The README quotes Uncle Bob's strategy directly: not reading agent-written code, and instead surrounding the agent with constraints such as unit tests, gherkin tests, QA procedures, quality metrics, mutation testing and coverage. The project's own framing of the consequence is blunt: if you are not going to read the code, the things you do read have to carry the trust instead.
So old-coder is not a linter, a test framework or a CI service. It is a set of instructions, written in markdown, that changes what an agent is expected to hand back. The audience is a developer who already delegates implementation to Claude Code, Codex CLI, Cursor, Aider or a custom agent loop and who wants a reviewable artifact rather than a diff. The repository is Python-based in its demo, licensed MIT, and the last push was on 2026-08-18.
SPEC, RED, GREEN, REFACTOR, GAUNTLET, EVIDENCE: the actual loop
The README's flowchart runs SPEC then RED then GREEN then REFACTOR then GAUNTLET then EVIDENCE, with a dotted arrow from REFACTOR back to RED for the next behavior. Two of those stages are documents you read. SPEC comes before any code and contains concrete examples of what the code must and must not do, plus the tools the agent wants to install. Approving it is described as the single yes/no you give. EVIDENCE comes after and holds real numbers from one final fresh run, which the README says you can rerun yourself with a single command.
The middle stage is a table of checks, each mapped to a question. Full test suite asks whether anything broke. Types, lint and complexity ask whether there are obvious mistakes or unreadable tangles. Changed-line coverage asks whether every new line is exercised by a test. Mutation testing plants bugs on purpose to see whether the tests catch them. Property-based tests push hundreds of random inputs at the rules. Real execution checks the code runs outside the test harness. Supply chain and secrets checks whether the agent pulled in risky packages or leaked a key. Suite health checks whether the tests themselves are stable, in any order.
Domain layers are picked per task from a risk model documented in references/gauntlet.md, and effort scales with risk: a typo fix runs a couple of checks, while anything touching money, logins, data or concurrency runs everything and the agent attacks its own code with hostile inputs first. The design choice worth noting is that the skill does not define new tooling. It orchestrates checks your project already has, which is why it is portable across agents and also why it is only as strong as the test suite underneath it.
Installing old-coder and reading your first SPEC
The primary install path is the skills CLI, which fetches the skill from the repository by URL and name. Run it from the project where you want the skill available.
npx skills add https://github.com/amazingang/old-coder --skill old-coderThe same repository ships a second skill, old-coder-api, for HTTP/JSON API design and review. Install it when you want gates for compatibility, authorization, idempotency, pagination, rate limits and operability, or install both in one command.
npx skills add https://github.com/amazingang/old-coder --skill old-coder --skill old-coder-apiWhen both apply, the README assigns responsibilities: old-coder owns workflow, approval and evidence, while old-coder-api owns the API contract, and its gate decisions become SPEC constraints and gauntlet checks.
For Claude Code there is a manual route that copies the skill directory into a skills folder, after which you invoke /old-coder or let it trigger on high-assurance requests. The README gives both a user-level and a project-level location.
cp -r skills/old-coder ~/.claude/skills/
# or copy it to <project>/.claude/skills/Other agents have no skill loader. The README says to add skills/old-coder/SKILL.md to your AGENTS.md, rules file or system prompt, and to keep its references/ directory alongside it. That last part matters: SKILL.md points into references/gauntlet.md for the risk model, so moving the file alone leaves the agent without the layer menu.
The most concrete first use in the repository is the demo, a rate limiter built end to end under the skill. Its requirements-dev.txt is installed into a virtual environment and the package installed in editable mode, then the gauntlet script runs.
cd demo-rate-limiter
python3 -m venv .venv && .venv/bin/pip install -r requirements-dev.txt -e .
./tools/gauntlet.shThe README reports the demo's evidence.md as 41 tests, 100% coverage (49/49 statements and 20/20 branches) and 22/22 planted bugs caught. Those are the project's own numbers for its own demo, not a benchmark of your codebase.
The honesty rules, and the limit the README states plainly
Because the agent grades its own homework, the skill imposes rules on the report. Never weaken a test to make it pass. Never report a check that did not run. Anything unverified is labeled unverified, never pass. If no human approved the spec, the report must say so and claim less confidence. These are instructions to the agent, not enforcement, and the README does not claim otherwise.
The stated limit is the most interesting sentence in the repository. The gauntlet turns the constraints expressed in the spec into executable evidence; it cannot prove the spec is complete, and it cannot authenticate its own checkers and mappings. That is the argument for the human approval step on SPEC, and it is why EVIDENCE is described as bounded, auditable confidence rather than absolute proof.
The demo makes the point sharper than prose can. Fresh-context verification of earlier green states still found real behavioral defects and an unsound mutation runner, according to the README. A green gauntlet is not self-authenticating. If your plan is to approve a SPEC once and then trust every subsequent report without reading it, the project's own demo is the counterexample.
Where old-coder is the wrong tool
The skill is a workflow wrapper. It cannot create the checks it runs. If your project has no test suite, no type checker, no linter and no way to execute the code outside a test harness, the gauntlet has almost nothing to invoke, and EVIDENCE will be a short document full of unverified labels. That is honest, but it is not useful.
Risk scaling cuts the other way too. A one-line typo fix runs a couple of checks; there is little reason to route that through SPEC approval and a full report. The overhead is the point on high-assurance work and pure friction on trivial work.
There is also a portability cost. The install path through npx skills is clean, but the manual path for other agents means pasting SKILL.md into a rules file and keeping references/ next to it. That is a copy you now maintain by hand, and the README does not document an upgrade command for the manual route. The repository also does not describe a rollback procedure for a skill that has been copied into a project. Finally, the skill cannot verify that the agent followed its instructions. The rules against weakening tests and reporting unrun checks are constraints on behavior, and the README's own admission that checkers and mappings are not authenticated is the boundary.
old-coder-api versus the main skill, and the mutation-testing layer versus plain coverage
The clearest internal alternative is old-coder-api. Both ship from the same repository and install through the same command, but they answer different questions. old-coder owns the workflow, the approval gate and the evidence report. old-coder-api owns the API contract and supplies gates for compatibility, authorization, idempotency, pagination, rate limits and operability. Its gate decisions are fed back in as SPEC constraints and gauntlet checks. If you are building an HTTP service, installing only the main skill leaves the contract-level checks on the table.
A second comparison lives inside the gauntlet table. Coverage and mutation testing are listed as separate checks because they answer different questions. Changed-line coverage asks whether every new line is exercised by a test. Mutation testing plants bugs on purpose and asks whether the tests catch them. A suite can reach full line coverage and still miss a planted bug, which is why the demo reports both 100% coverage and 22/22 planted bugs caught rather than coverage alone. If your existing pipeline stops at coverage, the mutation layer is the part of old-coder that adds information you do not already have.
For teams already running spec-driven development, the difference is where the spec lives. Here it is a document the agent writes and a human approves before RED, and the approval is recorded in the EVIDENCE report's confidence claim. That is a narrower and more auditable arrangement than a spec that exists only in a ticket.
Maintenance, upgrade cost and the MIT licence
The repository is not archived, and the last push was on 2026-08-18. The only release listed is v0.1.0, described as the first stable protocol, published on 2026-08-15. A single release three days before the last push is a young project by any measure, and the README does not describe a deprecation policy, a versioning scheme for the skill files, or a migration path between releases.
Upgrade cost depends on how you installed it. Through npx skills, updating means rerunning the add command against the repository. Through the manual copy into ~/.claude/skills/ or <project>/.claude/skills/, or through pasting SKILL.md into an AGENTS.md, you are the upgrade mechanism. The README does not document a rollback for either route, so keeping your own copy under version control is the only recovery path the README supports.
The licence is MIT, stated in the README and present as a LICENSE file at the repository root. MIT is permissive and imposes no copyleft obligation on your project, but this is a description of the licence text, not legal advice. If you redistribute the skill inside a commercial product, read the LICENSE file yourself. The repository also contains CONTRIBUTING.md, so contributions are accepted under whatever terms that file sets.
Editorial conclusion
Adopt old-coder if you already run a coding agent on work where a wrong answer costs more than a review cycle, and if you are willing to sign off on a SPEC before any code exists. Skip it if your tasks are one-line edits, or if you have no test runner for the gauntlet to invoke, because the skill produces evidence from checks that must already exist in your project. Before trusting a report, rerun the demo with ./tools/gauntlet.sh from demo-rate-limiter and read the disclosed defects in its evidence.md, which is the clearest statement of what a green gauntlet does and does not prove.
Frequently asked questions
What is the old-coder skill and which agents does it work with?
It is a markdown skill that makes a coding agent produce a SPEC before coding and an EVIDENCE report after, so you review two documents instead of the code. The README says it works with any agent that follows instructions, naming Claude Code, Codex CLI, Cursor and Aider.
How do I install old-coder?
The README's primary path is npx skills add https://github.com/amazingang/old-coder --skill old-coder. For Claude Code you can instead copy skills/old-coder into ~/.claude/skills/ or a project-level .claude/skills/; for other agents you add skills/old-coder/SKILL.md to your AGENTS.md, rules file or system prompt and keep its references/ directory alongside it.
Does old-coder prove that agent-written code is correct?
No. The README states that the gauntlet cannot prove the spec is complete and cannot authenticate its own checkers and mappings, which is why a human approves the SPEC. EVIDENCE is described as bounded, auditable confidence rather than absolute proof.
What is the difference between old-coder and old-coder-api?
Both ship in the same repository. old-coder owns workflow, approval and evidence; old-coder-api owns the API contract and adds gates for compatibility, authorization, idempotency, pagination, rate limits and operability. When both apply, the API skill's gate decisions become SPEC constraints and gauntlet checks.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/amazingang-old-coder)