# Ponytail puts a seven-rung YAGNI ladder in front of your coding agent

> This is a set of review instructions packaged as a plugin, a skill directory and a set of harness hooks, whose whole job is to make a coding agent stop building things the task does not need. Its headline benchmark was corrected once after a bad baseline, and one of its own caveats says the cost saving reverses on reasoning models.

**DietrichGebert/ponytail** — Ponytail adds YAGNI-focused review instructions that push coding agents to remove unnecessary abstractions and avoid speculative code.

- Repository: https://github.com/DietrichGebert/ponytail
- Website: https://ponytail.dev
- Stars: 147,444 · Forks: 7,917
- Language: JavaScript
- License: MIT
- Published: 2026-08-08 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/dietrichgebert-ponytail

## Two prompts install it, and a missing node leaves it silently inactive

The install is two commands in Claude Code, and they have to be sent as two separate prompts.

```
/plugin marketplace add DietrichGebert/ponytail
```

```
/plugin install ponytail@ponytail
```

The same shape exists for Codex, where the two lifecycle hooks are reviewed and trusted through the /hooks screen before you start a new thread.

```bash
codex plugin marketplace add DietrichGebert/ponytail
codex plugin add ponytail@ponytail
```

Now the caveat, and it is the one that will waste your afternoon. The Claude Code and Codex plugins, and the Cursor hooks, run two small Node.js lifecycle hooks, so node has to be on your PATH. The note is specific about who this bites: for Nix and nvm users, it must be on the non-interactive shell's PATH.

If it is not, the failure mode is silence. The skills still work. The always-on activation just stays quiet instead of erroring on every prompt. You get a session with no obvious symptom and no automatic behaviour, and the only way to find out is to check the hook yourself.

## The npm tarball ships four harness directories out of a dozen

The manifest has a files array, and it is a per-harness allowlist rather than a whole-tree publish.

```json
"files": [
  "AGENTS.md",
  "hooks/",
  "skills/",
  ".opencode/",
  ".qoder/",
  ".qoder-plugin/",
  "pi-extension/",
  "scripts/uninstall.js",
  "scripts/cursor-hooks.js",
  "assets/",
  "LICENSE"
]
```

The git tree at the top level holds a lot more than that: .claude-plugin/, .codex-plugin/, .devin-plugin/, .grok-plugin/, .kiro/, .windsurf/, .cursor/, .clinerules/, .openclaw/, .agents/, plus gemini-extension.json, plugin.json, plugin.yaml, opencode.json, commands/, docs/, benchmarks/ and examples/.

None of those are named in the array, so a registry install is a smaller tree than a clone. That matters because the documented install path is a git marketplace, not npm, so the two routes do not give you the same files. Anyone filing a bug against Claude Code or Codex support should say which route they used, because the same symptom has two different sets of files behind it.

## The root test script chains three suites, two of which live in sub-packages

There is one test script, and it is a chain.

```json
"test": "node --test tests/*.test.js && npm test --prefix pi-extension && npm test --prefix ponytail-mcp"
```

Read it as three suites that must all pass. The first is the repository's own tests, run by the Node test runner against a glob in tests/. The second recurses into pi-extension with its own test command. The third recurses into ponytail-mcp, which is the component that carries a root-level __init__.py in a package whose recorded primary language is JavaScript, so the MCP side is not the same language as the plugins.

Two consequences for anyone picking this up. A fresh clone where you installed only the root dependencies fails at step two, because the sub-packages have no node_modules of their own, and the error points at a sub-package rather than at the install you skipped. And the chain is joined with &&, so the first failure stops the run and the other two suites never execute. A red root test tells you one thing went wrong, not how many.

## Rung two is reuse, and reuse only works if the agent read your code

The mechanism is a ladder, and the agent is supposed to stop at the first rung that holds.

```
1. Does this need to exist?   → no: skip it (YAGNI)
2. Already in this codebase?  → reuse it, don't rewrite
3. Stdlib does it?            → use it
4. Native platform feature?   → use it
5. Installed dependency?      → use it
6. One line?                  → one line
7. Only then: the minimum that works
```

The stated ordering rule is that the ladder runs after the agent understands the problem, not instead of it. It reads the code the change touches and traces the real flow before picking a rung. The summary of that is that it is lazy about the solution and never about reading.

This is where the real cost of the skill sits. Rungs two through five are all reuse decisions, and a reuse decision is only as good as the search behind it. In a codebase the agent cannot navigate, the ladder tells it to prefer what is already there, and it will either miss a helper and write a duplicate, or find the wrong one and use it. The benchmark that produces the headline number was scored on one repository, a FastAPI plus React template, so the reading step was in a codebase of a size the authors chose.

## On GPT-5.5 the cost saving reverses, and the benchmark is all Haiku

The README is unusually direct about where the economics stop working, and the sentence is worth quoting because it undercuts the headline.

The rule was never fewest tokens. It is to write only what the task needs and never cut validation, error handling, security or accessibility. Lower cost and latency are described as a side effect on the models that follow the ladder, and then the caveat: a terse reasoning model that spends thinking tokens deliberating the rungs can go the other way, and on GPT-5.5 it does.

So the cost argument is model-specific, and the project says so. If your agent is a reasoning model, the seven rungs are deliberation you pay for in thinking tokens, and the net can be a loss.

Now put the measurement against that. The benchmark is twelve feature tickets, the same agent with and without the skill, n=4, on Haiku 4.5, scored on the git diff against one real repository. Every efficiency figure in the README comes from one non-reasoning model on one codebase with four runs per arm. If you are not running Haiku, the percentages do not transfer, and the project has told you the direction they go.

## A plain YAGNI prompt is close, and it costs you a safety guard

The comparison table has three arms, and the third one is the interesting one.

Against a no-skill baseline, ponytail comes out at -54% lines of code, -22% tokens, -20% cost, -27% time, 100% safe. A terse-prose control called caveman gets -20% lines but +7% tokens, +3% cost and +2% time. A prompt that simply says YAGNI plus one-liners gets -33% lines, -14% tokens, -21% cost, -30% time, and 95% safe.

That last cell is the point. The naive instruction, the one a person types in ten seconds, produces numbers close enough to the skill's on the efficiency columns and is the only arm that fails a safety check. The README states the reason elsewhere: ponytail keeps every safety guard while a bare write one-liners prompt drops one.

The override list is trust-boundary validation, data-loss handling, security and accessibility, and those are named as never on the chopping block. So if you were going to solve this with a system prompt, you would be reproducing most of the token savings and giving up the guard. What the plugin adds is not the idea, it is the part that survives contact with a model under a deadline.

## The old 80-94% headline was a baseline artefact, and the fix is stated in the badge

The repository corrected its own headline, and the correction is visible rather than buried.

The earlier single-shot run reported 80-94% less code as a flat figure, measured on five everyday tasks across three models with ten runs, counting the lines of a single answer. Issue 126 pointed out that the bare-model baseline pads its answer with prose and options, so part of that gap was an artefact of how a chat model answers a one-shot prompt rather than a difference in code. The README now says that against a fair agentic baseline, 94% is the per-task ceiling and not the average, and that the agentic numbers are the corrected and defensible version.

The single-shot run is still reproducible, with one command and an API key.

```bash
npx promptfoo eval -c benchmarks/promptfooconfig.yaml
```

Two things follow. The project is willing to retract a number in its own README, which is the reason to believe the number now in the README. And the reproduction is gated on .env.example, which asks for an ANTHROPIC_API_KEY and notes that promptfoo reads it automatically, so verifying the claim costs money and time rather than being a free check. The cheap check is examples/, twelve short before-and-after write-ups, if you want to know what the instruction does to your own code.

## Conclusion

Adopt it if your agent habitually over-builds on frontend-shaped tasks, which is the exact case the benchmark was measured on, and if you run a non-reasoning model where the token saving is real. Do not adopt it for a codebase that is already minimal, where the README puts the reduction near zero, and do not adopt it for the cost argument if you are on a reasoning model, because the project itself reports the cost going the other way on GPT-5.5. Check two things first: that node is on your non-interactive shell's PATH, since a missing node leaves the skill silently inactive rather than erroring, and that the harness directory you rely on is in the package.json files array, because the npm tarball ships a smaller set of plugin directories than the git tree holds.

## FAQ

### What is ponytail in AI?

It is a set of review instructions for coding agents, described as a lazy senior dev mode whose goal is to remove unnecessary abstractions and speculative code. The agent works down a seven-rung ladder, from does this need to exist through reuse what is already in the codebase, standard library, native platform feature, installed dependency and one line, before writing the minimum that works.

### How do I use ponytail in Claude Code?

Send two separate prompts: `/plugin marketplace add DietrichGebert/ponytail` and then `/plugin install ponytail@ponytail`. The plugin runs two small Node.js lifecycle hooks, so node has to be on your PATH, and for Nix or nvm users it has to be on the non-interactive shell's PATH. Without it the skills still work and the always-on activation simply stays quiet.

### How do I install the ponytail extension?

Each harness has its own two-step plugin install, using the plugin slug ponytail@ponytail. Claude Code uses the slash commands above, Codex uses `codex plugin marketplace add DietrichGebert/ponytail` followed by `codex plugin add ponytail@ponytail` and then trusts the two lifecycle hooks through /hooks, and GitHub Copilot CLI uses the equivalent `copilot plugin` commands.

### What does the ponytail skill do to my code?

It makes the agent skip work the task does not require and prefer things that already exist. In the README's example, asking for a date picker produces a native input element with a comment noting the browser already has one, instead of installing a date library and writing a wrapper component with a stylesheet.

### How do I use ponytail in Codex?

Run `codex plugin marketplace add DietrichGebert/ponytail` and `codex plugin add ponytail@ponytail`, then launch codex, open /hooks, review and trust its two lifecycle hooks, and start a new thread. Restarting the Codex desktop app after installing also picks the plugin up.

## Sources

- [Official documentation](https://ponytail.dev)
- [Official README](https://github.com/DietrichGebert/ponytail#readme)
- [Project repository](https://github.com/DietrichGebert/ponytail)
- [Release notes](https://github.com/DietrichGebert/ponytail/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/dietrichgebert-ponytail
