Ponytail: a YAGNI review skill that makes coding agents delete code
Ponytail adds YAGNI-focused review instructions that push coding agents to remove unnecessary abstractions and avoid speculative code.
At a glance
- What is it?
- Ponytail installs as a plugin for Claude Code, Codex, Copilot CLI, OpenCode and other agent harnesses, injecting a seven-rung ladder that pushes the agent toward native features and existing dependencies. The repository reports a 54% mean reduction in lines changed across twelve agentic tasks, with the caveat that a reasoning-heavy model can spend more tokens deliberating the ladder.
- Who is it for?
- Adopt Ponytail if your agent habitually reaches for a library when the platform already ships the feature, and you accept that the plugin's hooks need Node.js on the non-interactive shell's PATH. Skip it if you run a reasoning-heavy model that burns thinking tokens on the ladder, or if your codebase is already minimal and the cut approaches zero.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 15 days ago.
- What is it written in?
- Mainly JavaScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
The over-build problem Ponytail targets
Ask a coding agent for a date picker and the README's own example shows what tends to happen: it installs flatpickr, writes a wrapper component, adds a stylesheet, and opens a discussion about timezones. Ponytail's answer is a single line, an `<input type="date">`, because the browser already has one. The project's stated rule is not "fewest tokens" but "write only what the task needs, and never cut validation, error handling, security, or accessibility." The target user is an engineer running an agent inside a real repository who is tired of reviewing speculative abstractions, unused configuration flags and dependencies pulled in for a capability the standard library or the platform already provides. The repository frames this as YAGNI-focused review instructions rather than a code generator: Ponytail does not write the code, it changes what the agent considers acceptable before writing.
The seven-rung ladder and where it runs
The mechanism is a decision ladder the agent is told to stop at, at the first rung that holds. Rung one asks whether the thing needs to exist at all; rung two asks whether the codebase already has it; then standard library, then native platform feature, then an already-installed dependency, then a one-liner, and only then the minimum that works.
The ordering matters more than the list. Reuse sits above the standard library, and the standard library sits above native platform features, which is why the date picker collapses to an HTML input rather than a small date utility. The README is explicit that the ladder runs after the agent understands the problem, not instead of it: the agent reads the code the change touches and traces the real flow before picking a rung. That distinction separates Ponytail from a bare "write one-liners" prompt, which the benchmark table scores at 95% safe against Ponytail's 100%, meaning the bare prompt dropped a safety guard somewhere in the run. The repository also states that four categories are never on the chopping block: trust-boundary validation, data-loss handling, security and accessibility.
Installing Ponytail in Claude Code and running a first review
Claude Code installs from the plugin marketplace. The README notes that you have to send two separate prompts for the install to work, which is a real friction point rather than a stylistic note: issuing both lines in one message does not complete the install.
/plugin marketplace add DietrichGebert/ponytail/plugin install ponytail@ponytailThe same two commands work in the Claude Code Desktop app's Code tab, typed into the prompt box, or reached through the + button next to it under Plugins, then Add plugin. After installing, the plugin activates the ruleset for subsequent turns. The repository does not document a rollback command, so removing it means going back through the same plugin management surface.
Codex uses a CLI path instead, followed by a hook review that the agent host requires before the plugin runs.
codex plugin marketplace add DietrichGebert/ponytail
codex plugin add ponytail@ponytailAfter running those, the README says to open `codex`, run `/hooks`, review and trust the two lifecycle hooks, and start a new thread. The Codex desktop app picks up the same install after a restart.
For OpenCode the plugin is declared in `opencode.json` rather than installed through a marketplace, and the ruleset is injected every turn at the active level.
{ "plugin": ["@dietrichgebert/ponytail"] }A checkout-based variant points at the local file and reuses the repository's own `hooks/` and `skills/` directories.
{ "plugin": ["./.opencode/plugins/ponytail.mjs"] }GitHub Copilot CLI follows the same marketplace shape as Codex, and namespaces its commands by plugin name, so an invocation reads `/ponytail:ponytail ultra` or `/ponytail:ponytail-review`. The Pi agent harness installs with `pi install git:github.com/DietrichGebert/ponytail`. The Claude Code and Codex plugins run two small Node.js lifecycle hooks, so `node` must be on your PATH; the README calls out Nix and nvm users specifically, because the hook needs the non-interactive shell's PATH, not the interactive one. If Node is missing, the skills still work and only the always-on activation goes quiet.
What the benchmark numbers do and do not cover
The headline figure is a 54% mean reduction in lines changed across twelve feature tickets, measured on a headless Claude Code session editing tiangolo's full-stack-fastapi-template, with Haiku 4.5 and n=4. The same table reports 22% fewer tokens, 20% lower cost and 27% less time, against a no-skill baseline. Two things about that table deserve attention. First, the README corrects its own earlier claim: a previous single-shot benchmark reported 80-94% as a flat figure, and the repository acknowledges in issue #126 that the bare-model baseline padded its answer with prose and options, so part of that gap was a conversational artifact. The agentic numbers replace it. Second, the reduction is not uniform. The README states it reaches 94% where an agent over-builds, citing a date picker falling from 404 lines to 23 and a color picker from 287 to 23, and is near zero where the code is already minimal. A 54% mean over a distribution with a 94% ceiling and a near-zero floor tells you the average is carried by the over-build cases. If your tickets look like the floor, the plugin will not do much.
Where Ponytail makes things worse
The README names its own failure mode: a terse reasoning model that spends thinking tokens deliberating the rungs can go the other way, and it states this happens on GPT-5.5. That is a genuine cost regression, not a rounding error, and it follows from the design. A ladder with seven ordered questions is more instruction to reason about than "be concise," so a model that thinks out loud before every edit pays for the structure it is being asked to follow. The second limitation is scope. The measured cut is near zero on code that is already minimal, so a repository with disciplined contributors and no speculative abstractions has little to gain and still takes on two lifecycle hooks, a plugin dependency and a Node.js PATH requirement. Third, the two-prompt install in Claude Code is easy to get wrong, and the README does not document an uninstall or rollback path, which matters if you are evaluating the plugin on a shared machine. Finally, the benchmark covers one repository, one model family and twelve tickets; the README points readers to benchmarks/results/2026-06-18-agentic.md for per-task tables and limitations rather than claiming the numbers generalize.
How it differs from a terse-prompt control like caveman
The benchmark includes a control arm called caveman, described as a terse-prose control, and the comparison is instructive because both projects try to shrink agent output but attack it from opposite ends. Caveman targets how the agent talks: shorter prose, fewer explanations. Ponytail targets what the agent decides to build, with the ladder as the decision procedure and the prose style left alone. The measured result separates them. Caveman cut lines by 20% but increased tokens by 7%, cost by 3% and time by 2% against the no-skill baseline, while Ponytail cut all four. That pattern makes sense: suppressing explanation does not stop an agent from installing flatpickr, and it can make the agent spend effort compressing text that was never the expensive part. Ponytail's cut comes from skipping the dependency and the wrapper component entirely. If your complaint is that the agent rambles in chat, a terse-prose skill addresses that; if your complaint is the diff, the ladder is the relevant lever.
Licence and the cost of keeping up
Ponytail is MIT licensed, which permits commercial use, modification and redistribution provided the copyright notice and permission notice are retained. The repository ships a LICENSE file at the top level, and package.json declares `"license": "MIT"` alongside a publishConfig that makes the npm package public. Note that the npm package name is `@dietrichgebert/ponytail` while the repository directory is `DietrichGebert/ponytail`; the two are not interchangeable in commands. Licence terms cover the plugin code and the skills, and nothing in the repository suggests a separate commercial tier or a hosted service. On maintenance, the last push was on 2026-08-07, the same date as the v4.9.0 release described as 53 commits of doing less, so the project is being changed rather than frozen. The upgrade cost is the interesting part: because the plugin injects instructions into every turn, a version bump can change agent behaviour across your whole team without any code change in your repository. Release notes for v4.8.3 and v4.8.4 mention making the plugin lazy in subagents and in Hermes respectively, which is the kind of change that alters output on runs you did not touch. Pinning a version and reading the release notes before bumping is the practical posture, and the repository's tests run through `node --test tests/*.test.js` plus the `pi-extension` and `ponytail-mcp` suites if you want to check a checkout before rolling it out.
Editorial conclusion
Adopt Ponytail if your agent habitually reaches for a library when the platform already ships the feature, and you accept that the plugin's hooks need Node.js on the non-interactive shell's PATH. Skip it if you run a reasoning-heavy model that burns thinking tokens on the ladder, or if your codebase is already minimal and the cut approaches zero. Before trusting it, open benchmarks/results/2026-06-18-agentic.md and check whether the twelve FastAPI and React tickets resemble your own work, then run the two /plugin commands and watch one diff.
Frequently asked questions
What is Ponytail in AI?
Ponytail is a plugin that adds YAGNI-focused review instructions to coding agents, pushing them to remove unnecessary abstractions and avoid speculative code. It works by having the agent stop at the first rung of a seven-step ladder before writing anything.
How do you use Ponytail in Claude Code?
Add the marketplace with `/plugin marketplace add DietrichGebert/ponytail`, then install with `/plugin install ponytail@ponytail`. The README states you have to send the two commands as two separate prompts for the install to work.
How do you install the Ponytail extension for Claude?
In Claude Code, run the two `/plugin` commands from the README in sequence. The same steps work in the desktop app's Code tab, either typed into the prompt box or through the + button, Plugins, then Add plugin.
How do you use the Ponytail skill in Codex?
Run `codex plugin marketplace add DietrichGebert/ponytail` and `codex plugin add ponytail@ponytail`, then open `codex`, run `/hooks`, and review and trust the two lifecycle hooks before starting a new thread.
Does Ponytail work in Cursor?
The repository contains a `.cursor/` directory and a `scripts/cursor-hooks.js` file in the published package, but the README section on Cursor is truncated in the available material, so the exact setup steps are not confirmed here.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/dietrichgebert-ponytail)
Community notes