# Turn one feature into five pull requests an agent can follow

> Shotgun is a spec-first tool for coding agents whose premise is that agents fail large features by forgetting context, rebuilding what exists and drifting off spec halfway through, so it indexes the repository with tree-sitter, proposes a plan you confirm step by step, and exports file-by-file instructions per staged pull request, with a Rust-pinned dependency dictating a Python ceiling and analytics keys embedded in the container image at build time.

**shotgun-sh/shotgun** — Spec Driven Development 🤠 Write codebase-aware specs for AI coding agents so they don't derail.

- Repository: https://github.com/shotgun-sh/shotgun
- Website: https://shotgun.sh/
- Stars: 686 · Forks: 38
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/shotgun-sh-shotgun

## The failure mode is a ten-thousand-line pull request

The problem statement is the product, and it is specific about how agents fail.

The claim is that agents are good at small tasks and derail on large features, and the three failure modes are named: they forget context, they rebuild things that already exist, and they go off spec halfway through.

All three are symptoms of the same cause, which is a single prompt trying to carry a whole feature. There is nowhere to re-read the plan, and nobody is reviewing the diff until the end.

So the tool does the planning first. It reads the codebase, plans the full feature up front, and then splits that plan into staged pull requests, each carrying file-by-file instructions an agent can actually follow. The stated outcome is that instead of one enormous pull request nobody reviews, you get five focused ones that ship.

That framing explains the whole product. The export format is not documentation, it is instructions addressed to an agent, and the unit of work is a pull request rather than a task. Everything else in the tool is arranged around making that split safe.

It works with the main coding agents, and the licensing is either bring your own key or prepaid usage credits at cost, so there is no per-seat model decision to make.

## Planning asks before it touches a file

There are two execution modes, and the difference is whether you get asked.

Planning is the default. The tool proposes an execution plan, shows you each step, and asks for confirmation before running agents that change files. You get checkpoints, you can refine the plan, and you can confirm or skip a cascaded update when one change affects other documents.

That last clause is the interesting one, because it names a real failure mode in spec-driven work. A plan that changes one document usually implies changes to others, and a tool that silently rewrites them is a tool that destroys work you wrote by hand. Confirming or skipping each cascade puts a human between the agent and a file you care about.

Drafting runs the whole plan in one pass with no intermediate confirmations. Progress is still tracked internally, but you are not prompted at each step.

The interface is a terminal application that opens automatically, with a key combination to switch modes and a slash key for a command palette. The documentation recommends the terminal interface over the command line, and there is a separate command line reference for people who need it.

So the safe default is the slow one, and the fast one is a deliberate opt-in with a keypress rather than a flag you forget about.

## Five sub-agents you never configure

The internal pipeline is five stages, and the documentation is emphatic that you do not select or manage them.

A research stage explores and understands. A specification stage defines requirements. A planning stage produces a roadmap. A task stage breaks the roadmap into steps. An export stage formats the result for an agent to consume.

That is the tool's whole argument in one table. Each stage has a different job and a different failure mode, so giving each one a focused prompt is what keeps a research step from trying to write a plan.

And then the sentence that defines the surface: planning and drafting are the only execution modes you control, everything else is handled by the router.

So there is exactly one knob, and it is a knob about how much autonomy you want rather than about what the agent does.

The usage tips section reinforces the same idea from the human side. Ask it to research how authentication is handled rather than jumping to building. Tell it to ask you questions first. Say that you are working on payments and need refunds, rather than adding refunds with no context. The tool is designed to be talked to, and the three examples are all about supplying context you have and it does not.

The keyboard table adds a usage-statistics shortcut, which tells you cost is something the tool expects you to watch.

## The index is tree-sitter, so the language list is fixed

The first step of every run is codebase indexing, described as building a searchable graph of the whole repository. The implementation is in the dependency list, and it is more specific than the documentation.

There are parser packages for five languages: Python, JavaScript, TypeScript, Go and Rust, all built on the same parser library. So the index is a syntax tree index rather than an embedding index, and the languages it handles well are the five with grammar packages installed.

The reason this matters is the problem the tool claims to solve. Rebuilding things that already exist is what an index prevents, and an index built from syntax trees is very good at finding definitions, references and structure. It is much weaker at finding something that is conceptually the same but named differently. So the tool removes one class of agent failure thoroughly and mitigates another partially, and a repository in a language without a grammar package gets less of both.

A file-watching dependency is also present, which suggests the index can follow changes rather than being a one-shot snapshot.

And there is one pinned exact dependency that shows up as a constraint on the whole project: a database library pinned to a specific version, with the packaging metadata carrying an upper Python bound attributed to that package lacking wheels for the newest Python. The same ceiling appears in the support table, alongside a 32-bit row marked unsupported.

## One dependency is pinned for a patch file, and it says so

Two dependency lines in the manifest carry comments that explain themselves, and both are worth reading.

The agent framework is pinned to an exact version with the reason given as a bug workaround, and the comment names the file implementing that workaround by its filename. So the pin is not conservatism: there is a patch in the tree that stops working when the framework moves.

That is an honest and common kind of pin, and naming the file is better than saying workaround. It also tells you the exact thing to watch when you upgrade, because the pin and the patch have to move together.

The other model client is bounded rather than exact, with a floor and a ceiling. A ceiling on a model client is unusual and is usually about a breaking change in a specific generation, so the upper bound is a deliberate stop rather than a lag.

Two more entries in the list explain themselves without comments. A token-counting library and a sentence-piece tokeniser, which together mean the tool counts tokens itself rather than trusting a provider's count. A pricing data package, which is how the usage-statistics shortcut can show cost rather than token counts. A retry library, so the model calls back off.

And a gitignore-pattern library, commented as being for gitignore matching, which tells you the workspace respects ignore files. That matters for an indexing tool: a repository that ignores its own build output should not have that output in the index.

## Analytics keys are baked into the image, and a flag decides if that fails

The container build tells you more about the project's posture than the feature list does.

Three build arguments are declared, and all three concern analytics and validation: two keys for the analytics platform and one flag that decides whether a missing key fails the build. The comment says the keys are used during the build process to embed them, so they end up inside the image rather than being supplied at runtime.

That design has a consequence worth stating plainly. An image built once and pushed to a registry carries its analytics identity forever, and anyone pulling it shares that identity unless they override it. The environment template says as much, describing the prefixed variables as runtime overrides.

So there are two escape hatches, both documented: build with your own keys, or override at runtime. Neither is optional in a regulated environment, which means the default build is not usable there.

The validation flag is the part that makes the arrangement tolerable. It is skipped by default so local development and continuous integration jobs do not need secrets, and enabled in production builds so a misconfigured release fails rather than shipping an image with no analytics. That is a reasonable default, and it is worth noticing that the skip is the default and the strictness is opt-in.

The observability tool's token is handled differently again: it is only embedded in builds whose version string marks them as pre-release, so a stable release does not contain it. The rest of the image is unremarkable and careful, with a non-root user, a frozen dependency sync, and only the specific source directories copied in.

## One command to run, and a migration note from pipx

Installation is a single command on macOS and Linux, run through an execution tool rather than a package manager.

You install that tool, with a package manager or a piped installer, and then run the project through it at a pinned channel:

```bash
brew install uv
uvx shotgun-sh@latest
```

The reason given for using that tool rather than pip is speed and binary wheels, with the specific claim that it avoids compile and build-tool errors.

That claim is the whole reason a project with a native dependency at several exact pins chooses this route. A user who does not have a compiler installed should still be able to run it.

Windows gets a longer path: a one-time execution policy change scoped to the current user, the tool installed from a PowerShell installer, the binary directory prepended to the path, and then the same run command. There is also an optional administrator step that fetches a Visual C++ redistributable, described as being for code indexing.

And a warning that matters more than any of it: run it in PowerShell, not in the older command prompt and not in a developer shell. That is the kind of instruction that exists because a specific failure was reported often enough to write down.

The repository also carries a migration note from the tool the project used to be installed with, which dates the switch to this execution tool. The lockfile is committed and the container syncs with it frozen and without development dependencies, so a given commit produces the same tree.

## Conclusion

It fits a team already using a coding agent on features larger than a single prompt, where the failure is not the agent's ability but the absence of a plan it can re-read and a diff small enough to review. Three things to check first. Planning is the default and it asks before running agents that change files, so leave it there until you trust the plan; Drafting removes those checkpoints. The indexer is built on tree-sitter with grammar packages for five languages, so a repository in something else will index less well than the documentation implies. And the container image embeds analytics keys at build time, with a build flag that decides whether a missing key fails the build, so read the environment template before you build it yourself.

## FAQ

### What is the shotgun CLI?

It is a MIT-licensed spec-driven tool for coding agents such as Codex, Cursor, Claude Code and Antigravity. It reads your codebase, plans a full feature up front, and splits it into staged pull requests with file-by-file instructions an agent can follow, with the stated aim of producing several focused pull requests instead of one enormous one. It works with your own model keys or with prepaid usage credits.

### What is the difference between Planning and Drafting mode in shotgun?

Planning is the default: the tool proposes an execution plan, shows each step and asks for confirmation before running agents that change files, and lets you confirm or skip cascaded updates when one change affects other documents. Drafting runs the whole plan in one pass with no intermediate confirmations, though progress is still tracked internally. A keypress switches between them.

### How does shotgun research a codebase?

The first step builds a searchable graph of the entire repository. The implementation uses a parser library with grammar packages installed for five languages: Python, JavaScript, TypeScript, Go and Rust, plus a file-watching dependency so the index can follow changes. Being a syntax-tree index rather than an embedding index is what lets it find existing definitions and references.

### What Python versions does shotgun support?

Python 3.11 to 3.13. The packaging metadata carries an explicit upper bound of 3.14, with a comment attributing it to one pinned dependency lacking wheels for that release, and the support table in the documentation marks 3.14 and later as unsupported because no wheels exist yet. Thirty-two-bit Python is also listed as unsupported.

### How does shotgun install and run?

Install an execution tool, then run the project through it at a pinned channel with a single command. The reason for that route rather than pip is faster installation and reliable handling of binary wheels, avoiding compiler errors. On Windows there is a one-time execution policy change, a path update and an optional administrator step for a runtime redistributable, and you must run it in PowerShell rather than the older command prompt.

## Sources

- [License: MIT](https://github.com/shotgun-sh/shotgun/blob/main/LICENSE)
- [Project website](https://shotgun.sh/)
- [README](https://github.com/shotgun-sh/shotgun/blob/main/README.md)
- [Releases](https://github.com/shotgun-sh/shotgun/releases)
- [shotgun-sh/shotgun on GitHub](https://github.com/shotgun-sh/shotgun)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/shotgun-sh-shotgun
