# cavekit: one spec file, seven commands, and a project that froze itself

> A Claude Code plugin that distills spec-driven development down to a single SPEC.md and a loop of spec, build and check, then froze in favour of its successor in August 2026.

**JuliusBrussee/cavekit** — Frozen — compressed spec-driven development plugin for Claude Code. Still works; active development moved to JuliusBrussee/caveman.

- Repository: https://github.com/JuliusBrussee/cavekit
- Website: https://caveman.so/
- Stars: 1,149 · Forks: 90
- Language: Unknown
- License: MIT
- Published: 2026-10-07 · Updated: 2026-10-07 · Language: en
- Canonical page: https://hysenlabs.com/projects/juliusbrussee-cavekit

## The status line is the first thing in the README

The project opens with a status block rather than a logo. It says, in a quoted callout, that cavekit is frozen as of August 2026, that it is no longer in active development, and that everything below still installs and works but should be expected to bring no new features or fixes. Active development of the family has moved to `caveman` and `caveman-browse`, both linked.

This is worth pausing on because it is unusual. Most frozen projects freeze quietly and let a user discover the staleness when something breaks. This one leads with the news, names where to go instead, and separates what still works from what will not happen. The repository description says the same thing more tersely: frozen, compressed spec-driven development plugin for Claude Code, still works, active development moved to JuliusBrussee/caveman.

The last push to master was 2026-08-14, and there are three visible releases: v2.0.0 and v3.0.0 and v3.1.0, all on 2026-04-11 and 2026-04-18. So the whole active period compressed into about ten days in April 2026, followed by a freeze. Licence is MIT and the repository is at 1,149 stars.

For anyone evaluating this today, the practical read is that the design below is a historical artifact worth learning from, and the tool to install is whatever the successor has become.

## The argument is that most spec frameworks cost more than they save

The README states its thesis in a few sentences and they are blunt. Plan-then-execute forgets. Spec-driven development remembers, but most SDD frameworks bury that value under agent swarms, dashboards, and ceremony that costs more tokens than it saves.

So the pitch is deliberately the opposite of a framework pitch. cavekit describes itself as the simplest full loop: grill, then spec, then research, then review, then build, operating over one `SPEC.md` file, with no sub-agents. The README quantifies the gap. Three commands are run every time; four more are reached for only when the change earns it.

The spine is described as three properties. The durable spec is `SPEC.md` at the repository root, which survives context resets and acts as the agent's long-term memory, so you lose the window, reload the spec and keep going. The caveman encoding is a compressed notation using symbols, fragments and pipe tables, claimed at roughly 75% fewer tokens than prose. And the backprop reflex turns every test failure into an entry, so classes of bug become invariants the spec never forgets.

The token comparison is the argument's sharpest edge. The README claims all nine of its skill descriptions cost about 1.1k of context, which it describes as 16 times lighter than spec-kit's 18.6k. Whether or not that ratio survives your own setup, it names the right thing: in agentic workflows, your instructions are competing for the same context window as your actual task.

## Three commands for every change, four for the ones that earn it

The command table is short enough to hold in your head, which is the point.

The loop you run every time has three entries. `/ck:spec` creates, amends or backpropagates `SPEC.md`, and it is described as the sole mutator of the file. `/ck:build` plans and then executes against the spec, naming which test proves each invariant, and auto-backpropagates on failure. `/ck:check` is read-only and produces a drift report listing invariant, interface and task violations, which the README calls the drift detector.

Then there are four reach-for commands. `/ck:grill` interrogates a fuzzy idea into a sharp goal and constraints section, one question at a time, before you spec. `/ck:research` gathers external knowledge into a research section so the build is grounded in facts rather than guesses, with every finding citing a source. `/ck:review` is an adversarial senior review of the spec before build, which refutes and hardens the invariants and ends in a go or no-go gate. `/ck:deepen` is a spare-budget design pass that makes one shallow module deep while holding behavior and keeping tests green.

The rule that keeps this from becoming the thing it criticizes is called right-sizing. A one-line fix is just `/build`. The full chain is for genuinely uncertain or high-blast-radius work, and never for a typo. That single sentence is a better statement of workflow design than most frameworks manage in their entirety, because it ties ceremony to blast radius rather than to process preference.

There is also a stated non-goals list that functions as a design document: no sub-agents, no dashboards because `cat SPEC.md` is the dashboard, no parallel workers, no JSON or YAML spec bodies, no hooks, no orchestration binaries and no TypeScript helpers.

## A sectioned spec where each verb owns its own sections

The format is defined in `FORMAT.md` at the repository root, and it is the part you would want to steal regardless of which plugin you use.

`SPEC.md` is organized into seven section markers, each with a letter: goal, constraints, interfaces, research, invariants and tasks and bugs. Research, tasks and bugs are each described as pipe tables. The load-bearing rule is that each verb owns specific sections, and no verb rewrites a section it does not own.

That constraint is the mechanism that makes a multi-agent-friendly file safe even in a single-agent workflow. `/ck:spec` is the sole writer of the file, and `/ck:check` only reads. A read-only drift detector cannot corrupt the thing it is measuring, so you can run it as often as you like, which is the property you want from a verification step.

The backprop protocol is what closes the loop. Every test failure becomes a bug entry, and classes of bug become invariants the spec never forgets. So a bug found during implementation does not just get fixed, it changes the specification so the same class cannot recur silently. The `skills/backprop` module is documented as a six-step protocol.

The repository tree is consistent with the README's minimalism. `commands/` holds seven thin slash-command entry points that delegate to the skills, and `skills/` holds `spec`, `build`, `check`, `grill`, `research`, `review`, `deepen`, `caveman` and `backprop`. There is also `.claude-plugin/` and a `plugin.json` for marketplace packaging, plus `CHANGELOG.md`, `UPGRADE.md`, `SECURITY.md` and a `LAUNCH-POST.md`.

## Install paths, and the older generation still frozen at v3.1.0

Installation is one line through the `skills` CLI:

```bash
npx skills add JuliusBrussee/cavekit
```

That installs nine skills into `~/.claude/skills/`: spec, build and check as the loop, grill, research, review and deepen as the reach-for set, plus caveman and backprop as utilities. Claude activates each when its trigger context matches, so a request like writing a spec invokes spec, a fuzzy idea invokes grill, and a risky change before build invokes review. The README notes that Claude Code picks them up on the next launch, which is a real detail if you are scripting around it.

Two alternatives are given. A marketplace install adds the slash commands as well:

```bash
/plugin marketplace add juliusbrussee/cavekit
/plugin install ck@cavekit
```

Or clone directly into the plugins directory:

```bash
git clone https://github.com/juliusbrussee/cavekit.git ~/.claude/plugins/cavekit
```

Then there is the previous generation, which the README handles with unusual care. v3.1.0 and earlier is not deprecated, it is frozen at that tag and remains a fully working plugin. Its feature list is a direct contrast to the current one: a four-command Hunt lifecycle of sketch, map, make and check, plus ship, review, revise, status, design, research, init, config, resume and help, sixteen slash commands total, twelve named sub-agents, per-task token budgets, a stop-hook state machine, model-tier routing, auto-backpropagation from test failures, tool-result caching, Codex peer review, knowledge-graph integration, design-system enforcement, parallel wave execution and team mode.

The README then tells you how to choose. Pick v3.1.0 if you want the full autonomous loop, parallel agents and peer review. Pick v4 if you want the distilled loop, fewer moving parts and smaller token bills. Installing the older version is a marketplace add pinned to the tag, or a clone with the `-b v3.1.0` branch flag.

That contrast is the most useful thing in the repository. It is a controlled experiment on what happens when you remove orchestration from an agent workflow, run by the person who built the orchestrated version.

## Conclusion

cavekit is worth reading even though it is frozen, because it is a clear-eyed argument about what makes agentic development workflows expensive. The claim is that durable state in a single plain-Markdown file beats agent swarms and dashboards, and the evidence offered is a token count: nine skill descriptions at roughly 1.1k of context against 18.6k for a comparable framework. What it actually shipped is a disciplined design. One file, one writer per section, no sub-agents, no hooks, no orchestration binaries, and a right-sizing rule that keeps the heavy commands off the path of a one-line fix. If you want to use it, install the frozen version and expect no fixes, or move to the caveman repository the README points at, which is where development continued. UPGRADE.md frames the move as a two-way door, and it genuinely is, because `SPEC.md` is ordinary Markdown that outlives whichever plugin wrote it.

## FAQ

### Is there a caveman skill in Claude Code?

Yes. cavekit ships a `caveman` skill as one of its nine, described as an encoding utility: a compressed notation using symbols, fragments and pipe tables, claimed at roughly 75% fewer tokens than prose. Note that the project's active development has since moved to a separate repository also named caveman, so the current version of that skill lives elsewhere.

### What is cavekit and what does SPEC.md do?

cavekit is a Claude Code plugin for spec-driven development, built around a single `SPEC.md` file at the repository root. That file is the agent's durable memory: it survives a context reset, it is divided into sections for goal, constraints, interfaces, research, invariants, tasks and bugs, and each command owns the sections it is allowed to rewrite.

### Is cavekit still maintained?

No. The README leads with a status block saying the project is frozen as of August 2026, that everything still installs and works, and that no new features or fixes are expected. Active development moved to the author's caveman and caveman-browse repositories. The last push to master was 2026-08-14.

### What is the difference between cavekit v3.1.0 and v4?

v3.1.0 is the older, feature-heavy generation: sixteen slash commands, twelve named sub-agents, per-task token budgets, a stop-hook state machine, model-tier routing and parallel wave execution. v4 is the distilled version: one spec file, three loop commands, four reach-for commands, no sub-agents and no orchestration. The README says to stay on v3.1.0 for the autonomous loop and move to v4 for fewer moving parts and smaller token bills.

### How do I install cavekit?

Run `npx skills add JuliusBrussee/cavekit`, which installs nine skills into `~/.claude/skills/` and activates them by trigger context on the next Claude Code launch. You can also add it through the marketplace to get the `/ck:` slash commands, or clone the repository into `~/.claude/plugins/cavekit`.

## Sources

- [JuliusBrussee/cavekit on GitHub](https://github.com/JuliusBrussee/cavekit)
- [License: MIT](https://github.com/JuliusBrussee/cavekit/blob/main/LICENSE)
- [Project website](https://caveman.so/)
- [README](https://github.com/JuliusBrussee/cavekit/blob/main/README.md)
- [Releases](https://github.com/JuliusBrussee/cavekit/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/juliusbrussee-cavekit
