claude-code-harness: a plan, work, review loop with a runtime floor
Claude Code Dedicated Development Harness - Achieving High-Quality Development Through an Autonomous Plan Work Review Cycle.
At a glance
- What is it?
- Claude Code Harness turns agent coding into an approved-spec pipeline with a Go adjudication engine in front of every tool call. It is for teams that keep losing the plan, the tests and the review trail.
- Who is it for?
- Adopt it if you already run Claude Code on a repo with a real test command and you want the plan, the acceptance criteria and the review verdict to survive the session. Do not adopt it if you want a faster single-shot edit loop, if your project has no tests to gate on, or if you need the safety layer to be configurable down to nothing, because the five runtime floor categories have no disable switch.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 9 days ago.
- What is it written in?
- Mainly Shell, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The drift problem Harness was built around
The README opens with a list of failures that will be familiar to anyone who has handed a multi-file task to a coding agent: plans live in chat and disappear, tests become optional under deadline, review happens after the merge, and release evidence gets reconstructed from memory. The project's stated fix is procedural rather than model-level. It replaces "ask the agent to code" with one path: write the spec, implement only the approved slice, verify, review independently, package evidence. The README is explicit that it does not make the model smarter and instead fixes the boundary around the model, so the procedure keeps working when the model changes.
That framing tells you who it is for. A solo developer doing exploratory edits in a scratch repo gets little from a contract approval step. A team that has to answer "what was approved, what was verified, and who signed off" is the target. The repository ships a Plans.md and a spec.md at the top level, which is the same shape the loop produces per task, so the project runs its own method on itself.
One claim in the README is worth isolating because it is unusual: the project says its claims are machine-checked, that CI gates verify described components are wired, that the task ledger stays consistent, and that shipped binaries rebuild from source. It states a feature appears in the README only after a gate proves it is reachable, and closes with "Written is not working." Treat that as a design commitment rather than a guarantee about any specific line you are reading.
Five verbs, and the gate attached to each
The loop is five skills: plan, work, review, sync, release. Each stage leaves the material the next stage needs, and each has its own gate. The README presents this as a table, and the gates are the substance.
/harness-plan turns intent into spec.md and Plans.md covering scope, acceptance criteria, dependencies, unknowns and stop conditions. The gate is you: the README says your job is not to write the plan but to approve or correct the generated contract before execution continues.
/harness-work implements one approved task and adds tests when the task requires them, with TDD required when the task says so. /harness-work all runs the whole approved plan, with the same TDD gate applied task by task; the README recommends it once the plan is clear and the repo baseline is known.
/harness-review reviews the result separately from implementation, and major findings block completion. The README's phrasing here is the sharpest line in the document: PR-ready is not release-ready.
/harness-sync compares the plan against what is actually implemented and reports drift. No gate is listed for it. /harness-release packages only verified evidence into CHANGELOG, tag and release, and the release preflight must pass.
There is one more rule that shapes everything downstream: data the agent has not seen stays unknown instead of being quietly invented. That is a constraint on the artifacts, not a promise about the model's reasoning.
Installing the plugin and running your first approved plan
The README's own install path is four lines inside Claude Code. The first opens the CLI, the next two add the marketplace and install the plugin, and the last runs the one-time setup.
claude
/plugin marketplace add Chachamaru127/claude-code-harness
/plugin install claude-code-harness@claude-code-harness-marketplace
/harness-setupAfter setup, hand it something small. The README's example is a README onboarding flow, which is a good first task because the acceptance criteria are easy to judge.
/harness-plan Improve the README onboarding flowThe documented result is that Harness drafts spec.md and Plans.md for you, and you approve or correct them before execution continues. Read both files before you type anything else. If the acceptance criteria are vague, that vagueness propagates into /harness-work and /harness-review, because those stages gate against the contract the plan stage produced.
For the other tools, the README documents separate routes and warns that four install routes are not four identical guarantees. Codex CLI uses scripts/setup-codex.sh --user, and the README says to rerun it after Harness updates and then restart Codex. Cursor uses scripts/setup-cursor.sh. The README labels Claude Code, Codex CLI and Cursor as supported tier. The repository also carries .grok-plugin/, .grok/, opencode/, skills-codex/ and hosts/ directories, so the host surface is broader than the four-tool table, but the README's install section is where the supported routes are named.
The Go engine that adjudicates before the call runs
This is the part that distinguishes Harness from a prompt template, and the README says so directly. Every tool call is adjudicated by a Go engine before it runs, not reviewed after the fact. The stated reason is concrete: a file diff cannot see a network send or a deletion.
The design uses two layers with deliberately different strength. The runtime floor has five categories and denies outright. It is not overridable by any config, env var or permission mode. Those categories are billing, network egress, secret reads, production deploys, and destruction outside the task worktree. The README says the floor sits on an isolated code path with no disable switch, so an autonomous run cannot talk itself past it. That is a strong architectural claim, and it is the one to verify yourself if you plan to run unattended agents near production credentials.
The second layer, guardrails R01 through R15, returns deny, confirm or warn, and is partly tunable by project config. The README names direct pushes to main, writes to protected paths, forced pushes and history rewrites as examples, each with a defined verdict. The repository carries claude-code-harness.config.example.json, claude-code-harness.config.schema.json, harness.toml and hosts.toml, which is where that per-project tuning would live.
The interaction design is the interesting trade-off. Confirmations move to plan time. Rather than interrupting a run, Harness collects the risky operations a plan will need and asks once, up front, and approvals carry an expiry, a task scope and a use limit. That prevents one approval from becoming a permanent hole, at the cost of a planning step that has to anticipate what the run will need. If your plan is wrong about which operations are risky, you will find out mid-run rather than at the prompt.
Every stop is recorded. Rule id, category and verdict land in a JSONL log. Command text is never written, only a hash and a length, and for secret-read and billing stops not even that. The practical effect is that you can count what blocked you, but you cannot reconstruct the exact command from the log alone.
Cross-session messaging and its optional verification gate
Open three agents on one repo and they normally work blind to each other. The README cites CooperBench numbers to argue this matters: two agents editing the same file succeed about half as often as one agent working alone, and 63% of the failures trace to a false belief about what the other one changed. Harness keeps a roster and a message path between local sessions.
The roster command is bin/harness session list, and the README notes it shows every live session on the machine including sessions in other worktrees, because the store resolves from git --git-common-dir. Each row carries the team and agent a sender needs. Sending uses bin/harness inbox send with --team, --from, --to and --subject flags, or the session-send skill. Receiving happens at the receiving session's turn boundary, wrapped as data with an explicit non-instruction envelope: a peer's message is a report to verify, never an order to follow.
Sending is unfiltered by default. Setting [livemsg] verification = "on" adds a gate that checks a message's factual claims, including whether mentioned files exist, whether mentioned commits resolve, and whether a clean worktree claim matches git status. When a claim fails, the reason returns to the sender instead of a false message being delivered. While the setting is off, the README states the send path does not call the gate at all, so the cost is zero and so is the protection.
This subsystem is local-only and does not depend on harness-mem. If harness-mem is installed alongside, the README says its roster entries are preserved untouched. The README's own summary of the trade-off is worth keeping: the pipeline is the product, and what you route through it is your call.
Where Harness is the wrong tool
The approval gate is the main cost. Every task begins with a contract you have to read and correct. For a one-line fix, a throwaway script, or a spike you intend to delete, that is pure overhead, and the README's own advice to start with something small suggests the authors know the first run feels heavier than a bare prompt.
The TDD gate is conditional on the task declaring it, so a repo with no test command cannot get much value from /harness-work. The gate has nothing to run. Similarly, /harness-release packages only verified evidence, which means a project that has never had a release process will spend its first cycle building one.
The runtime floor is not configurable, by design. If your workflow legitimately needs an agent to read a secret, deploy to production, or delete files outside the task worktree, Harness will deny it and no config, env var or permission mode will change that answer. You would have to do that step yourself outside the harness. That is the intended boundary, but it is a boundary, and it will stop work you consider routine.
The four tool routes are also not equal. The README states plainly that a setup script means a tool has an entry path, not a shared product promise. Codex CLI users are told to rerun scripts/setup-codex.sh --user after Harness updates and restart Codex, which is a manual upgrade step that plugin users do not have. If you standardize on a non-Claude host, budget for that difference.
Finally, the README does not document rollback. There is no described procedure for undoing a Harness-initiated change or reverting the harness setup itself, so plan to rely on your own version control for that.
How this compares with opencode and Cursor routes
The closest comparison the README itself invites is between Harness and the host tools it wraps. Cursor and opencode are editors and agent runtimes; Harness is a procedure and an adjudication layer that sits on top of a host. Installing Harness does not replace Cursor or opencode, it changes what happens between the plan and the merge. The repository even carries an opencode/ directory, so opencode appears as a host rather than a rival.
The real difference in approach shows up in where control lives. In a plain agent session, the model decides what to run and you review the diff afterward. Harness moves the decision in front of execution, into a Go engine that returns deny, confirm or warn before the call happens, and moves the risky confirmations back to plan time as scoped, expiring approvals. A diff-based review cannot see a network send or a deletion, which is the gap the README uses to justify the engine.
The comparison against Codex CLI is narrower than it looks. Codex CLI is a supported host with its own setup script, so the two are not alternatives in the usual sense. The meaningful question is which host you already use and whether you accept that the Codex route requires a rerun of the setup script after each Harness update.
If you want a lighter option, a plain prompt template plus your own CI is the honest alternative. It will be faster to start and cheaper to maintain, and it will not give you the runtime floor, the JSONL stop log, the drift report from /harness-sync, or the three HTML decision surfaces (Plan Brief, Progress and Acceptance) that the README describes for non-engineer sponsors.
Licence, upgrade cost and what the repository implies
The project is MIT licensed, with LICENSE.md and LICENSE.ja.md at the top level. MIT is permissive: you can use, modify and redistribute it, including commercially, provided the copyright notice and permission notice are retained. That is the standard reading of the licence text, not legal advice, and if you are embedding Harness in a product you should have your own counsel review the actual file rather than this summary.
The upgrade picture differs by host. Claude Code users install through the plugin marketplace, so updates arrive through that channel. Codex CLI users are told in the README to rerun scripts/setup-codex.sh --user after Harness updates and then restart Codex. Cursor users have scripts/setup-cursor.sh. The release cadence visible in the repository is fast: v5.12.0, v5.13.0 and v5.13.1 all landed between 2026-08-24 and 2026-08-25, with the last push on 2026-08-25. A version number in the 5.x line moving three times in two days means you should expect to re-run setup steps and re-read the changelog rather than treat an install as a one-time event.
The repository layout suggests a project that maintains its own claims: a scorecard.yml, a benchmarks/ directory, a tests/ directory, .githooks/, and a go.work file alongside go/. The README's statement that CI gates verify described components are wired and that shipped binaries rebuild from source is consistent with that layout. If you are evaluating it for a regulated environment, the artifacts to inspect first are SECURITY.md, docs/CLAUDE_CODE_COMPATIBILITY.md, and the guardrail rule set behind R01 through R15.
Editorial conclusion
Adopt it if you already run Claude Code on a repo with a real test command and you want the plan, the acceptance criteria and the review verdict to survive the session. Do not adopt it if you want a faster single-shot edit loop, if your project has no tests to gate on, or if you need the safety layer to be configurable down to nothing, because the five runtime floor categories have no disable switch. Before installing, read docs/CLAUDE_CODE_COMPATIBILITY.md and confirm that /harness-setup completes against your Claude Code version; the README states that a setup script is an entry path, not a shared product promise across the four supported tools.
Frequently asked questions
What is claude-code-harness?
It is an MIT-licensed development harness for Claude Code, Codex CLI, Cursor and Grok that replaces ad hoc agent prompting with one path: write the spec, implement only the approved slice, verify, review independently, and package evidence. It also adds a Go engine that adjudicates every tool call before it runs.
How do I use claude-code-harness?
Install the plugin through the Claude Code marketplace, run /harness-setup once, then start with /harness-plan on a small task. Harness drafts spec.md and Plans.md, and your job is to approve or correct that contract before execution continues.
How does claude-code-harness work?
The loop is five skills: plan, work, review, sync and release, each with its own gate. Separately, every tool call is adjudicated by a Go engine before it runs, split into a five-category runtime floor that cannot be overridden and guardrails R01 through R15 that are partly configurable per project.
Is claude-code-harness open source?
Yes. The repository is public, the default branch is main, and it ships under the MIT licence with LICENSE.md and LICENSE.ja.md at the top level.
Is claude-code-harness free?
The project is MIT licensed, which permits use, modification and redistribution including commercial use, provided the copyright and permission notices are retained. That covers the harness itself; it does not cover whatever you pay for the underlying model or host tool.
How does claude-code-harness compare with Cursor?
They are not the same kind of thing. Cursor is a host the README lists as a supported tier with its own scripts/setup-cursor.sh install route, while Harness is the procedure and adjudication layer that sits on top of a host and gates the plan, the work and the review.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/chachamaru127-claude-code-harness)