agent-spec freezes 23 requirement subcommands and five verdicts, and its overview stops mid-word twice
`agent-spec` is an AI-native BDD/spec verification tool for task execution.
At a glance
- What is it?
- A Rust intent compiler that lowers human requirements into Task Contracts and checks code against them, with every gate between the drafting model and the verification machine declared deterministic. The interesting parts are the lint failures it names, the baseline it retired, and the trace record that deliberately stores no identity.
- Who is it for?
- agent-spec fits a team that already writes requirements down and wants a machine to notice when the code drifts from them, and it fits badly as a place to start learning spec-driven development, since the contract surface is a promise rather than a tutorial and the architecture details sit in a separate document. Before adopting it, check three things for yourself.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 35 days ago.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Two parts of the project's own overview stop mid-sentence
The architecture section draws the pipeline as a flowchart of named boxes, and the drawing breaks off inside a node label at `lint-knowledge --gate`, so the diagram never shows the layers that follow the governance gate. The pipeline claim therefore rests on the prose after it rather than on the picture. The same pattern shows up at the end of the overview, where a section titled What agent-spec doesn't solve opens a sentence about telling readers something up front and breaks off after the letters up-fro. The list of non-goals never arrives. The nearest statements of limits are scattered elsewhere: the detailed contracts for Requirement Governance, Code Graph IR, the Intent-Code Linker, Quality Planning, and Execution Bundles are deferred to `docs/intent-compiler/architecture.md`, the CLI never authors a question candidate, and derived code facts never become governed truth.
AI shows up at two ends, and every gate between them is model-free
The central claim is narrow enough to check: AI participates only at the edges. On the way in, an `agent-spec-intent-compiler` skill has a model draft Candidate Requirement Blocks while a human reviews and accepts them, and a second intake path, `agent-spec requirements import`, is deterministic, taking marked blocks or the YAML dialect. A Requirement Governance Gate then moves a block from proposed to accepted or rejected. On the way out, an agent implements against the contract and the machine checks whether the code satisfies it. The gates in between are named and declared deterministic and model-free: `lint-knowledge`, `graph`, `plan`, `lifecycle`, `trace`. Human acceptance sits at both ends, requirement review on the way in and Contract Acceptance on the way out, which is the whole review loop the project describes: humans review the contract, agents implement, the machine verifies.
The governance walk has one exit per layer and four id prefixes
Governance is described as flowing in one direction, with each layer having exactly one exit, and the walk is short enough to print. Proposals ask whether to do something and why, decisions record the ruling together with its alternatives, requirements carry MUST clauses plus scenarios, and specs are the executable, verifiable contract. The id prefixes are fixed by directory: `LEP-NNN` in `knowledge/proposals/`, `ADR-NNN` in `knowledge/decisions/`, `REQ-*` in `knowledge/requirements/`, and `task-*` in `specs/`. Two fields carry the handoffs, the `## Produces` line under a proposal and the `satisfies:` list under a contract. `knowledge/standards/operational/id-registry.md` is the stated authority for which prefix lives where, and `agent-spec knowledge new <kind> <id>` exists so the enum values, filename, and exit pointer are right the first time.
knowledge/proposals/ LEP-NNN should we do this, and why
│ ## Produces: ADR-NNN
knowledge/decisions/ ADR-NNN the ruling, with alternatives
│ a governed requirement
knowledge/requirements/ REQ-* MUST clauses + scenarios
│ satisfies: [REQ-*]
specs/ task-* executable, verifiable contractThree named lint failures mark where the chain breaks
`lint-knowledge --gate` is the enforcement point and it names three failures. A task contract with no `satisfies:` line raises `orphan-spec` at Warning severity, so a contract wired to nothing is a warning rather than a hard stop, which is worth knowing before you rely on a pipeline to catch it. An `ADR-*` id appearing under a requirement's `## Dependencies` raises `dependency-kind-mismatch`, catching a decision id used where a requirement belongs. An accepted proposal whose produced decision does not link back raises `produces-link-integrity`, closing the walk in the other direction. The 1.3.0 migration left its own mark: it used a shrink-only baseline for contracts that already existed, and that baseline is now retired and empty, so any non-empty `.agent-spec/orphan-baseline.json` is an Error rather than an exemption. The migration escape hatch became a hard failure.
Humans stop the machine at three points, and the trace keeps no identity
Three stages stop for a human, and all three emit the same machine-readable envelope so an agent harness can render the pause as a choice instead of improvising the question. Candidates are drafted by the agent from source text and validated by the CLI, at most four, each with a label and a one-sentence description, and the CLI never authors one itself. An empty candidate list is valid and means the question stays free-form rather than silently taking the nearest stored option.
| Stage | Command | What it asks | |---|---|---| | Reverse interview | `requirements questions` | Ambiguity a requirement lint found | | Governance | `knowledge questions <id>` | A proposal's unresolved questions, a decision's alternatives | | Acceptance | `verify --emit-questions` | Scenarios the machine could not settle, with their evidence |
Answers to acceptance questions convert directly into the decisions JSON that `resolve-ai` already consumes, so a human's answer lands on a path the tooling already reads. When a human settles a scenario, the trace record carries the judgment's class and an evidence digest, never an identity, because approval binding stays with the external system that can actually attest it. Liveness tracing keeps compiled knowledge honest after the fact, and the 1.0 promise commits to derived liveness being never stored, so the record is recomputed rather than frozen.
Two IR tracks, one of them derived and rebuildable
The design keeps two intermediate representations apart on purpose. Requirement IR records what the system must do, and accepted KLL requirements remain the governed source of truth. Code Graph IR records what the current program is, and that graph is derived and rebuildable. A provider-neutral Intent-Code Linker then grounds accepted work units in code without turning derived code facts into KLL truth, which is the rule that stops a re-derived observation from being promoted into a requirement. Quality Planning resolves deterministic tools and required agent skills into an Execution Bundle, with skills guiding generation and tool results providing acceptance evidence. The primary planning surface for the whole loop is the Task Contract, rendered by `agent-spec contract`. The five verdicts a verification run can return, `pass`, `fail`, `skip`, `uncertain`, and `pending_review`, sit alongside an `is_passing` flag, and both are part of the promised machine formats.
The 1.0 promise names 19 commands, two families, and 23 requirement subcommands
Version 1.0 turns a long list of surfaces into a compatibility promise, with breaking changes only at a major version. The CLI list is 19 single commands running from `init` through `plan`, plus `wiki` and `atlas` families, plus a `requirements` family with 23 subcommands. The overlap is part of the shape: `plan` exists both as a top-level command and as `requirements plan`, and tracing appears three ways as `trace`, `requirements trace-graph`, and `requirements traceability`. Machine formats are promised as firmly as the flags, covering lifecycle and verify JSON top-level keys, schema `$id` URIs under `agent-spec/intent-compiler/`, YAML dialect v1.1, compilation provenance manifests in v1 and replayable v2 forms, the requirement-traceability projection, compile bundle layouts `agent-spec-v1` and `arc-v1`, and the atlas graph `schema_version`. Deprecations get one minor release with a notice and then leave at the next major, and `brief`, the legacy alias of `contract`, was deprecated in 0.4.0 and is now removed, with `contract` rendering the identical output.
Cargo.toml ships 1.4.0 with the MIR feature off and three fixtures excluded
The crate manifest names version 1.4.0 on edition 2024 under an MIT license, with a workspace of three members: the root crate, `crates/rust-atlas`, and `crates/code-graph-provider`. The notable default is `default = []`, with `mir` mapped to `rust-atlas/mir`, so the mid-level IR work is opt-in rather than switched on. Three fixture directories, `fixtures/requirements-noteapp`, `fixtures/atlas/basic`, and `fixtures/atlas/concurrent-query`, are excluded from the workspace, and `exclude = [".superpowers/**"]` names a directory that does not appear at the top level of the repository. The lint table denies `unsafe_code`, `unwrap_used`, and `expect_used`, which fits a tool whose job is to fail loudly rather than guess. `blake3` sits in the dependency list, matching the evidence digests the trace records carry. The manifest's `homepage` points at the repository, while the project site address appears only in the repository metadata.
Editorial conclusion
agent-spec fits a team that already writes requirements down and wants a machine to notice when the code drifts from them, and it fits badly as a place to start learning spec-driven development, since the contract surface is a promise rather than a tutorial and the architecture details sit in a separate document. Before adopting it, check three things for yourself. The repository's own list of non-goals is not in the text, so decide what you expect it not to do. The `orphan-spec` lint fires at Warning severity, so an unwired contract will not stop a pipeline on its own. And `blake3` digests plus a class-only trace record mean the tool proves consistency, not identity, which is a deliberate limit rather than a bug to route around.
Frequently asked questions
What is ZhangHanDong/agent-spec?
An intent compiler for AI agent coding. Human intent, captured as structured requirements, lowers into Task Contracts that agents implement against, and the machine then verifies whether the code satisfies them. BDD and spec verification are the backend of that compiler.
Does agent-spec run a language model while it verifies code?
Not in the gates. The project states that AI participates only at the edges, drafting candidate requirements and implementing contracts, while `lint-knowledge`, `graph`, `plan`, `lifecycle`, and `trace` stay deterministic and model-free.
Which verdicts can an agent-spec verification run return?
Five: `pass`, `fail`, `skip`, `uncertain`, and `pending_review`, together with `is_passing` semantics. Those names, the lifecycle and verify JSON top-level keys, and the YAML dialect version 1.1 are all part of the 1.0 compatibility promise.
How do I add a new requirement or contract in agent-spec?
Scaffold it with `agent-spec knowledge new <kind> <id>`, which sets the enum values, filename, and exit pointer. `knowledge/standards/operational/id-registry.md` is the authority on prefixes: LEP-NNN for proposals, ADR-NNN for decisions, REQ-* for requirements, and task-* for specs.
What happened to the brief command in agent-spec?
It is gone. `brief` was the legacy alias of `contract`, deprecated in 0.4.0 and now removed, with `contract` rendering the identical output. The stated policy keeps a deprecated surface working for one minor release with a notice, then drops it at the next major.
Does agent-spec have a section on what it does not solve?
It opens one, and the first sentence breaks off after the letters up-fro, so the list of non-goals is not in the text. The closest statements of limits are elsewhere: detailed contracts live in `docs/intent-compiler/architecture.md`, the CLI never authors a question candidate, and derived code facts never become governed truth.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/zhanghandong-agent-spec)