Spec-Driven Develop: an architecture-first Markdown workflow for AI coding agents
Spec-driven development workflow for AI coding agents: architecture-first planning, task decomposition, GitHub Issue/PR tracking, Deep Discuss, and adaptive control for Claude Code, Codex, Cursor, and other Markdown-capable agents.
At a glance
- What is it?
- Spec-Driven Develop is a skills-only plugin that turns large agent tasks into a seven-phase spec loop with GitHub Issue tracking and tiered review. It ships as Markdown plus small helper scripts, and it expects your agent to be able to read them.
- Who is it for?
- Adopt Spec-Driven Develop if you run Claude Code, Codex, OpenCode or Cursor on migrations, rewrites and multi-module refactors where you want the plan, the Issues and the review trail in the repository rather than in chat history. Skip it for one-file fixes and throwaway scripts: seven phases, S.U.P.E.R health evaluation and delivery batches cost more than the change is worth.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 66 days ago.
- What is it written in?
- Mainly Shell, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem: agents lose the plan halfway through a large change
Ask an AI coding agent to rewrite a service in another language and the failure is rarely a single bad function. It is drift. The agent forgets which modules it already touched, invents a new directory layout in hour two, and leaves no artifact that a second session or a second engineer can resume from. Spec-Driven Develop exists for that class of work: migrations, rewrites, refactors, architecture changes and implementation plans that span many files and more than one sitting. The README describes the target user as developers using AI coding agents for exactly those tasks. The unit of work is not a prompt. It is a spec with phases, tasks, acceptance criteria and a tracker.
The project is platform-agnostic by design. It is pure Markdown plus small helper scripts, with no SDK and no third-party runtime dependencies, so anything that reads custom skills can use it: the README lists Claude Code, Codex, OpenCode, Cursor, Windsurf, Cline, Aider, Continue, Roo Code and Augment. That breadth is the point. The workflow lives in files you can diff, not in a vendor's runtime.
How the seven-phase pipeline actually runs
The README lays out Phases 0 through 6 as a linear pipeline. Phase 0 captures intent in one or two sentences. Phase 1 does deep analysis: architecture, module inventory, risk assessment, with a S.U.P.E.R health evaluation attached. Phase 2 refines intent by asking targeted questions grounded in that analysis, confirming scope, priorities and constraints. Phase 3 decomposes the work into phases, tasks and parallel lanes, annotating each task with S.U.P.E.R design drivers, then plans delivery batches and creates GitHub trackers. Phase 4 generates MASTER.md as a GitHub index or a local tracker. Phase 5 presents the plan for confirmation, executes tasks in parallel or sequentially, integrates delivery-batch PRs and runs an adaptive control feedback loop. Phase 6 archives the artifacts for traceability.
The interesting design decision is in Phase 5, and v1.15 sharpened it. Execution is orchestrator-centric: the main agent owns progress and quality, and dispatching a sub-agent is described as an economic decision rather than a default. Orchestrator-direct execution is Tier 0 and the norm. A single task-executor coder is Tier 1, for large or context-heavy batches. Parallel lanes are Tier 2, and they are gated: disjoint file sets, at least L effort per lane, independent verifiability, and no more than four lanes. Review is tiered the same way. Machine validation (L1) and the orchestrator's own diff review (L2) are the default; an independent code-reviewer agent (L3) is reserved for Tier 2 lanes and high-risk changes. Reviewer agents check lane diffs against acceptance criteria and commit fixes directly to lane branches, and only APPROVED or FIXED lanes integrate.
That is a real constraint, not decoration. If your change touches one shared config file across every lane, the disjoint-file-set rule blocks parallelism and you fall back to sequential execution. The workflow would rather serialize than merge conflicting lane branches.
GitHub Issues, milestones and the one-batch-PR default
When the workflow detects a GitHub repository, it creates a GitHub Issue per task, organized with Milestones (one per phase), Labels for priority, size and lane, and optionally a GitHub Projects board. Before implementation it reviews the full phase Issue set and groups related work into delivery batches based on dependencies, shared files and contracts, validation, review scope and rollback boundaries. Issues stay atomic as acceptance and telemetry records; the delivery batch owns the integration branch, aggregate validation and the PR.
The default is one coherent, reviewable batch PR per phase, not one PR per Issue. If your team's review process assumes a tight one-Issue-one-PR mapping, this is the setting that will surprise you first. The upside is fewer half-finished PRs and a rollback boundary that matches the phase rather than a single task. The cost is a larger diff in a single review, which is precisely the thing review tooling handles worst.
Installing Spec-Driven Develop and running your first phase
The repository ships the skills under plugins/ and a sync script under scripts/. The v1.15 release notes describe scripts/install-agents.sh as syncing the bundled skills to ~/.agents/skills for agents that read that shared directory. Run it from the repository root:
./scripts/install-agents.shThe README does not spell out flags for this script, so treat a plain invocation as the documented path and check the script's own output for where it wrote. Agents that read ~/.agents/skills should then see the bundled skills. Claude Code, Codex and OpenCode also have plugin integration, which the README lists as optional distribution alongside the plain Markdown skills.
The plugin is skills-only, which is the part that trips people up. Workflows are invoked through skills, not slash commands. So the first real use is not a command; it is a request in the agent's own language, for example asking it to migrate a module to another runtime. Phase 0 captures that intent, Phase 1 analyses the project, Phase 2 asks you targeted questions, and Phase 3 produces the task breakdown. If the repository is GitHub-hosted, expect Issues and Milestones to appear at that point. Confirm the plan in Phase 5 before any code is written, because that is the gate the workflow defines.
Where Spec-Driven Develop is the wrong tool
The clearest limitation is the shape of the work it accepts. A seven-phase pipeline with deep analysis, S.U.P.E.R health evaluation, task decomposition, delivery batches and an archive step is overhead for a one-line fix, a dependency bump or a throwaway script. The README positions the project for large-scale complex tasks, and nothing in it suggests a lightweight mode for small ones. Using it on a typo fix means paying the full ceremony for no benefit.
The second limitation is the host agent. The workflows are Markdown that the agent must read and follow. An agent that does not read custom skills, or that reads from a directory other than the one scripts/install-agents.sh writes to, gets nothing from the install. There is no runtime that enforces the phases; compliance depends on the model following instructions.
The third is GitHub coupling. Issue, Milestone, Label and optional Projects creation is described for detected GitHub repositories. The README mentions MASTER.md as a GitHub index or a local tracker, so a non-GitHub setup is contemplated, but the tracker integration is where the workflow's value concentrates. On GitLab or a bare remote you are using a thinner version of it.
Finally, the tiering rules can refuse to parallelize. Disjoint file sets, at least L effort per lane, independent verifiability and a four-lane cap are hard conditions. Work that fails them runs sequentially, which is correct but slower than a naive fan-out.
Compared with plain spec-driven prompting and with Deep Discuss
The nearest alternative is not another plugin. It is what most people already do: write a spec by hand, paste it into the agent, and ask for a plan. That approach has no phase gate, no Issue tracker, no delivery batches and no archive, but it also has no install step and no assumption that your agent reads skills from a shared directory. If your change fits in one session and one reviewer, hand-written prompting is the cheaper path.
Inside this repository the comparison is between its own three skills. Spec-Driven Develop automates the full pipeline for large-scale complex tasks. Deep Discuss is a structured deep-discussion workflow for problem analysis, brainstorming and solution design through disciplined multi-phase thinking; it produces thinking, not a task tracker. Review SPD (review-spd) is a findings-first code review workflow for uncommitted changes, date-range commits and branch or PR diffs, focused on bugs and regressions. Picking the wrong one is a common first mistake: running Deep Discuss on a migration gets you analysis without Issues, and running the full pipeline to review a three-commit branch is overkill when review-spd is the narrower fit.
Maintenance, upgrade cost and the MIT licence
The last push to the default branch was on 2026-07-26, and v1.15.0 was released the same day. The repository is not archived. Recent releases are close together: v1.13.1 on 2026-07-01, v1.14.0 on 2026-07-14, v1.15.0 on 2026-07-26. That cadence matters for upgrade cost because the release notes describe structural changes, not just fixes. v1.14 introduced GitHub-native task tracking and batch PRs; v1.15 moved to orchestrator-centric execution, single-sourced prompts and a skills-only surface. A skills-only surface is a breaking change for anyone who built habits around slash commands, and single-sourced prompts mean the canonical text for a topic lives in exactly one reference, so local edits to a copy can be overwritten by the next sync.
The licence is MIT, which is permissive and imposes no copyleft obligation on your own code. That is a statement about the licence text, not legal advice; if your organisation has rules about vendoring third-party workflow files into a repository, run it past whoever owns that policy. The practical upgrade cost is low in one sense (Markdown files, no dependency tree to resolve) and non-trivial in another: your agent's behaviour changes when the prompts change, so a version bump is worth a re-read of the release notes before you sync.
Editorial conclusion
Adopt Spec-Driven Develop if you run Claude Code, Codex, OpenCode or Cursor on migrations, rewrites and multi-module refactors where you want the plan, the Issues and the review trail in the repository rather than in chat history. Skip it for one-file fixes and throwaway scripts: seven phases, S.U.P.E.R health evaluation and delivery batches cost more than the change is worth. Before you commit, verify three things in your own checkout: that your agent actually reads from ~/.agents/skills after scripts/install-agents.sh runs, that the GitHub Issue and Milestone creation in Phase 3 matches how your repository is configured, and that your team accepts one batch PR per phase instead of one PR per Issue, because that is the default and it changes your review queue.
Frequently asked questions
What is spec-driven development in Spec-Driven Develop?
The project describes it as an architecture-first workflow that turns large software changes into a spec-driven loop: project analysis, task decomposition, GitHub Issue and PR tracking, progress continuity, and adaptive control. It is implemented as Markdown skills rather than a runtime.
Does spec-driven development actually work, according to Spec-Driven Develop?
The repository does not publish benchmarks or outcome data, so there is no measured answer here. What it does document is a seven-phase pipeline with confirmation gates in Phase 2 and Phase 5, tiered review, and an archive step, and it states that only APPROVED or FIXED lanes integrate.
What does "spec-driven development" mean in this project's own terms?
Spec-Driven Develop defines it as an architecture-first workflow for AI coding agents, with spec-driven development listed as a key concept alongside task decomposition, S.U.P.E.R principles and adaptive control. The spec lives in Markdown artifacts and, on GitHub repositories, in Issues and Milestones.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/zhu1090093659-spec-driven-develop)