bkit-claude-code targets 100% match and gates at 90%
bkit Vibecoding Kit - PDCA methodology + Claude Code mastery for AI-native development
At a glance
- What is it?
- A Claude Code plugin of 44 skills, 34 specialist agents and 11 quality gates that splits work into context-budgeted sprints and self-repairs design drift. The interesting numbers are the ones that disagree with each other: a stated 100% match target behind a 90% gate, a Claude Code floor of v2.1.143 against bkit's own v2.1.40, and eight lifecycle phases shown in a six-row table.
- Who is it for?
- Use bkit-claude-code when you are already inside Claude Code and your recurring failure is that generated code quietly stops matching the design you approved, because measuring that gap and repairing it automatically is the specific thing it does. Do not adopt it expecting a spec-to-code guarantee, since the loop's stated target is 100% match while the gate that stops it is 90%, and the sprint table merges two of the eight lifecycle phases into one row.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 9 days ago.
- What is it written in?
- Mainly JavaScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 6, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The stated target is 100% match and the gate that stops the loop is 90%
This is the pair of numbers worth holding on to. The Iterate row of the PDCA table says the target is 100% match between design and code, that a `gap-detector` measures the match rate, and that a `pdca-iterator` self-repairs until quality gate M1 passes, with the gate defined inline as matchRate of 90% or more and a maximum of five cycles. The same 90% threshold appears in the before-and-after table, where a match rate below 90% triggers auto-repair for up to five cycles. So the loop is written to stop ten points short of its own stated target, and after five repair attempts it reports and moves on rather than escalating. That is a defensible design for a loop that has to terminate, but it means a passing run is evidence of 90% alignment at best, not of a design that was fully realised.
Claude Code v2.1.143 or newer, against bkit's own v2.1.40
The install requirement is unusually specific and worth copying exactly if you hit it. bkit requires Claude Code v2.1.143 or later, because the strict plugin-manifest path recognises the official `displayName` field only from that version. On anything older, `claude plugin install` fails with `Validation errors: Unrecognized key: "displayName"`. The stated fix is to run `npm install -g @anthropic-ai/claude-code@latest`, with a compatibility guide in `docs/06-guide/cc-compatibility.guide.md`. The awkward part is the numbering. bkit itself is at v2.1.40 while the host has to be at v2.1.143, two projects whose version ranges overlap almost entirely, so reading a version number off either project without knowing which one you are looking at sends you the wrong way. A single manifest field is what the entire floor rests on, and there is no stated fallback path for clients that never gain support for it.
One human decision across an eight-phase lifecycle
The autonomy boundary is drawn in one row of the table. In the Design phase a `cto-lead` agent proposes three architecture options and you pick one, which the table marks as the only required user input in the whole lifecycle. Everything else runs without you: a `pm-lead` orchestrates four product-management agents, a `/pdca team` call spawns four to six implementation specialists, a `qa-lead` runs four QA agents through five test levels plus a data-flow integrity check, and a `report-generator` writes the completion report. Each phase is gated, and the README says so plainly: if the match rate is too low, a critical issue turns up or the data flow is broken, the run pauses and tells you why, so you never have to remember to verify. The one dial that governs how far any of this goes is `/control level N`, described as deciding how far the orchestrator runs before stopping and as controlling the same thing for sprints and agents alike. That paragraph is cut off mid-sentence in the visible documentation.
Sprints are sized to 75K tokens by topological sort and greedy bin-packing
The unit of work is a sprint, and the sizing has an actual algorithm behind it. You type `/sprint master-plan my-release --features auth, billing, reports`, and a `sprint-master-planner` agent investigates in depth rather than quickly, reading the existing codebase or researching the web, before splitting your features into context-budgeted sprints. The default ceiling is 75K tokens per sprint, and the split is dependency-aware through Kahn topological sort combined with greedy bin-packing, which is a specific and sensible choice: it respects feature dependencies and then packs what is left into the fewest sessions. Since it is a default, the ceiling is adjustable, which matters more than the number itself. The phase table shows six rows for what the surrounding text calls an eight-phase lifecycle, and the difference is explainable, with prd and plan sharing one row and archived having no row at all. The workflow diagram above it stops partway through Step 3.
Picks up where you stopped, including after the laptop crashes
The durability claim is the strongest sentence in the documentation and the least qualified. On approving the master plan, every sprint is registered in the task management system in dependency order, Kahn-topologically sorted, and memory is written; the result is described as bkit picking up exactly where you stopped if your laptop crashes, if your session clears, or if you start a new Claude Code session next week. Two of those three are about sessions and follow from files surviving. The first is about hardware, and it holds only to the extent that wherever the `memory/` directory lives also survives. There is a `memory/` directory at the repository root of the tree, and whether it is inside the project you commit is the decision that determines it. The audit trail is the other half: the Docs = Code philosophy has every feature produce a PRD, a plan, a design, an analysis and a completion report, and the README calls that history the audit trail.
44 skills, 34 agents, 11 gates, and two test directories
The headline counts are 44 skills, 34 specialist agents and 11 quality gates, with the gates named as match rate, critical issues, convention, test coverage, security and data-flow integrity among others. The repository tree is the honest version of that number, and it is wide: `skills/`, `agents/`, `hooks/`, `commands/`, `lib/`, `memory/`, `templates/`, `servers/`, `bkit-system/`, `evals/`, `refs/`, `skill-creator/` and `output-styles/` all sit at the top level, alongside both a `test/` and a `tests/` directory. Two test directories is the kind of small redundancy that tells you which convention won, and the README does not say. There are also generated artefacts committed at the root, an HTML and a PDF of the infrastructure architecture, plus a `README-FULL.md` beside the README you are reading and a `bkit.config.json`.
Eight languages, 43 frameworks, five test levels and an uneven release cadence
The internationalisation claim is concrete: eight-language auto-detection, with the examples given in Korean and Chinese, and an intent-router plus auto-trigger that picks the skill or agent so you do not have to know the command names. The agent fan-out is also specified rather than implied. `/pdca pm` runs four product-management agents, discovery, strategy, research and prd, against 43 frameworks. `/pdca team` spawns four to six specialists from a list of six roles, developer, QA, frontend, backend, security and architect. `/pdca qa` runs a five-agent team, and the QA phase itself works through five test levels with a seven-layer hop check on data-flow integrity. The release history is the least tidy part. Three recent tags carry titles that read like principles rather than changelogs, and the gap between two of them is thirty-four days: v2.1.38 on 2026-08-17, v2.1.39 on 2026-09-20, v2.1.40 on 2026-09-25.
Editorial conclusion
Use bkit-claude-code when you are already inside Claude Code and your recurring failure is that generated code quietly stops matching the design you approved, because measuring that gap and repairing it automatically is the specific thing it does. Do not adopt it expecting a spec-to-code guarantee, since the loop's stated target is 100% match while the gate that stops it is 90%, and the sprint table merges two of the eight lifecycle phases into one row. Two operational checks before you install: confirm your Claude Code is v2.1.143 or newer, because the plugin manifest's `displayName` field is only recognised from that version and older clients reject the install outright, and confirm where the `memory/` directory lives, since the recovery promise depends on that file surviving whatever your session does not.
Frequently asked questions
What does bkit-claude-code require to install?
Claude Code v2.1.143 or later, because the strict plugin-manifest path recognises the official `displayName` field only from that version. On older clients `claude plugin install` fails with `Validation errors: Unrecognized key: "displayName"`, and the stated fix is `npm install -g @anthropic-ai/claude-code@latest`.
How does bkit-claude-code measure whether generated code matches the design?
A `gap-detector` agent computes a match rate between the design spec and the generated code, and a `pdca-iterator` self-repairs until quality gate M1 passes. That gate is a match rate of 90% or more, with a maximum of five repair cycles, even though the stated target is 100% match.
How large are bkit-claude-code sprints?
The default ceiling is 75K tokens per sprint, so a single Claude Code session can finish one without overflowing the context window. The split is dependency-aware, using Kahn topological sort plus greedy bin-packing, and it happens when you run `/sprint master-plan` with a `--features` list.
How much human input does the bkit-claude-code lifecycle need?
One required decision. A `cto-lead` agent proposes three architecture options in the Design phase and you pick one, which the documentation marks as the only required user input. Every other phase runs on gates, with the run pausing and explaining itself if match rate, critical issues or data-flow integrity fail.
Does bkit-claude-code work outside English?
It claims eight-language auto-detection, with Korean and Chinese given as the worked examples, and an auto-trigger plus intent-router that selects the right skill or agent so you do not need to know the command names. Memory and task state are what let a later session resume the plan.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/ww-w-ai-bkit-claude-code)