boss-skill: gates that are checkable rather than unavoidable, by the project's own admission
Boss Skill - BMAD 全自动研发流水线(多 Agent 编排)
At a glance
- What is it?
- A marketplace skill that turns one coding agent into a nine-role team with event-sourced state, quality gates and a zero-network claim enforced by a source-level test. Its documentation is unusually candid that a gate can still be ignored, and its npm manifest carries a different name than the project.
- Who is it for?
- This is worth reading for its evidence model rather than its role names, because the event log, the recorded gate verdicts and the deterministic evals are the parts that survive a change of agent. The honesty about enforcement limits is the strongest signal here, and a team that understands its gates are checkable rather than binding can decide what to compensate for.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gates are checkable, and the project admits they are not binding
One sentence in the design notes is worth more than the rest of the feature list, because it tells you what this tool is not. QA, deployment and final checks run as real commands whose verdicts are recorded as events, and the final gate command plus a doctor command fail when a completed stage carries a failed gate. Then the caveat, stated plainly: enforcement still relies on the orchestrating agent honouring the protocol, so the CLI makes the verdict checkable, not unavoidable. That distinction is the whole argument for event-sourced pipeline state, and most projects on this subject hide it. Here you get a verdict you can check in a script, in continuous integration, or in a code review, which catches the careless case and leaves the determined case to your process.
Zero network is enforced by a test that fails CI
The privacy claim is the strongest thing in the repository and it is backed by a mechanism rather than a promise. There is no telemetry, no phone-home and no remote proxy for language models, and the only network surface is an opt-in preview server bound to the loopback address. More unusually, a source-level test on the network boundary fails continuous integration if anyone reintroduces an outbound client, which means the guarantee is maintained by the build rather than by good intentions. The doctor command then checks the same boundary at runtime, alongside the resolved runtime, install locations and event-stream health. The project's own phrasing is that the claim is auditable rather than merely asserted, and having a command that proves it is the difference between a privacy policy and a property you can verify.
No shell anywhere: verification goes through structured argv
The second guarantee is about command injection, and the reasoning is concrete. Wave verification runs through a structured argument list read from a JSON file, never through a shell, and the stated consequence is that cloning a malicious repository cannot smuggle commands through the task file. That is a real threat model for a workflow tool, because the workflow definition is exactly the file a stranger controls when you clone something, and any step that builds a shell string out of task content is an execution primitive. Passing an argument array removes the string-building step entirely. The interesting part is that this only works because the pipeline state is data rather than prose: the events are appended to a log, the execution state is projected from them, and the gates read that state instead of parsing prose an agent wrote.
State is an append-only log projected into read-only state
The artifact layout is the data model. Everything a run produces lives under a per-feature directory: a design brief, a product requirements document, an architecture document, a task list and a QA report, with a hidden metadata directory holding an events file, an execution state file and a workflow plan. The events file is append-only, which is what makes the run replayable and the gates verifiable, and the execution state is a projection of that log rather than something maintained alongside it. The project's own selection advice follows from this design: if you do not need a traceable folder of artifacts for a feature, you probably do not need the full pipeline, and a single role against an existing project is the cheap first run. The read-only review role writes one review file and is promised to touch nothing else in the repository.
Marketplace only, so the npm manifest is not the product
The distribution statement is emphatic: this is a skill you install into a coding agent, not a tool that installs other skills, and distribution is marketplace-only with no npm package and no global binary. Three routes are given, per host. Claude Code adds a marketplace and installs the plugin. Codex adds the marketplace through its own command and installs from a plugin browser. Any other agent goes through a separate skills command line tool:
npx skills add echoVic/boss-skillIt discovers the skill, prompts for the target agent, the scope and the install method, and writes a lock file you are told you can commit. For hosts that copy files rather than install plugins, the hooks have to be wired by hand with the bundled command line entry point. Yet the repository root carries a package manifest, a lock file, a formatter config and a test config, named after a different scope than the project and marked private. That manifest exists to run the project's own tests, not to be installed.
Nine roles, four stages, and the mapping is not printed
The headline sells a team, listing product management, architecture, interface design, tech lead, scrum master, frontend, backend, quality assurance and operations. The command table then describes a four-stage pipeline. Nowhere on the visible page is the mapping from nine roles onto four stages, and the only role selection shown is a core set of four, product, architecture, development and quality, paired with a flag that stops the run after implementation and test evidence. So the granularity you can actually request today is coarse, and the fuller cast is a description of the intended shape rather than nine separately invocable steps. Individual stages do have their own commands, including a read-only review, a quality stage with gates, a ship stage for build and deployment checks, and an extend command for adding a custom agent, pack or gate.
The typecheck script runs the compiler twice with different rules
The project manifest is the clearest statement of engineering priorities in the repository. The typecheck is not one invocation but two: first against the project configuration, then a second run against a small set of test and configuration files with the flags spelled out on the command line, including strict checking, unchecked index access, skipping library checks, permission to import TypeScript extensions, and a flag requiring that all syntax be erasable so the files can be type-stripped. So the test harness is held to a stricter standard than the library itself, which is an unusual and defensible inversion. Around it sit separate test suites by directory, for faults, for the harness, for the install matrix and for skills, a coverage variant, a lint and format pair from a single formatter, and evaluation scripts that are shell rather than JavaScript, including a release mode for the same run.
Deterministic evals score a transcript without calling a model
The claim that makes the rest testable is that captured transcripts can be scored without calling a real language model. That turns an evaluation suite from a nightly cost into a deterministic check, and it is why the evaluation scripts are shell scripts with a release mode rather than something that calls out to a provider. Combined with the event log and the recorded gate verdicts, it means the pipeline has three sources of truth a reviewer can read after the fact: what happened, what the gates said, and what a fixed rubric scored. The published cost of that design is volume. The repository ships a plugin manifest for several hosts, a privacy document, a security document, a design document, seven translated readmes, and a directory of examples, and the model manifest's own description calls the thing a harness engineer rather than the name the project uses everywhere else.
Editorial conclusion
This is worth reading for its evidence model rather than its role names, because the event log, the recorded gate verdicts and the deterministic evals are the parts that survive a change of agent. The honesty about enforcement limits is the strongest signal here, and a team that understands its gates are checkable rather than binding can decide what to compensate for. Two practical notes. Distribution is marketplace-only, so budget for the copy-based install path and its manual hook wiring. And if the zero-network claim matters to you, run the doctor command first, since the project makes it auditable rather than merely stated.
Frequently asked questions
What is the boss-skill project?
An auditable agent-team workflow for coding agents. It turns one agent into a structured team of product, architecture, interface design, tech lead, scrum master, frontend, backend, quality and operations roles, adding runtime state, append-only events, quality gates, deterministic evals, hooks and replayable artifacts that a prompt-only setup cannot produce.
How do I install boss-skill?
Through a marketplace rather than a package. For Claude Code, add the marketplace and install the boss plugin. For Codex, add the marketplace with its own command and install from the plugin browser. Any agent can use `npx skills add echoVic/boss-skill`, which prompts for target agent, scope and method and writes a lock file. There is no npm package and no global binary.
What files does boss-skill write into a project?
A per-feature folder holding a design brief, a product requirements document, an architecture document, a task list and a QA report, plus a metadata directory with an append-only events log, an execution state file and a workflow plan. The read-only review role writes just its review file and is promised to touch nothing else.
Can an agent ignore a failed quality gate in boss-skill?
In practice, yes, and the project says so. The final gate command and the doctor command fail when a completed stage carries a failed gate, but enforcement still depends on the orchestrating agent honouring the protocol, so the CLI makes the verdict checkable rather than unavoidable.
Does boss-skill make network calls?
By default none: no telemetry, no phone-home and no remote model proxy. The only network surface is an opt-in preview server bound to the loopback address, a source-level network-boundary test fails continuous integration if an outbound client reappears, and the doctor command checks the boundary at runtime.
Which coding agents does boss-skill work with?
Claude Code, Codex, OpenClaw, Antigravity and Hermes, each with its own install path, and an upgrade command that updates the skill and re-merges its hooks. Hosts that copy files instead of installing plugins need the hooks wired manually with the bundled command line entry point.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/echovic-boss-skill)