# SkillSpec: turning SKILL.md files into testable contracts for coding agents

> SkillSpec is a Rust CLI that audits a skill folder for agent follow-through risk and compiles it into a structured contract. It is aimed at teams whose agents skip late safety rules and report done without evidence.

**modiqo/skillspec** — SkillSpec makes agent skills followable, testable, and provable with Doctor risk reports, guided imports, structured contracts, and alignment proof.

- Repository: https://github.com/modiqo/skillspec
- Website: https://skillspec.sh
- Stars: 724 · Forks: 58
- Language: Rust
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/modiqo-skillspec

## The failure mode SkillSpec is built around

The README states the problem plainly: a SKILL.md is just text, the harness loads it, and the model reads whatever it reads. Three specific failures are named. A never do X rule sitting around line 400 gets skipped because models are most reliable at the start and end of context, not the middle. Each miss tends to produce another paragraph of prose, which makes the next miss more likely. And the only artifact you get back is the final answer, with no durable record of which route ran or what was skipped.

That framing matters because it points at a different fix than most agent tooling. SkillSpec does not try to make the model smarter or wrap it in a new runtime. The README is explicit that there is no new agent runtime and no orchestration platform. Instead it moves the load-bearing parts of a skill out of prose and into a small structured contract: when to use the skill, which route to take, what is forbidden, what dependencies must exist, what checks must pass, and what proof should exist at the end. The audience is whoever owns a skill that other people rely on, which in practice means platform or developer-experience teams inside companies running Claude Code or Codex.

## What Doctor actually scores, and what it cannot know

The first command most people will run is Doctor against a skill folder or a public GitHub URL. The README shows the output shape: a target, a shape label such as simple_skill, a risk score rendered as HIGH (74/100), then findings, a likely consequence, and a next step. The findings in the example are concrete and checkable: a short generic description that may hurt automatic discovery, an active skill load of 8,482 tokens described as above the balanced target, 14 must/never obligations appearing after 60 percent of the body, tools and commands used but dependencies never declared, and no tests or progress surface so completion cannot be checked.

Those categories tell you what the tool measures. It is counting and locating obligations, estimating context load, looking for declared dependencies, and looking for a trace or test surface. It is not executing the agent and watching it fail. The score is a static assessment of the skill text and its structure, so a HIGH score is a prediction about follow-through risk, not a record of an actual miss. The README's own example ends with a suggested prompt asking an agent to import the skill, compile it, test it, install it, and print the alignment summary, which is where the dynamic side of the loop lives.

## Install and a first run

The README gives two ways to get the CLI. The install script verifies a checksum and writes to ~/.local/bin by default. Cargo is the alternative if you already have a Rust toolchain. Both are followed by a version check, which is the quickest way to confirm the binary landed on your PATH.

```bash
curl -fsSL https://skillspec.sh/install.sh | sh
skillspec --version
```

```bash
cargo install skillspec
skillspec --version
```

The README also documents pinning a version and choosing an install directory through environment variables, and notes prebuilt archives for macOS, Linux x86_64, and Windows x86_64 on the releases page.

Once the CLI is present, point Doctor at a skill folder. This is the no-plugin path, and the README says no install is required to try it if you use the hosted page instead.

```bash
skillspec doctor ./my-skill
```

Expect a report in the shape shown in the README: risk level and score, a findings list, a likely consequence, and a next step. If your skill has a long body with obligations near the end, the finding about late must/never rules is the one to read first, because it maps directly to the failure mode the project was built around.

For harness integration, the README documents plugin installs for Claude Code and Codex, and a skillspec install skill subcommand with targets codex, agents, and claude-local. The --retire-existing flag appears in the README's local development examples.

```bash
claude plugin marketplace add modiqo/skillspec --sparse .claude-plugin plugins/skillspec
claude plugin install skillspec@skillspec
```

```bash
skillspec install skill skills/skillspec --target codex --retire-existing
```

## The contract file and why it is YAML next to SKILL.md

The design choice worth noting is placement. The contract is a skill.spec.yml that lives next to your SKILL.md, not a separate registry entry or a server-side object. That keeps the contract reviewable in the same pull request as the skill text, and it means a skill folder remains self-describing when copied between repositories.

The README lists what the contract carries: trigger conditions, route selection, forbidden actions, required dependencies, required checks, and end-of-run proof. That is a narrower surface than the prose it supplements. The trade-off is real. Anything you move into the contract has to be expressible as a field, and the README does not describe how to express judgment calls that resist structured form. A skill whose correctness depends on tone, or on a conditional that only a human would recognize, will still live mostly in prose. SkillSpec reduces the chance that a load-bearing rule is buried, but it does not eliminate the prose, and it does not claim to.

The workspace layout supports the split the README describes. Cargo.toml lists ten crates, including skillspec-doctor, skillspec-boundary, skillspec-harness, and skillspec-workspace. Separating the doctor from the harness layer is consistent with a tool that wants to assess a skill without owning the runtime that executes it.

## Where SkillSpec is the wrong tool

If your skills are short prompts you rewrite weekly, the contract is overhead. You will spend more time maintaining skill.spec.yml than you save, and the risk findings will mostly restate what you already know about a 40-line file.

The second limit is structural. SkillSpec measures the skill, not the agent. A skill can pass Doctor and still fail in practice because the model ignored a well-placed rule, or because the harness truncated context before the rule was reached. The README's own framing acknowledges this: the alignment summary and the test step exist precisely because the static report is not proof of behavior. Anyone expecting Doctor to be a verdict on whether their agent will behave correctly is reading it wrong.

The third limit is that the README does not document rollback. The install subcommand shows a --retire-existing flag, and the plugin commands install into a harness, but there is no described procedure for cleanly removing a skill or reverting a contract change after it has been installed. If you are deploying to a fleet of developer machines, plan for that gap before you standardize on it.

## How it differs from prompt linters and eval harnesses

Two adjacent categories are worth separating. Prompt linters check text style and structure. Eval harnesses run a model against a dataset and score outputs. SkillSpec sits between them. Like a linter, it reads the skill without executing it, and Doctor's findings are structural. Like an eval harness, it cares about whether the agent followed the instructions and wants a record at the end, which is what the contract and alignment summary provide.

The difference in approach is that SkillSpec does not own the model call. It attaches to an existing harness through a plugin or an install target, and the proof comes from the run the harness already performed. That is a lighter integration than standing up a separate evaluation service, and it is also a narrower one: SkillSpec can only report on what the harness surfaces. If your harness does not expose a trace, the proof surface the README describes has nothing to read. That constraint follows directly from the no-new-runtime decision, and it is the main thing to check against your own setup.

## Licence, releases, and upgrade cost

The repository is Apache-2.0, and the Cargo.toml workspace declares license = "MIT OR Apache-2.0" with LICENSE-APACHE and LICENSE-MIT both present at the top level. That dual licensing at the crate level, alongside the Apache-2.0 repository licence, is worth confirming with your own legal review before you redistribute binaries; this is a description of the files, not legal advice.

Versioning is moving. The recent releases are v0.2.0, v0.2.1, and v0.2.2, all dated 2026-07-28 and 2026-07-29, which is a patch cadence inside a single week. The last push to main was on 2026-08-09. Pre-1.0 with a fast patch series means you should pin the version in CI rather than tracking latest, and the README documents exactly that: SKILLSPEC_VERSION and SKILLSPEC_INSTALL_DIR as environment variables on the install script. The README also notes that tagged releases publish crates in dependency order, and that PR CI uses package file-list checks instead of cargo publish --dry-run because a same-version split crate graph cannot dry-run downstream crates until their siblings are on crates.io. That is a real constraint on anyone trying to vendor or republish the crates themselves.

## Conclusion

Adopt SkillSpec if you maintain skills that carry safety rules or undeclared tool calls and you want a contract plus an alignment summary instead of prose alone. Skip it if your skills are throwaway prompts, or if you cannot run an extra CLI step in your build. Before adopting, run skillspec doctor on one real skill, check that the risk score matches what you have seen fail, and confirm your harness is one of the targets the install command supports.

## FAQ

### What does SkillSpec actually do with my SKILL.md file?

It reads the skill and produces a risk report through the doctor command, then lets you compile the skill into a skill.spec.yml contract that sits next to the original file. The contract records trigger conditions, routes, forbidden actions, required dependencies, required checks, and end-of-run proof.

### Do I need to install a new agent runtime to use SkillSpec?

No. The README states there is no new agent runtime and no orchestration platform. SkillSpec is a CLI plus a contract file, and it attaches to an existing harness through a plugin or an install target such as codex, agents, or claude-local.

### How do I install the SkillSpec CLI?

The README gives two paths: the install script at https://skillspec.sh/install.sh, which verifies a checksum and writes to ~/.local/bin by default, or cargo install skillspec. Both are followed by skillspec --version to confirm the binary is on your PATH.

### What does a HIGH score from SkillSpec Doctor mean?

It is a static assessment of follow-through risk based on findings such as late must/never obligations, undeclared dependencies, and the absence of a test or trace surface. The README presents it as a prediction, with the contract and alignment summary as the steps that provide actual proof.

## Sources

- [License: Apache-2.0](https://github.com/modiqo/skillspec/blob/main/LICENSE)
- [modiqo/skillspec on GitHub](https://github.com/modiqo/skillspec)
- [Project website](https://skillspec.sh)
- [README](https://github.com/modiqo/skillspec/blob/main/README.md)
- [Releases](https://github.com/modiqo/skillspec/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/modiqo-skillspec
