# The sample output argues with itself about level L3

> harness-score is a deterministic scanner that grades the context files, rules, skills, hooks and CI wiring around an AI coding agent. Its selling point is that every check is a filesystem fact, and its most instructive detail is a printed example whose numbers do not agree with the hint printed underneath them.

**paladini/harness-score** — Your AI coding agent is only as reliable as the harness around it. Measure that harness in seconds with harness-score.

- Repository: https://github.com/paladini/harness-score
- Website: https://paladini.github.io/harness-score/
- Stars: 545 · Forks: 46
- Language: TypeScript
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/paladini-harness-score

## The sample run argues with its own hint

The block of sample output at the top of the README is the most quoted part of the project, and it does not add up. The run reports Maturity L2, a score of 70 out of 108, and six dimension lines: Context and Guides at 16 of 20, Skills and Commands at 11 of 17, Hooks and Guardrails at 0 of 14, Sensors and Feedback at 16 of 20, CI Feedback at 10 of 14, and Hygiene and Safety at 17 of 23. Those six numbers do sum to 70, so the arithmetic is sound.

The closing hint does not follow. It reads `To reach L3: sensors ≥ 60%; ci ≥ 50%`, while the same block reports Sensors at 80% and CI at 71%. Both conditions are already met, and the guidance says the scanner always names the requirement that blocks the next level, so the only way to read it is that something not printed is also gating the climb. Anyone copying this example as a template should check the real requirement rather than assume the hint is complete.

## Levels gate on shape, so points alone cannot lift you

The maturity model refuses to let raw points decide the level. The stated example is blunt about it: eighty points of well written documentation with zero tests is not maturity, it is L1. Climbing means covering a new dimension at each rung rather than deepening one you already have.

```
  harness-score v1.0.0  ~/my-app

  Maturity: L2 · Guided   Score: 70/108 (65%)
  Detected: Cursor, Claude Code

  Context & Guides     ████████████████░░░░  80%  16/20 pts
  Skills & Commands    █████████████░░░░░░░  65%  11/17 pts
  Hooks & Guardrails   ░░░░░░░░░░░░░░░░░░░░   0%   0/14 pts
  Sensors & Feedback   ████████████████░░░░  80%  16/20 pts
  CI Feedback          ██████████████░░░░░░  71%  10/14 pts
  Hygiene & Safety     ███████████████░░░░░  74%  17/23 pts

  To reach L3: sensors ≥ 60%; ci ≥ 50%
```

The ladder runs L0 Unharnessed, L1 Documented, L2 Guided, L3 Sensing, L4 Self-correcting, each with one priority to climb. The top rung's advice is to gate CI on `--min-level 4`, and the header of that sample still reports v1.0.0 while the published releases are at v1.7.5.

## Thirty-six checks, 108 points, and no judgment anywhere

The scoring model is small enough to read in one table. Thirty-six checks spread across six dimensions add up to 108 points: Context and Guides at 20, Skills and Commands at 17, Hooks and Guardrails at 14, Sensors and Feedback at 20, CI Feedback at 14, and Hygiene and Safety at 23. Hygiene carries the most weight of any single dimension, and the two smallest are hooks and CI.

Each check names a concrete artifact instead of an aspiration: AGENTS.md substance, .cursor/rules scoping and frontmatter, packaged skills and slash commands, gate hooks and feedback hooks, a test runner with actual test files, a pipeline that runs on every push, and a gitignore with no leaked secrets. The determinism is the point rather than a limitation to apologise for. Zero LLM calls, zero network requests, and the same score for the same repository and commit on a laptop or in CI is what makes the number usable as a build gate.

## Four things it refuses to claim

The section on what the scanner deliberately does not measure is the most honest part of the documentation, and it is short enough to quote in full. It does not tell you whether your tests are good, only that they exist, run and gate. It does not tell you whether your rules are true, so a stale rule scores exactly like a fresh one. It cannot verify functional correctness, because no static scan can. And it says nothing about team practice, since branch protection and review culture live outside the tree.

The framing in that section is that a high score means the infrastructure for reliable agent work exists, and calls it necessary but not sufficient, the honest ceiling of what any deterministic scanner can claim. That framing matters for the maturity ladder as well, since the rungs describe capability rather than correctness. A repository at L4 has the control system in place; nothing in the model says the agent inside it writes good code.

## The scan reads ignored files on purpose

The filesystem behaviour is specified in detail, and one line in it deserves attention. The scan walks the complete relevant tree, including tracked, untracked and ignored files. It skips known dependency and generated directories, keeps followed symlink targets inside the scan root, and never reads file bodies larger than 512 KiB.

Reading ignored files is not an oversight, it is a requirement of the design. The hygiene dimension scores whether secrets have leaked, and it cannot answer that without opening the .env files a gitignore was meant to keep out of the history. The consequence is that the scanner touches exactly the files you would least like a tool to read, which makes it worth knowing before you point it at a checkout containing real credentials.

Two guards bound the walk: followed symlinks must resolve inside the scan root, and an emergency fuse stops pathological trees above 1,000,000 files. Three failure conditions are named, fuse hit, out-of-root symlink, and an unreadable relevant path, and the sentence breaks off after the third one without stating what the scanner does about it.

## The npm scan script runs a prebuilt binary

The repository is a private monorepo whose root package is named harness-score-monorepo at version 0.1.0, with workspaces pointing at packages/*. The scan script is one line: it runs node against packages/cli/dist/cli.js with the current directory as the argument.

That dist path is worth pausing on. The build script is separate, so the scan uses whatever was last compiled into it, and a stale build means scoring the repository with an older checker. The release pipeline treats scanning as one of its own gates: release:prepare runs the versioning step, the tests, the lint check, the scan, the docs build and the release notes script in that order. Versioning goes through changesets followed by a sync-version script, and the test command itself includes a separate node test run over scripts/release.test.mjs, so the release tooling has its own tests.

Other scripts reveal the scope: a devin verification script, plugin generators with a check mode for CI, a bench script in the workspace, and husky wired up through prepare so git hooks install with the dependencies.

## Two registries, a docs site, an action and a plugin

The distribution surface is wider than the npm badge suggests. Badges point at both the npm package and a JSR scope, and the documentation is a VitePress site under a GitHub Pages URL, with dev, build and preview scripts for it. The repository also carries action.yml and an action directory, so the scanner can run as a GitHub Action, and a .claude-plugin directory with a plugins tree and a PLUGINS-ROADMAP.md, so it can also ship as an agent plugin. The plugin generators have a check mode, meaning the committed plugin output is verified in CI rather than regenerated by hand.

The declared engine floor is Node 18 or newer, the formatter and linter are biome, versions move through changesets, and husky manages the pre-commit side, which is also one of the things the CI Feedback dimension scores. For a tool whose argument is that your pipeline should enforce standards, it is built out of the same parts it measures.

## Three patch releases in two days, one with a cut title

The recent release history is worth reading closely. v1.7.3 is titled around a Claude Code guide scoring change and is dated 2026-09-30. v1.7.4, dated 2026-10-01, is titled as a release attribution correction, which is to say a published release was credited to the wrong thing. v1.7.5 lands the same day with a title that stops mid-word, Recognize Astral ty.

Those tags are dated later than the 2026-09-15 push recorded for the default branch, so the releases and the branch history in the repository metadata do not line up in time. Two days and three patch versions is also a fast cadence for a tool that measures maturity and preaches measured process, and a title that needs completing is exactly the kind of detail the tool's own hygiene dimension would flag in someone else's repository.

## Conclusion

Use it as a shape check, not a quality grade, and read the ceiling section before you trust a number. Two things to verify in your own repository first: the sample output in the docs does not reconcile with its own threshold hint, so confirm what the tool actually requires for the next level rather than copying the example, and remember that the scan reads ignored files on purpose, because the hygiene dimension has to open .env files to know whether secrets are in them.

## FAQ

### What does harness-score actually measure?

The harness around an AI coding agent, meaning context files, rules, skills, hooks, sensors and guardrails. It reports a maturity level from L0 to L4, a 108-point breakdown across six dimensions built from 36 checks, and a ranked list of what to fix next.

### Does harness-score call an LLM or use the network?

No. Every check is a filesystem fact such as a file existing, parsing or matching a pattern, and the tool states it makes zero LLM calls and zero network requests, producing the same score for the same repository and commit.

### What does harness-score not measure?

It does not tell you whether your tests are good, whether your rules are true, whether the code works, or how your team practises review. A high score means the infrastructure exists, which the project calls necessary but not sufficient.

### How do I use harness-score in a pipeline?

Because the score is deterministic for a given commit, it can gate a build. The scanner names the requirement that blocks the next level, and the top rung of the maturity ladder suggests gating CI on `--min-level 4`.

### Which files does harness-score read from my repository?

It walks the complete relevant tree including tracked, untracked and ignored files, skips known dependency and generated directories, keeps followed symlink targets inside the scan root, and never reads file bodies larger than 512 KiB.

## Sources

- [License: MIT](https://github.com/paladini/harness-score/blob/main/LICENSE)
- [paladini/harness-score on GitHub](https://github.com/paladini/harness-score)
- [Project website](https://paladini.github.io/harness-score/)
- [README](https://github.com/paladini/harness-score/blob/main/README.md)
- [Releases](https://github.com/paladini/harness-score/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/paladini-harness-score
