# xiaolai/nlpm: linting natural-language artifacts across Claude Code, Codex CLI and Antigravity

> NLPM scores the markdown files that drive AI behaviour, and its standalone validator checks a bug class the README says other validators miss: artifacts that exist on disk but never make it into a manifest.

**xiaolai/nlpm** — Natural-Language Programming Manager — scan, lint, and score NL artifacts with Claude-native quality scoring

- Repository: https://github.com/xiaolai/nlpm
- Website: https://nlpm.com/
- Stars: 144 · Forks: 40
- Language: HTML
- License: ISC
- Published: 2026-08-24 · Updated: 2026-08-24 · Language: en
- Canonical page: https://hysenlabs.com/projects/xiaolai-nlpm

## The bug class nlpm was built around

A SKILL.md can sit on disk, be perfectly written, and still be invisible after `claude plugin install`, because the plugin manifest never listed it. Nothing crashes. The skill simply is not there. The README frames this as the problem NLPM exists to solve, and states that the project is the only multi-tool NL artifact validator that systematically checks manifest-vs-disk consistency. It also states this was verified across 8+ tools including Anthropic's official `plugin-validator` and the Linux Foundation's `skills-ref`, with the research written up in `analysis/ecosystem-gap.md`.

That claim is the whole reason to care. Most linting for AI artifacts looks at one file at a time. This one compares what a manifest declares against what the repository actually contains, which is a cross-file property no single-file linter can express. The audience follows from that: plugin and skill authors who publish through a marketplace, and who have already been bitten by a component that silently failed to load.

## Eight commands, one rubric, three tool overlays

NLPM treats natural language artifacts as programs that can be linted. The README draws the comparison directly: ESLint scores JavaScript, ruff scores Python, NLPM scores the markdown that drives AI behaviour, covering skills, agents, commands, rules, hooks, prompts, CLAUDE.md and memory files.

Eight slash commands each do one thing. `/nlpm:ls` discovers and inventories artifacts. `/nlpm:score` scores quality on a 100-point scale. `/nlpm:check` runs cross-component consistency checks. `/nlpm:fix` auto-fixes what is fixable. `/nlpm:trend` tracks score history. `/nlpm:test` runs artifact tests against spec files in a TDD style. `/nlpm:init` sets up a project. `/nlpm:security-scan` scans plugins for risks in executable artifacts.

Scoring is deterministic by design: scores start at 100, every issue carries a fixed penalty, and the same artifact produces the same number. The bands run from 90-100 (Excellent, described as production-ready) down to below 60 (Rewrite, for fundamental problems), with a default pass threshold of 70 configured in `.claude/nlpm.local.md`. Penalty tables live in `skills/nlpm/scoring/SKILL.md` and the 50 Rules of Natural Language Programming in `skills/nlpm/rules/SKILL.md`.

The multi-tool part is a tier-aware overlay system: one universal floor plus per-tool adjustments for Claude Code, Codex CLI and Antigravity, documented in `analysis/multi-tool-design-2026-05.md`. That is a reasonable answer to a real problem, since the same skill file can be valid for one host and wrong for another. The trade-off is that overlays multiply the rubric surface you have to keep current as each host changes its conventions.

## Installing nlpm and scoring a repository for the first time

There are two install paths and the README says both reach the same code. The difference is update latency: the community marketplace is curated and its updates lag the maintainer's marketplace by up to about 24 hours.

Via Anthropic's community marketplace:

```bash
claude plugin marketplace add anthropics/claude-plugins-community
claude plugin install nlpm@claude-community --scope project
```

Via the maintainer's marketplace, where the README says the latest version lands first. Project scope is recommended there, with a user scope available for all projects:

```bash
claude plugin marketplace add xiaolai/claude-plugin-marketplace
claude plugin install nlpm@xiaolai --scope project
```

The README documents one failure mode for the second path: if install reports that the plugin was not found in marketplace 'xiaolai', your local clone is stale, because `plugin install` does not auto-refresh. The fix is `claude plugin marketplace update xiaolai`, then retry. The community marketplace does not have that caveat.

Once installed, the first real use is inventory then score. Run these inside Claude Code:

```
/nlpm:ls
/nlpm:score
/nlpm:score --changed
```

The first lists the NL artifacts in the repository. The second scores all of them. The third narrows scoring to git-changed files, which is the form you want once a repository has already been brought up to standard and you only care about the diff. From there, `/nlpm:check` runs the cross-component consistency checks and `/nlpm:fix` applies fixes for the issues it can repair automatically.

## The standalone validator, and where it fits in CI

The slash commands require Claude Code. The standalone binary does not. `bin/nlpm-check` is a single Python 3.11+ file with no external dependencies, and the README positions it for pre-commit hooks, CI and pre-publish gates. It runs the deterministic subset of `/nlpm:check`, including the manifest-vs-disk consistency check.

The README gives a one-line install that fetches the file directly from the repository:

```bash
curl -fsSL -o /usr/local/bin/nlpm-check \
  https://raw.githubusercontent.com/xiaolai/nlpm/main/bin/nlpm-check
chmod +x /usr/local/bin/nlpm-check
nlpm-check .
```

Run against a repository, it exits 1 on high-confidence findings, which is what makes it usable as a gate. Templates ship in `templates/`: `pre-commit-nlpm.sh` as a drop-in git pre-commit hook and `workflows/nlpm-check.yml` as a drop-in GitHub Actions workflow. The author guide is `docs/for-authors.md`.

Note the asymmetry the README itself implies. The standalone path is the deterministic subset, so the richer, judgement-based parts of the scoring live only in the Claude Code plugin. If you want everything, you install the plugin. If you want a gate that runs anywhere without an assistant in the loop, you take the single file and accept the narrower rule set.

## The self-auditing pipeline and its drift problem

Beyond linting, the repository runs as a self-evolving GitHub Actions pipeline that audits real plugin repositories, contributes fix pull requests, harvests teaching examples from clean ones, and feeds results back into its own rule catalog. The README dates several stages: the exemplar pipeline arrived in v0.8.17, rule-citation auto-PR in v0.8.18, and the two-stage drift detector across v0.8.15 and v0.8.16.

The exemplar pipeline is the clearest mechanism. Repositories that audit clean at a score of 90 or above produce a teaching artifact under `auditor/exemplars/`; the README reports 62 published so far, covering 31 of the 50 Rules with real-world positive references. A weekly workflow, `auditor-cite-exemplars.yml`, opens a human-gated pull request that adds `> Real-world example: [<repo>]` links to `skills/nlpm/rules/SKILL.md`, so each rule documents both a bad case in the rule body and a good case in a real repository.

The drift detector is the honest part of the story. `auditor/scripts/validate-rule-ids.py` re-validates each audit's `rule_id` against the rubric for type drift and against the rule's title keywords for semantic drift. A sweep dated 2026-05-13 found 990 mislabeled `rule_id`s across 128 historical audits. The validator is now wired as a soft-warn telemetry step in every new audit. That number is worth sitting with: a scoring system that had been running long enough to accumulate 128 audits was also mislabelling citations at scale, and the correction is soft-warn rather than blocking. `auditor/scripts/rule-health.py` reports `validated_hits` per rule (raw hits minus drift hits) alongside `exemplars_count`, so the rule health view is calibrated against real violations rather than scorer noise.

## Where nlpm is the wrong tool

The rubric is structural. It checks whether artifacts are consistent, complete and well-formed across components. It is not a prose critic and the README does not present it as one, so a team hoping to improve the clarity of a long CLAUDE.md will not get that from a penalty table.

There is a second, sharper limit. The deterministic score is only as good as the rules behind it, and the repository's own drift sweep shows how easily rule identifiers and their intended meanings separate over time. A penalty table is a maintenance commitment, not a one-time artifact. Anyone adopting NLPM inherits that commitment, and the README does not describe a deprecation or migration process for rules that stop being relevant.

Third, the standalone validator is deliberately the narrower path. If your gate must run without Claude Code, you get the deterministic subset, not the full scoring. Teams that need the complete rubric in CI have no documented option here.

Finally, the tooling assumes a repository layout with recognizable manifests and artifact directories. A project that keeps its prompts inline in application code, or generates them at runtime, has little for `/nlpm:ls` to inventory in the first place.

## Alternatives and how they differ

The README names two other validators directly. Anthropic's official `plugin-validator` and the Linux Foundation's `skills-ref` both appear in the ecosystem-gap research, and the README's position is that neither covers manifest-vs-disk consistency. The difference in approach is scope: those tools validate what a single artifact should look like, while NLPM compares the manifest against what is actually present on disk. If your failure mode is a malformed skill file, either of the others is a reasonable fit. If your failure mode is a skill that never loads, the check you want is the cross-file one.

There is also the general-purpose option of running a markdown linter plus a link checker in CI. That combination catches formatting and broken references, and it costs nothing to add. It will not tell you that a component is missing from a manifest, because it has no model of the manifest at all.

NLPM's own differentiator is breadth of host coverage. The tier-aware overlays mean one rubric covers Claude Code, Codex CLI and Antigravity artifacts, so a repository that publishes to more than one host does not need a separate validator per host. That is a real advantage for multi-host authors and irrelevant for anyone working in a single ecosystem.

## Licence and upkeep

The repository is licensed under ISC, a permissive licence that permits use, modification and redistribution provided the copyright notice and permission notice are retained. The README does not discuss commercial support, warranties or contributor terms beyond that, and nothing here is legal advice: read the LICENSE file in the repository if the distinction matters to your organisation.

Upgrade cost has two parts. The plugin path updates through the marketplace, and the README is explicit that the maintainer's marketplace carries new versions first while the community marketplace lags by up to about 24 hours. The standalone path is a single file fetched from `main`, so it has no version pinning in the documented install command. Pulling from `main` on every CI run means your gate can change without a commit in your repository. Pinning to a specific revision is the obvious mitigation, but the README does not document one.

The rubric itself is the larger ongoing cost. Fifty rules, penalty tables, per-tool overlays and a drift validator all need attention as the three host ecosystems evolve. The repository's own tooling acknowledges this by measuring drift rather than assuming it away.

## Conclusion

Adopt nlpm if you author Claude Code plugins or skills and want a deterministic gate before publishing, especially the manifest-vs-disk check that the README says Anthropic's plugin-validator and the Linux Foundation's skills-ref do not cover. Do not adopt it as a general prose-quality tool: the rubric is about artifact structure and consistency, not writing. Before you commit, verify two things yourself: that your local marketplace clone is current, since plugin install does not refresh it, and that the default pass threshold of 70 matches what your repository can actually reach today.

## FAQ

### Does nlpm require Claude Code to run?

No. The slash commands ship as a Claude Code plugin, but the standalone validator at bin/nlpm-check is a single Python 3.11+ file with no external dependencies and runs in pre-commit hooks or CI on any tool's artifacts. It covers the deterministic subset of the checks rather than the full scoring rubric.

### Why does claude plugin install nlpm@xiaolai fail with a plugin-not-found error?

The README attributes this to a stale local marketplace clone, because plugin install does not auto-refresh. Run claude plugin marketplace update xiaolai and retry. The community marketplace does not have this caveat.

### What is the default pass threshold for an nlpm score?

The default pass threshold is 70, and it is configured in .claude/nlpm.local.md. Scores start at 100 and go down by fixed penalties, so the same artifact always produces the same number.

### What does the manifest-vs-disk consistency check actually catch?

It catches the case where an artifact such as a SKILL.md exists on disk but is missing from plugin.json, which means it is invisible after claude plugin install. The README states this check is not covered by Anthropic's official plugin-validator, the Linux Foundation's skills-ref, or third-party tools.

## Sources

- [Official documentation](https://nlpm.com/)
- [Official README](https://github.com/xiaolai/nlpm#readme)
- [Project repository](https://github.com/xiaolai/nlpm)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/xiaolai-nlpm
