# AutoHarness: A Self-Maintaining Skill Layer for Claude Code

> AutoHarness is a MIT-licensed Claude Code plugin that distills reusable skills from real working sessions, merges near-duplicate skills rather than stacking them, and prunes skills that stop being used. It requires no third-party Python dependencies and touches only the skills it generated itself.

**tigerless-labs/autoharness** — Autoharness — a self-learning skill layer for Claude Code — distills skills from your real sessions, updates them as you work, and prunes the ones that stop getting used. No daemon, no benchmark.

- Repository: https://github.com/tigerless-labs/autoharness
- Website: https://github.com/tigerless-labs/autoharness
- Stars: 5,831 · Forks: 490
- Language: Python
- License: MIT
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/tigerless-labs-autoharness

## What AutoHarness Does to the Skill Layer

Claude Code's skill system lets users define reusable instruction sets that the model loads when relevant to a session. In practice, a skill library built up over many sessions can accumulate near-duplicate skills for similar scenarios, skills that were used once and never again, and skills that worked for an older workflow and are now stale. Cleaning that up requires manual review.

AutoHarness is a plugin that automates this maintenance. It watches running Claude Code sessions, distills what was done into a skill candidate when enough tool calls have passed, compares the candidate against the existing library, and either adds a new skill or folds it into an existing one that covers the same scenario. Skills that stop being adhered to in later turns are eventually pruned. The README frames this as the skill layer maintaining itself.

The plugin is described as running entirely as Python with zero third-party dependencies. It installs via Claude Code's plugin marketplace. The last push to the repository was on 2026-09-04, and the project is MIT licensed.

## How Distillation, Merging, and Pruning Work

The README describes four named behaviors. Learning happens automatically: every session counts tool calls, and when the count crosses AUTOHARNESS_REFLECT_EVERY_N (default 50), a background reflection fires at the end of that turn. The reflection is a subagent that receives a compressed summary of the session and produces a skill candidate.

Grouping, not just adding: the reflector compares the candidate against the existing library and decides whether the scenario matches an existing skill closely enough to fold into it rather than creating a new entry. A fold records which skill absorbed which, so a merge is not confused with a pruning event.

A skill survives by being adhered to in subsequent turns. The README states that adherence is measured by how often the skill is loaded over the turns it was available for. A skill that is never followed in practice is removed. No separate held-out evaluation is run; the production sessions are the evaluation.

The /learn command provides a manual trigger. Typing it in the Claude Code input box distills the current session immediately, without waiting for the background cadence. This is useful when a session just solved a tricky problem that should be captured as a skill right away.

The consolidation pass (AUTOHARNESS_CONSOLIDATE_EVERY_N, default 250 tool calls) runs less frequently and reviews the library as a whole, merging across all categories rather than just the latest distillation.

## Installing AutoHarness from the Plugin Marketplace

AutoHarness installs via Claude Code's plugin marketplace. Type these commands in the Claude Code input box:

```
/plugin marketplace add tigerless-labs/autoharness
/plugin install autoharness@autoharness
```

Then run /reload-plugins, or restart Claude Code, to activate the plugin. From that point, no further setup is needed. The plugin watches sessions and writes learned skills into .claude/skills/ in the background.

To update the plugin, run these commands in a terminal (not the Claude Code input box) and then restart Claude Code:

```
claude plugin marketplace update autoharness
claude plugin update autoharness@autoharness
```

The README notes that the catalog refresh must come first. Without it, the update command checks a stale local catalog and may report that the plugin is already at the latest version when a newer release has shipped.

To uninstall:

```
claude plugin uninstall autoharness@autoharness
claude plugin marketplace remove autoharness
```

Uninstalling stops the plugin from running but leaves its state and the skills it wrote on disk. To remove those, delete ~/.claude/autoharness/ and <repo>/.claude/autoharness/ along with any skills that carry a self-authored ledger marker in .claude/skills/.

## Configuring Cadence and Recall with Environment Variables

Every configuration option is an AUTOHARNESS_* environment variable with a default. Nothing needs to be changed to get started.

For learning cadence, AUTOHARNESS_REFLECT_EVERY_N (default 50) sets how many tool calls must occur before a background reflection fires. A lower value means the plugin learns faster and spawns more background subagents. AUTOHARNESS_CONSOLIDATE_EVERY_N (default 250) sets the cadence for the full-library curator pass. The README recommends keeping this well above the reflection cadence since consolidation is rarer.

AUTOHARNESS_DIGEST_EXCHANGES (default 20) controls how many exchanges before the episode window are compressed into a text digest for the reflector's prior-context. AUTOHARNESS_CARRIER (default "bundle") controls how the reflection subagent receives the session context. The bundle mode sends a redacted window plus the digest to a fresh subagent. The fork mode resumes the session and forks it, giving the reflector access to the real conversation on the parent's warm cache. The README notes that fork mode remains experimental.

For recall, AUTOHARNESS_INDEX_SUSPENDED (default 0) controls whether the session-start skill index is injected. Setting it to 1 stops the index from appearing at the start of each session while leaving all other autoharness behavior intact. The README describes this as the switch for measuring what the index actually contributes. AUTOHARNESS_INDEX_DESC_MAX_CHARS (default 60) controls how many characters per skill description appear in the index.

## What AutoHarness Does Not Touch and Where Its State Lives

The README is explicit about scope: autoharness only touches the skills it generated through the plugin. Any skill the user wrote, or any skill installed from another source, is left completely alone. The boundary is enforced through a self-authored ledger marker that each autoharness-generated skill carries in its metadata.

State lives in two locations. The global state directory is ~/.claude/autoharness/. Per-project state lives at <repo>/.claude/autoharness/. The skills autoharness writes appear in .claude/skills/ using the same file format as manually written skills, distinguished by the ledger marker.

The per-skill ledger records the scenario and decision for each create and update event. The README describes this as the raw material for building a benchmark from real usage if the user wants one later, though no benchmark-generation tooling is included in the repository.

The MCP server registered by autoharness is named stage_skill, but within the plugin context the fully qualified name is mcp__plugin_autoharness_stage_skill__stage_skill. The README notes that this translation is automatic when installed as a plugin. Testing the tool directly with the claude CLI without the plugin context gives a different name, and the agent allowlists still reference the plugin-namespaced form, so the README recommends always installing via the plugin path.

## Compared to Maintaining a Skill Library Manually

The alternative to autoharness is maintaining .claude/skills/ by hand: writing skill files when a pattern becomes obvious, editing them when they drift from current practice, and deleting them when they stop applying. This approach gives complete control over every entry and produces a version-controlled library that changes only when the developer decides to change it.

The trade-off is attention. A manual library stays accurate only as long as the developer reviews it regularly. In practice, skill files tend to accumulate without pruning, and useful skills from early in a project may be superseded by later ones without the older versions being removed.

AutoHarness bets that automated distillation from real sessions produces skills that reflect actual practice more accurately than skills written in the abstract. The README cites a performance improvement on CORE-Bench from 42% to 78% when the harness is active versus without it (referenced from the HAL paper at arxiv.org/abs/2510.11977). That number applies to the broader harness concept, not to autoharness in isolation.

Developers who want deterministic, reviewed skills with no background agent activity should stay with a manual library. Developers comfortable with automated skill management and who are willing to monitor what the plugin generates during early adoption will find the reduced maintenance burden meaningful.

## Conclusion

AutoHarness is the right choice for Claude Code users who have noticed their .claude/skills/ directory filling with redundant or stale skills and want the maintenance to happen without intervention. The plugin touches only the skills it generated, so manually written skills are safe. Developers who want full control over every skill entry, who do not want any background subagent activity during sessions, or who prefer to version-control a curated skill library without automated changes will find the manual approach simpler. The AUTOHARNESS_INDEX_SUSPENDED=1 environment variable suspends the index injection without stopping distillation, which provides a clean way to measure what the skills are actually contributing before committing to the full setup.

## FAQ

### What is AutoHarness and how does it work?

AutoHarness is a Claude Code plugin that automatically distills skills from your real sessions. It counts tool calls during a session, and when the count crosses a threshold, a background reflection fires that either adds a new skill or merges the insight into an existing one. Skills that stop being followed in practice are eventually pruned.

### How do I install AutoHarness for Claude Code?

Type /plugin marketplace add tigerless-labs/autoharness and then /plugin install autoharness@autoharness in the Claude Code input box, then run /reload-plugins or restart Claude Code. No further configuration is required to start using it.

### What environment variables configure AutoHarness behavior?

All configuration uses AUTOHARNESS_* environment variables with built-in defaults. AUTOHARNESS_REFLECT_EVERY_N (default 50 tool calls) sets the learning cadence. AUTOHARNESS_INDEX_SUSPENDED set to 1 stops injecting the skill index at session start without stopping distillation.

## Sources

- [Issues](https://github.com/tigerless-labs/autoharness/issues)
- [License: MIT](https://github.com/tigerless-labs/autoharness/blob/main/LICENSE)
- [Project website](https://github.com/tigerless-labs/autoharness)
- [README](https://github.com/tigerless-labs/autoharness/blob/main/README.md)
- [tigerless-labs/autoharness on GitHub](https://github.com/tigerless-labs/autoharness)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/tigerless-labs-autoharness
