# audit-harness installs hooks into every project and a Python file that gets imported, not parsed

> A three-layer attempt to stop AI agents from skipping audit records: hardcoded skills, a mandatory output block, and hooks that intercept tool calls. The layering idea is sound and the storage schema is unusually well specified. The deployment surface is the problem, since the installer edits your global instruction file, installs hooks that fire in every project on the machine, and makes project configuration a Python file the runtime imports.

**RickyTong1/audit-harness** — Three-layer audit enforcement framework for AI agents — hooks, skills, context recovery, and audit-driven daily reports

- Repository: https://github.com/RickyTong1/audit-harness
- Stars: 471 · Forks: 3
- Language: Shell
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/rickytong1-audit-harness

## The clone command has a placeholder where the URL should be

The quick start opens like this:

```bash
# 1. Clone/download this repository
git clone <repo> ~/src/audit-harness
```

Angle brackets around the word repo. That is the entire first instruction, and it is a placeholder that was never filled in. There is no clone URL anywhere in the document, so the first thing a new user has to do is supply a URL the documentation declines to provide.

The destination path is a real instruction though. The repository is expected to land at ~/src/audit-harness, and the next step depends on it:

```bash
# 2. Run the installer in the target project (smart mode)
cd /path/to/your/project
bash ~/src/audit-harness/install.sh
```

That hardcoded path is load-bearing. Both the clone command and the installer invocation assume the same location, and the later examples in the document assume it again, so a user who clones elsewhere has to rewrite three commands rather than one.

A related naming detail: the slash commands are unnamespaced. The four skills install as /start, /end, /recover and /report-daily, none of which carries the project name. In a machine where several skills are installed into the same directory, /end in particular is a collision waiting to happen with any other plugin that picked the same obvious name.

The documentation itself is written in Chinese. There is a README.md and a README_zh.md in the repository, and both are Chinese; no English version exists, so an English-speaking user is reading a translation or nothing at all.

## One install edits your global instruction file and hooks that fire everywhere

What the installer does is described in two parts, and the first part is the one to read carefully.

On first run it performs a global install into ~/.claude/, placing the core code, the skills, the hooks and a global CLAUDE.md. Then, for the current project, it creates .claude/audit_config.py and .claude/runs/ and injects the audit specification into that project's CLAUDE.md.

So a per-project tool rewrites your user-level instruction file. The global CLAUDE.md applies to every session you start on that machine, in every project, whether or not this framework was ever initialised there. Nothing in the visible text describes a backup, a diff preview or a confirmation prompt, which is the difference between appending to a config file and doing it unasked.

The hooks land in the same global area, ~/.claude/audit-harness/hooks/, and there are three of them, bound to three Claude Code events:

| Hook | Event | Behaviour |
|---|---|---|
| `post_tool_audit.sh` | PostToolUse | Records Write, Edit and Bash operations to a buffer file |
| `stop_flush.sh` | Stop | Archives the buffer and pending records after each reply turn |
| `prompt_inject_session.sh` | UserPromptSubmit | Injects the current session context on every prompt |

PostToolUse means every write, edit and shell command. UserPromptSubmit means every prompt you type. In projects you have never initialised, these hooks still fire, and the third one still injects context on each turn.

That is the intended behaviour of a global install and it is not hidden, but the name of the repository understates it. This is not a per-project audit harness that you add where you want it. It is a machine-wide layer, and the per-project initialisation only decides which directory gets the data.

There is also a performance shape worth knowing: each hook loads a shared shell helper and then runs `python3 audit_context.py update-index` to write the index. PostToolUse therefore starts a Python process after every tool call.

## Project configuration is a Python file the runtime imports

Initialising a project generates a file called .claude/audit_config.py, and it is Python:

```python
CORE_SCRIPTS = ["main.py", "pipeline.py"]   # key scripts (environment snapshot computes SHA256)
CORE_ASSETS = ["model.pkl"]                 # key assets
PROMPT_TEMPLATE_GLOB = "prompts/*.txt"      # prompt template glob

ALERT_RULES = [
    # business-relevant alert rules (see templates/audit_config.example.py)
]
```

The behaviour when the file is absent is stated as a fallback to module defaults, with everything empty. So a project with no configuration gets no core scripts, no assets, no template glob and no alert rules, and the framework proceeds rather than erroring.

The load order is two paths: first .claude/audit_config.py inside the project, then audit_config.py in the project root. The second location matters more than it looks. A file sitting in the repository root is a file that gets committed, shared with collaborators and read from the current working directory, so the configuration that controls environment hashing is reachable by anything else that can write to the repository.

The content is described as loading, not parsing, and the field names are module-level constants rather than a settings document. That is the difference: a Python file at either of those paths is code that the framework executes, not data it reads. A template is provided at templates/audit_config.example.py, so the intended workflow is to copy and edit it, which means every user is expected to maintain an executable config file by hand.

What the configuration controls is worth noting. Core scripts and core assets feed a SHA256 environment snapshot, so this is drift detection over the files that define your system's behaviour. Alert rules are described as business-relevant and there are none by default, so out of the box nothing raises.

One thing the documentation does get right on this axis is the schema discipline. The index file carries a version number, currently schema_version 1.1, and the design notes insist that the schema is defined once and shared between the library and the hooks to prevent the two from writing inconsistent state.

## Both layers rated fully reliable depend on you typing a slash command first

The three layers are described with reliability figures attached, and the figures are where a reader should slow down.

Layer one is skills: hardcoded, structured tasks, described as 100 percent reliable, with a built-in record, finalize and save flow and session-level auditing through the start and end commands. Layer two is an output format specification for unstructured tasks, described as about 90 percent reliable, where the project's instruction file requires every reply to contain a marked audit block. The stated principle is that a format constraint beats a behavioural instruction. Layer three is hooks, described as about 100 percent reliable, black-box interception: tool calls recorded automatically on PostToolUse, buffer and pending records archived on Stop, and the session id injected on UserPromptSubmit.

No method is given for any of those percentages. No sample, no definition of reliable, no way to reproduce them.

More interesting is the coupling between the layers. All the audit data is written under the project's .claude/runs/ directory, and the layout includes a file called .current_session, described as the handshake protocol between the start command and the hooks. So the hooks need a session that the skill creates. Skip /start and there is no session id to inject, and the PostToolUse hook has no session to write the tool calls into.

That means the two layers described as fully reliable both no-op silently when a session was never opened, and nothing in the documented flow reports a missing session. A run that never typed /start produces hooks that fire on every tool call and an audit trail that is empty, and the two look identical from the outside to a run that simply had nothing worth recording.

The per-session layout does have integrity machinery, which is the strongest part of the design: a session file, an archived audit trail, a manifest describing a structured batch, an anomalies file and a checksum file. The end command runs a completeness check before saving. So when auditing did happen, there is a way to tell whether the trail was tampered with.

## In-session recovery depends on the agent choosing to ask for it

There are two recovery paths, and they have very different reliability.

The reliable one is a new session. Starting a session automatically loads the index file and the most recent daily report, so opening /start is enough to bring back the state of previous work without the agent deciding anything.

The other path is compaction inside a live session. There, the mechanism is that the agent actively calls /recover. Nothing automatic fires. The framework's own first stated problem is that context windows get compressed or truncated, and its second stated problem is that this causes the agent to repeat mistakes it already made.

So the recovery path for the exact failure the project exists to address depends on the agent, in the same session, choosing to invoke a command, at the moment its context has just been degraded. That is the least reliable moment to ask anything of a model, and the framework does not document a hook that would fire on compaction. The third hook is bound to prompt submission, not to context compaction.

What recovery restores is well specified, and the priority order is the interesting design decision. User corrections come first, on the stated ground that losing them causes the agent to repeat the same mistake. Then task status, meaning where the work had got to. Then analysis conclusions, meaning what had been worked out. Last, environment configuration, meaning which version of the rules and which model were in play.

Ordering by consequence rather than by recency is the right call, and it means a compaction that destroyed most of the window can still be recovered in the part that matters. The audit trail serves as the external store that replaces the context window, which is the framing the project uses throughout: correct samples are kept too, so past conclusions can be re-evaluated later rather than merely trusted.

What is missing from that framing is any statement about how often the trail is read back automatically. The `/start` path loads the index and the latest daily report; whether the recovery path also pulls user corrections first or merely restores whatever the model asks for is not stated.

## Three record types, a ten-thousand-fold frequency spread, and one compaction type

The framework tracks three kinds of record, and they can all exist in one task. What distinguishes them is granularity and, more usefully, volume.

A data record is the processing chain for each row of business data. Row-level, and put at more than 10,000 per day. A change record covers modifications to code, configuration and prompts, at change level, put at 0 to 5 per day. A conversation record is each turn of dialogue between a person and the agent, at interaction level, put at 10 to 50 per day.

Those numbers span four orders of magnitude. A project with 10,000 data records a day and three change records produces an index dominated entirely by the first kind, and the change records, which are the ones that describe what a human or an agent actually altered, are the rarest thing in the store.

That imbalance is presumably why the core module defines a CompactRecord type, and one of the four stated design principles is that correct samples are stored compressed rather than dropped. But the document does not say what triggers compaction, what the retention threshold is, or whether compaction is per record type or applied uniformly. A uniform policy would compress the conversational records into irrelevance to keep the data records, which is the opposite of what the recovery priority order wants.

Two smaller details in the same area. The buffer written by the post-tool hook is cleared on every turn by the stop hook, so it is per-turn scratch space, and the archived trail is the durable copy. And a pending file exists for blocks the agent writes itself, which is how the layer-two format requirement reaches disk rather than staying in the transcript.

The daily output directory is separate again, holding what the daily report command generates. So the store is really four things, and the index is the only one with a documented schema version.

## There is no test directory and no CI directory in the tree

The repository listing is short: .claude-plugin/, .gitignore, CHANGELOG.md, CONTRIBUTING.md, LICENSE, README.md, README_zh.md, docs/, hooks/, install.sh, lib/, skills/ and templates/.

Two things are worth noting in that list by their absence. There is no tests directory and no .github directory, so there is no visible test suite and no visible continuous integration.

For most projects that would be unremarkable. For this one the omission lands in a specific place. The framework's own argument is that agents skip audit records because a written instruction is not binding, so it substitutes three mechanisms including checksums and a completeness check on save. Those are integrity mechanisms, and integrity mechanisms are exactly the kind of thing that regresses quietly: a change to the index schema or to the hook write path would not fail loudly, it would produce a subtly wrong trail.

A changelog, a contribution guide and a docs directory are all present, so the project is not unmanaged. The gap is narrower than that, and it is the gap between shipping a checksum and having something that checks it in an automated run.

Two more items in the tree. A .claude-plugin directory means the bundle can be installed as a Claude Code plugin rather than only by running install.sh, which is a second distribution path for the same three hooks and four skills. And the hooks live in a subdirectory named after the project rather than in the conventional hooks location, which means the installer also has to edit whatever configuration file points Claude Code at them.

Finally, provenance. The licence is MIT, there are no GitHub releases at all, and the last push to main was 2026-06-10, which is 114 days before 2026-10-02. So there is no tag to pin an install to, only a branch.

## Conclusion

Adopt audit-harness if you run one agent across many projects and want a single consistent audit trail with checksums, since the global install is the feature. Do not adopt it expecting project isolation, because the hooks and skills land in your user-level directory and fire everywhere, including projects you never audited. Verify first what your global instruction file looks like before running the installer, since it appends an audit section to it with no stated backup, and read .claude/audit_config.py before leaving it in a repository, because it is imported as code.

## FAQ

### What does audit-harness actually install?

On first run it installs globally into ~/.claude/, placing the core code, four skills, three hooks and a global CLAUDE.md. It then creates .claude/audit_config.py and .claude/runs/ in the current project and injects the audit specification into that project's CLAUDE.md. Running it again in another project recognises the global install and only performs project initialisation.

### How does the three-layer enforcement work?

Layer one is hardcoded skills for structured tasks, with a built-in record, finalize and save flow driven by the start and end commands. Layer two is a required output format for unstructured tasks, where the project instruction file requires every reply to contain a marked audit block. Layer three is hooks that intercept tool use automatically, recording Write, Edit and Bash operations, archiving records on each turn, and injecting the session id on prompt submission.

### Where does audit-harness store its audit records?

Everything goes under .claude/runs/ in the project. That holds an index file carrying schema_version 1.1, a per-turn buffer cleared by the stop hook, a pending file for blocks the agent writes, a .current_session file used as the handshake between the start command and the hooks, and per-session directories holding session.json, the archived audit_trail.jsonl, a manifest, anomalies.json and checksums.json, plus a separate daily output directory.

### How do I configure which files audit-harness tracks?

By editing the generated .claude/audit_config.py, which is Python rather than a data file, and which is looked up first under the project's .claude/ directory and then in the project root, falling back to empty module defaults. It lists CORE_SCRIPTS, CORE_ASSETS, PROMPT_TEMPLATE_GLOB and ALERT_RULES; core scripts and assets feed a SHA256 environment snapshot, and no alert rules are set by default.

### Does audit-harness recover context when the agent loses it?

Two paths, with different reliability. A new session restores state automatically, since /start loads the index and the most recent daily report. Inside a live session the agent has to call /recover itself, and the restoration is ordered by consequence: user corrections first, then task status, then analysis conclusions, then environment configuration.

## Sources

- [Issues](https://github.com/RickyTong1/audit-harness/issues)
- [License: MIT](https://github.com/RickyTong1/audit-harness/blob/main/LICENSE)
- [README](https://github.com/RickyTong1/audit-harness/blob/main/README.md)
- [RickyTong1/audit-harness on GitHub](https://github.com/RickyTong1/audit-harness)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/rickytong1-audit-harness
