AutoHarness: a governance wrapper for AI agents that runs on a 6-step pipeline
AutoHarness: Automated Harness Engineering for AI Agents
At a glance
- What is it?
- AutoHarness is an MIT-licensed Python middleware that wraps an existing LLM client and inserts a governance pipeline between the model and the tools it calls. It is aimed at teams whose agents work but are not yet safe to leave unsupervised.
- Who is it for?
- Adopt AutoHarness if you already have a working agent loop and want permission checks, risk classification and an audit trail without rewriting the loop. Do not adopt it if you expect it to improve model reasoning: the README is explicit that the model reasons and the harness does everything else.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 180 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap AutoHarness targets: agents that demo well and fail in production
The README frames the problem in one line: the gap between an agent that is demo-ready and one that is reliable is enormous, and the missing pieces are context management, tool governance, cost control, observability and session persistence. AutoHarness is the project's answer to that list, packaged as what pyproject.toml calls "behavioral governance middleware for AI agents".
The intended user is a developer who already has an agent loop calling tools and wants a policy layer around those calls, not a new agent framework. The README's own equation is Agent = Model + Harness, with the model doing the reasoning and the harness doing everything else. That is a narrow claim, and it is worth taking literally: nothing in the repository suggests AutoHarness improves what the model decides, only what is allowed to happen after it decides.
The framing also explains the package name. "Harness engineering" here means the surrounding machinery (budgets, permissions, logging, profiles), not prompt engineering and not fine-tuning.
Inside the 6-step governance pipeline
The README gives the pipeline as a six-stage flow: Parse & Validate, Risk Classify, Permission Check, Execute, Output Sanitize, Audit Log. Every tool call passes through it. The Core mode uses exactly this sequence; Standard extends it to eight steps and Enhanced to fourteen, adding a turn governor, alias resolution and failure hooks among other stages.
Risk classification is pattern-based rather than model-based. The README says built-in risk patterns detect dangerous operations, secret exposure and path traversal. That is a design decision with consequences: pattern matching is deterministic, cheap and auditable, but it will not catch a novel dangerous operation that does not resemble a known pattern. The README does not describe a way to plug in a classifier model, so treat the risk layer as a denylist with structure rather than a semantic guard.
Policy lives in a YAML constitution file. The README shows a constitution controlling the pipeline mode, and the CLI exposes autoharness validate constitution.yaml to check a file before use. Audit output is JSONL, described as logging every decision with full provenance, and cost attribution is per call with model-aware pricing. The repository also ships examples/cedar_policy.cedar and examples/opa_policy.rego, which indicates the project expects policy to be expressed in external policy languages in at least some deployments. The README does not document how those two files are loaded, so read the examples before assuming they are wired in.
Installing AutoHarness and wrapping a client
The README's quick install clones the repository and installs it in editable mode. Python 3.10 or newer is required according to pyproject.toml, and the runtime dependencies are pydantic, pyyaml, click and rich. Provider SDKs are optional extras: anthropic, openai, langchain and all.
git clone https://github.com/aiming-lab/AutoHarness.git
cd AutoHarness && pip install -e .The first real use is the wrap call. It takes an existing client and returns one that routes through the harness, so the surrounding code keeps calling chat.completions.create as before.
from openai import OpenAI
from autoharness import AutoHarness
client = AutoHarness.wrap(OpenAI())
response = client.chat.completions.create(
model="gpt-5.4",
messages=[{"role": "user", "content": "Refactor auth.py"}],
)If you want the full loop instead of a wrapped client, the README shows AgentLoop taking a model name and a constitution path, then result = loop.run("Fix the failing tests in auth.py"). The CLI has an interactive setup path, autoharness init, which asks for agent type, LLM provider, security level and pipeline mode. Running autoharness mode prints the current mode, and autoharness mode enhanced switches it. The README states that Enhanced is the default, so a fresh install starts at the heaviest setting unless you change it.
Where AutoHarness is the wrong tool
The most direct limitation is stated in the README itself: Enhanced is the default mode, and the project recommends switching to Core for minimal overhead. Fourteen pipeline steps per tool call is a real cost in latency and complexity, and the documentation does not quantify it. If your agent makes many cheap tool calls, the default configuration is likely the wrong starting point.
Second, the governance is pattern-based. An agent doing legitimate but unusual work (a build script that rewrites files across a tree, a migration that touches many paths) may trip path traversal or dangerous-operation patterns. The README does not describe a tuning workflow for false positives beyond editing the constitution, so budget time for that.
Third, AutoHarness is not a sandbox. It intercepts tool calls that flow through the wrapped client or AgentLoop. If your agent can reach a shell, a database or a network by another path, the pipeline does not see it. The README's own example is an agent running rm -rf / and the pipeline blocking it, which presumes the command goes through the governed tool interface.
Finally, the project is young. pyproject.toml classifies it as Development Status 3 - Alpha, and the only tagged release is v0.1.0. The last push to the repository was on 2026-04-02, so the code you clone today is close to that release but not necessarily identical to it.
AutoHarness against LangGraph and Guardrails AI
The README's comparison table puts AutoHarness next to LangGraph, Guardrails AI and the OpenAI SDK. The differences it claims are structural rather than cosmetic. LangGraph is a graph-based orchestration framework: you describe the agent as a graph of nodes and edges, and governance is something you build into the graph. AutoHarness takes the opposite position, leaving the loop where it is and inserting a pipeline between the model and its tools. If you already have an orchestration layer you like, that is an argument for AutoHarness; if you are starting from nothing and want the graph model, it is an argument against.
Guardrails AI is described in the table as output-only validation. AutoHarness claims validation on both input and output rails, plus execution-time checks in the middle. That is a genuine difference in coverage: output validation cannot stop a destructive command from running, only flag the response afterwards. The table is the project's own summary, so treat the column values as claims to verify rather than measured results.
The OpenAI SDK column is the baseline: trimming for context and handoff for multi-agent, with no governance pipeline. AutoHarness positions itself as what you add on top of that baseline, and its wrap() call is designed to make that additive step small.
Maintenance, licence and upgrade cost
AutoHarness is MIT licensed, with the LICENSE file at the repository root and the license field set to MIT in pyproject.toml. MIT is permissive: you can use it commercially, modify it and redistribute it, provided the copyright notice and permission notice are retained. That is the standard reading; if your organisation has specific compliance requirements, have counsel review the actual LICENSE text rather than relying on the SPDX identifier.
The upgrade picture is short. The only tagged release is v0.1.0, dated 2026-04-02, and pyproject.toml already declares version 0.1.1, so the main branch is ahead of the last tag. The README advertises 958 passing tests in a badge, and the Makefile provides make test, make lint, make typecheck and make dev targets, which means you can run the suite yourself before pinning. There is no documented migration guide, no changelog beyond the v0.1.0 news entry, and no stated backward-compatibility policy. The README does not document rollback. Pin to a commit and read the diff before moving.
The last push was on 2026-04-02, roughly five and a half months before this writing. That is recent enough that the project is not abandoned, but there is only one release to judge stability by.
Editorial conclusion
Adopt AutoHarness if you already have a working agent loop and want permission checks, risk classification and an audit trail without rewriting the loop. Do not adopt it if you expect it to improve model reasoning: the README is explicit that the model reasons and the harness does everything else. Before committing, read the pipeline-mode table and decide whether Enhanced is tolerable, run autoharness validate against your own constitution.yaml, and check that the pyproject.toml version (0.1.1) matches the release you intend to pin, since the only tagged release is v0.1.0.
Frequently asked questions
What is harness engineering for language agents?
AutoHarness uses the term for the machinery around the model: context management, tool governance, cost control, observability and session persistence. The README's equation is Agent = Model + Harness, with the model reasoning and the harness handling the rest.
What does harness mean in AutoHarness?
It refers to the governance layer that wraps an LLM client, not to a test harness. In this project the harness is the pipeline that parses, classifies, checks permissions, executes, sanitizes output and writes an audit record for each tool call.
What is auto harness?
AutoHarness is an MIT-licensed Python package, described in pyproject.toml as behavioral governance middleware for AI agents. Its quickstart wraps an existing OpenAI client with AutoHarness.wrap(OpenAI()) so tool calls pass through a governance pipeline.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/aiming-lab-autoharness)