# FailproofAI review: observability and enforcement for AI agent harnesses

> FailproofAI hooks twelve agent harnesses, captures every run, and can block a tool call before it executes. Here is what the repository documents, where it stops, and who should not install it.

**FailproofAI/failproofai** — Observability and enforcement for AI agent harnesses. Capture every run and runtime reliability with policy enforcement.  40 built-in policies, a local dashboard, no account required with a generous free cloud plan

- Repository: https://github.com/FailproofAI/failproofai
- Website: https://befailproof.ai
- Stars: 5,218 · Forks: 501
- Language: MDX
- License: NOASSERTION
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/failproofai-failproofai

## The problem FailproofAI targets: agents that act before anyone looks

Coding agents and chat gateways now run shell commands, edit files and call tools on their own. The failure mode is not a crash. It is a run that finishes quietly after doing something nobody wanted, with no record of which policy, if any, was in force. FailproofAI is aimed at that gap. The README describes it as observability and enforcement for every harness your agents run in, and the package description says it hooks twelve of them, blocking the tool call before it runs.

The audience is narrow and specific. You need this if you already run agents through one of the twelve harnesses listed in the README: Claude Code, OpenAI Codex, GitHub Copilot CLI, Cursor Agent CLI, OpenCode, Pi, Hermes, OpenClaw, Factory Droid, Devin CLI, Antigravity CLI or Goose. If your agents live somewhere else, the project still offers something, but not the same thing, and the README is explicit about the difference. It is not a general-purpose LLM tracing product and it does not try to be a model gateway.

## How the hook model works across twelve harnesses

The design bets on a shared event shape. The README states that all twelve harnesses produce the same events, the same policies and the same session history, whichever one your agent runs in. That is the whole architectural claim: normalize the harness, then write policy once.

The repository layout backs this up. There are per-harness directories at the top level (.claude/, .codex/, .cursor/, .devin/, .factory/, .opencode/, .pi/), a pi-extension/ directory, an openclaw-plugin/ directory, and a docker-hook-sync/ directory. Alongside those sit a Next.js application in app/, a Rust workspace declared in Cargo.toml with members under crates/, and a separate sdk/ tree. The npm package ships a CLI binary named failproofai and a daemon shim named failproofaid, which suggests a long-running process that harness hooks talk to rather than a library you import per request.

Where the documentation gets thin is the wire format. The README says the events are the same; it does not publish the event schema or the policy evaluation order. If you plan to write custom policies, that is the first thing to check in the docs site rather than assume.

## Installing FailproofAI and running your first enforced session

The README's install section begins with an npm command and is truncated at that point, so the exact install line is not reproducible from the repository files. What the repository does confirm is the package name, `failproofai`, published on npm, and the engine requirements in package.json: Node `>=20.9.0` and Bun `>=1.3.0`. Check both before you start.

```bash
npm install -g failproofai
```

The package exposes two binaries. `failproofai` is the CLI entry point, and `failproofaid` is a daemon shim. The repository's own dev script runs the server on port 8020, which is the port the local dashboard is served from during development.

```bash
FAILPROOFAI_TELEMETRY_DISABLED=1 bun scripts/dev.ts --port 8020
```

That environment variable appears in the repository's `dev` and `start` scripts, so it is the documented way to turn telemetry off. If you are evaluating the tool on a machine that should not phone home, set it before the first run rather than after.

For policy authoring, the repository ships worked examples under `examples/`: `examples/policies-basic.js`, `examples/policies-advanced/`, `examples/policies-stop.js`, `examples/policies-notification.js` and `examples/convention-policies/`. The names tell you what each covers. `policies-stop.js` is the one to read first if your interest is blocking rather than logging, since stopping a call is the behaviour that distinguishes this project from a tracer.

The README claims zero latency and local execution. Treat both as claims to verify against your own harness, because the README gives no measurement methodology and no numbers.

## Where FailproofAI stops: unsupported harnesses and the SDK boundary

The honest limitation is in the README itself. Agents that run in none of the twelve harnesses report through the Python SDK, which the README says gives you tracing, sessions and audits. Enforcement there needs a hook in your own runtime, and the README's answer is to contact support so the maintainers map it. That is not a self-service path. If your agent is a custom loop, you get visibility and you do not get blocking until someone writes an integration for you.

A second boundary is the release line. The most recent releases listed are `v1.0.4-beta.4` and `v1.0.4-beta.2`, with a separate `failproofai-sdk-v0.0.1b2`. The package.json version is `1.0.6-beta.0`. Everything public here is pre-1.0 or beta-tagged, which means policy behaviour and event shapes can move between versions. Pin an exact version and read the changelog before upgrading.

Third, the licence is not a standard identifier. The repository reports `NOASSERTION`, and the README badge reads MIT plus a Commons Clause. The README's own count also drifts: the description says 40 built-in policies, package.json and the README body both say 39. Verify the number in the version you install.

## FailproofAI compared with general LLM tracing tools

The obvious alternative is a tracing platform such as Langfuse or LangSmith: you instrument your application, send spans to a backend, and read dashboards. The difference in approach is where the decision happens. A tracing platform observes after the fact. It can tell you that a tool call was dangerous once the call has returned. FailproofAI positions the hook inside the harness so the call is evaluated before execution, which is why the README frames the product as enforcement rather than monitoring.

That difference costs you portability. A tracing SDK works with any code you control, including a hand-rolled agent loop, because you are the one emitting spans. FailproofAI works with the harnesses it has hooked, and outside them you fall back to a Python SDK that the README describes as tracing-only. So the trade is real in both directions: enforcement inside a supported harness, or universal coverage with no blocking. If blocking is not something you need, a tracing platform is the simpler choice and does not tie you to a harness list.

## Licence, upgrade cost and what the repository does not document

The README badge describes the licence as MIT plus a Commons Clause, while the repository metadata reports `NOASSERTION`. Those two statements are not the same, and a Commons Clause typically adds a restriction on selling the software. Read `LICENSE` in the repository before you build anything commercial on top of it. This is a description of what the files say, not legal advice.

Upgrade cost is the beta cadence. Releases in the listed set land days apart, and the version in package.json leads the newest published release. Each bump can change policy semantics or event fields. The repository has a CHANGELOG.md at the top level, so the cost of upgrading is bounded by reading it, but you should assume a review step per version rather than a silent `npm update`.

What the repository does not document: rollback of a blocked call, the event schema, policy evaluation order, and any latency measurement behind the zero-latency claim. The README also does not state how the local dashboard authenticates, which matters if the daemon binds to a port on a shared machine. Check the docs site for those before you deploy beyond a laptop.

## Conclusion

Adopt FailproofAI if your agents already run inside one of the twelve supported harnesses and you want a local record of every run plus a way to stop a tool call before it executes. Skip it if your agents run in a harness outside that list: the README says tracing and sessions work through the Python SDK, but enforcement needs a hook in your own runtime that you have to arrange with the maintainers. Before rolling it out, verify three things in your own environment: that the Node and Bun versions in package.json meet your machines, that the policies you enable behave as you expect against a non-production project, and that the package version you pin matches the behaviour you reviewed, since the published release line is still beta.

## FAQ

### Which agent harnesses does FailproofAI support?

The README lists twelve: ten coding CLIs and two chat and assistant gateways. The named ones include Claude Code, OpenAI Codex, GitHub Copilot CLI, Cursor Agent CLI, OpenCode, Pi, Hermes, OpenClaw, Factory Droid, Devin CLI, Antigravity CLI and Goose. Agents outside that list report through the Python SDK, where tracing, sessions and audits work but enforcement needs a hook in your own runtime.

### How do I install FailproofAI?

It is published on npm as the package failproofai, and package.json declares the binaries failproofai and failproofaid. The engines field requires Node >=20.9.0 and Bun >=1.3.0. The README's install section is truncated at the npm command, so confirm the exact line on the docs site.

### Does FailproofAI need an account or send data to the cloud?

The package description says a local dashboard is included and no account is needed, with a free cloud plan also offered. The repository's dev and start scripts set FAILPROOFAI_TELEMETRY_DISABLED=1, which is the documented way to turn telemetry off.

### How many built-in policies does FailproofAI ship?

The counts in the repository disagree. The project description says 40 built-in policies, while package.json and the README body both say 39. Check the count in the specific version you install.

## Sources

- [FailproofAI/failproofai on GitHub](https://github.com/FailproofAI/failproofai)
- [Issues](https://github.com/FailproofAI/failproofai/issues)
- [Project website](https://befailproof.ai)
- [README](https://github.com/FailproofAI/failproofai/blob/main/README.md)
- [Releases](https://github.com/FailproofAI/failproofai/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/failproofai-failproofai
