CLI tool
perplexityai/numbat avatar
perplexityai/numbat

perplexityai/numbat: endpoint visibility for AI agent activity

Visibility into AI agent activity on endpoints, with on-device detection, optional pre-action blocking, and forensic reconstruction.

1,087 stars114 forksGoApache-2.0

At a glance

What is it?
numbat watches desktop, CLI, IDE and gateway agents through local hooks, OTLP logs and on-disk session artifacts, then evaluates everything with one CEL rule engine. Blocking exists but is off by default and limited to supported synchronous pre-action hooks.
Who is it for?
Adopt numbat if you need local, schema-versioned evidence of what agents did on developer endpoints and you are willing to check the coverage matrix before trusting any host. Skip it if you need guaranteed blocking: enforcement applies only to rules marked enforce: true on supported synchronous pre-action hooks, and every shipped rule is monitor-only, so you must copy a rule, keep its id, add enforce: true and bump its version before anything can be stopped.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Go, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap numbat fills: agents act, but nobody keeps the receipt

An agent running on a developer workstation reads files, calls tools, and makes network requests. Most of that leaves no durable, structured record on the machine itself. numbat targets that gap. It observes supported desktop, CLI, IDE and gateway agents through local hooks and plugins, OTLP/HTTP logs, and on-disk session artifacts, then normalizes live and at-rest activity into one event model.

The audience is narrow and specific: security engineers and platform teams who already have agents running on endpoints they control, and who want evidence rather than a cloud dashboard. Detection runs locally. Records can go to stdout or a local file, and HTTP delivery is optional, so an air-gapped or privacy-sensitive fleet is a supported shape rather than an afterthought.

The design choice worth noticing is that the same CEL rule engine evaluates both live hook events and reconstructed artifacts. That matters because it means a rule you write for real-time monitoring can also run against a session directory you found after the fact.

One event model, three ingestion paths, and a CEL rule engine on top

numbat has three ways in. Hooks and plugins capture activity while an agent runs. An OTLP/HTTP log exporter path accepts agent telemetry. On-disk session artifacts are parsed for forensic reconstruction, which the README frames as working without prior numbat instrumentation. All three converge on a single normalized event, and a CEL rule engine evaluates it.

Rules come in three forms: built-in CEL rules, multi-step sequence rules, and custom YAML rules. Sequence rules are the interesting case. The README shows a finding for the rule chain.secret_read_then_egress, version 1.4, with severity high, citing two event ids: a secret-file access followed by a proposed upload. The finding carries attack tags including attack.t1048, attack.t1552 and attack.t1567, and the observed command is preserved in the record.

Output is versioned NDJSON. Events, findings, enforcement decisions, indicators and scan summaries are separate record types, and JSON Schemas under docs/schema/v0.3.0/ define the wire format. The example records carry schema_version 0.3.0. Events and findings retain source references, so a finding points back at the events that produced it through cited_event_ids and evidence_refs.

Two constraints in the data model are deliberate and worth stating plainly. First, the README says a sequence finding does not prove either action completed. A finding is a claim about observed intent or attempt, not a completed effect. Second, normal record output never includes a complete raw transcript, and adding raw evidence files to a case bundle is opt-in. You get redaction by default and full fidelity only when you ask.

Installing numbat and running a first read-only scan

Releases are published for macOS, Linux and Windows on amd64 or arm64, each with SHA-256 checksums. The README also documents installing from source with Go 1.26.6 or newer:

bash
go install github.com/perplexityai/numbat/cmd/numbat@latest

If you prefer a checkout, the documented build disables cgo, which is consistent with the single-binary distribution claim:

bash
CGO_ENABLED=0 go build -trimpath -o numbat ./cmd/numbat

The safest first use is inventory, because these commands are read-only and do not install hooks or change agent configuration. Running numbat agents lists the agents it discovered:

bash
numbat agents

Then scan the parser-backed agents it found, or narrow the automatic discovery to one agent:

bash
numbat scan
numbat scan --agent codex

What you should see is NDJSON output. The scan path reads on-disk session artifacts and applies the same rule engine used for live events, so a scan can surface a sequence finding such as the secret-read-then-egress example without any hook ever having been installed. That is the property that makes numbat useful on a machine where agents have already been running unmonitored.

Turning on live monitoring, and why enforcement is a separate decision

Live capture is installed per agent, and the README uses Codex as the concrete example. Hooks start in monitor-only mode:

bash
numbat hook install --agent codex --emit all
numbat hook status --agent codex

The --emit all flag writes events, findings, indicators and applicable enforcement decisions to ~/.numbat/records.ndjson. Before relying on this, read the hook trust note: requirements vary by agent and scope, and for the Codex user hook you must review and trust its current definition in /hooks (CLI) or Settings > Hooks (app), including after changes such as --enforce. Codex hooks installed with --managed are trusted by policy. The README is explicit that hook status verifies configuration, not execution or delivery. That is a real operational caveat: a green status does not mean events are arriving.

Enforcement is opt-in twice over. Enforce mode is disabled by default and applies only to rules marked enforce: true, and all shipped rules are monitor-only. To enforce anything you copy the complete shipped YAML from rules/ into a controlled operator directory, keep the same id, add enforce: true, and bump its version:

bash
numbat rules check --rules-dir ./numbat-policy
numbat hook install --agent codex --emit all \
  --rules-dir ./numbat-policy --enforce

This is the right default. A monitoring tool that starts blocking tool calls on day one is a monitoring tool that breaks someone's workflow at 2am. The cost is that blocking is not a feature you get by installing numbat; it is a policy you author, validate, and install against a supported synchronous pre-action hook.

Where numbat is the wrong tool

The README points at docs/agent-coverage.md#matrix as authoritative for each host and surface, and that phrasing is a warning. Coverage is per surface, not universal. If your agent is not in the matrix, or the surface you care about is not a supported synchronous pre-action hook, numbat can observe but cannot stop anything. The README states blocking is limited to supported synchronous pre-action hooks, which means an asynchronous tool call has no enforcement path at all.

Detection quality is bounded by the rules you write. The shipped rules are monitor-only, and the README's own example finding is marked confidence: medium. A medium-confidence finding on a secret read followed by an egress attempt is a prompt to investigate, not an incident.

Reconstruction has its own boundary. It works from supported on-disk session artifacts, so an agent that does not persist sessions in a supported format leaves nothing to reconstruct. And because normal record output never includes a complete raw transcript, an investigation that needs the full conversation must explicitly opt in to adding raw evidence files to a case bundle.

Finally, this is a local tool with local storage. Records land in ~/.numbat/records.ndjson unless you configure HTTP delivery. On a shared workstation, that file is itself something you need to think about.

How numbat differs from a cloud agent-observability platform

The obvious alternative is a hosted observability platform that ingests agent traces and evaluates detections server-side. The difference in approach is where the rule engine runs and what leaves the machine. In a hosted model, telemetry is shipped out and policy is evaluated centrally; the endpoint is a collector.

numbat inverts that. Detection runs locally, the rule engine is CEL and the rules are files on disk (built-in CEL rules, sequence rules, and custom YAML in a rules directory you control), and HTTP delivery is optional rather than assumed. That inversion buys two things: rules can be evaluated against artifacts that were never instrumented, and an endpoint can be monitored without a network path to a backend.

It costs the things central platforms are good at. There is no team-wide correlation across endpoints described in the README, no shared dashboard, and no central place where a rule change propagates to every host. You distribute policy yourself, which is why the README's enforcement workflow is written as copy the YAML, keep the id, bump the version, then rules check and hook install with --rules-dir. If your requirement is fleet-wide detection engineering with a shared rule console, this is a different architecture than what you want. If your requirement is local evidence with a schema you can validate, numbat is aimed at exactly that.

Maintenance, licence and upgrade surface

The repository is not archived, and the last push was on 2026-09-09. Releases are frequent for a project at this stage: v0.1.1 on 2026-07-29, v0.1.2 on 2026-08-01, and v0.2.0 on 2026-08-17. The project is licensed Apache-2.0, and the repository carries a THIRD_PARTY_LICENSES.txt, which is the file to read before redistributing a build.

Upgrade cost is dominated by the schema, not the binary. Records are versioned NDJSON with JSON Schemas under docs/schema/v0.3.0/, and the example records declare schema_version 0.3.0. Any downstream pipeline that parses those records is coupled to that version, so a schema bump is the event that forces work on your side. The binary itself is a single static build per platform, so replacing it is cheap.

The other recurring cost is rule versioning. The enforcement workflow requires keeping the same rule id and bumping the version when you fork a shipped rule to add enforce: true. That is a deliberate discipline: it keeps your local policy traceable to the upstream rule it came from, and it means every enforcement change is a versioned change rather than an edit in place.

On licensing, Apache-2.0 permits commercial use and modification with the usual notice and attribution conditions, but this is a summary of the identifier in the repository, not legal advice for your deployment.

Editorial conclusion

Adopt numbat if you need local, schema-versioned evidence of what agents did on developer endpoints and you are willing to check the coverage matrix before trusting any host. Skip it if you need guaranteed blocking: enforcement applies only to rules marked enforce: true on supported synchronous pre-action hooks, and every shipped rule is monitor-only, so you must copy a rule, keep its id, add enforce: true and bump its version before anything can be stopped. Verify first that your agent appears in docs/agent-coverage.md and that you can review and trust the hook definition in /hooks or Settings > Hooks, because hook status confirms configuration, not execution or delivery.

Frequently asked questions

What is perplexityai/numbat?

It is a Go tool that gives endpoint visibility into AI agent activity, with local detection, optional pre-action blocking, and forensic reconstruction. It observes supported desktop, CLI, IDE and gateway agents through local hooks and plugins, OTLP/HTTP logs, and on-disk session artifacts.

How do I install numbat?

Download a release for macOS, Linux or Windows on amd64 or arm64, each with SHA-256 checksums, or install with Go 1.26.6 or newer using go install github.com/perplexityai/numbat/cmd/numbat@latest. The README also documents a static build with CGO_ENABLED=0.

Does numbat block agent actions by default?

No. Enforce mode is disabled by default and applies only to rules marked enforce: true, and all shipped rules are monitor-only. To enforce a detection you copy its shipped YAML, keep the same id, add enforce: true, bump its version, then run rules check and hook install with --rules-dir and --enforce.

Which agents does numbat support?

The README states that docs/agent-coverage.md#matrix is authoritative for each host and surface, and it names Codex and Claude Code in its examples. Blocking in particular is limited to supported synchronous pre-action hooks, so coverage must be checked per agent and per surface.

Does numbat store full agent transcripts?

Normal record output never includes a complete raw transcript, and adding raw evidence files to a case bundle is opt-in. Artifact scanning is read-only and applies secret redaction.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. perplexityai/numbat on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/perplexityai-numbat.svg)](https://hysenlabs.com/projects/perplexityai-numbat)