agent-inspect keeps the trace on your laptop, and admits its safety verdict is not a certification
Local evidence debugger and trajectory-test toolkit for TypeScript AI agents: inspect causal runs, catch wrong tool paths in CI, and share safe offline evidence.
At a glance
- What is it?
- A TypeScript toolkit that turns an agent run into a local JSONL trace, reads it as a nested execution tree, gates on it in CI with deterministic checks, and bundles redacted evidence offline. The honesty is in the caveats: trajectory checks refuse to invent outcomes, and the share verdict is explicitly not compliance.
- Who is it for?
- Use agent-inspect if you are writing TypeScript agents whose failures are spread across a dozen steps and you need one artefact that serves debugging, a CI gate and a review bundle. Do not use it as a hosted observability replacement or an evaluation platform, because the project scopes itself to the laptop-to-pull-request loop and says it complements those rather than replacing them.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 2 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
One local JSONL trace serves debugging, gating and sharing
The starting observation is that agent code does not fail in one place. It plans, retrieves, calls tools, invokes a model, retries, and produces side effects, and flat logs show you fragments of that rather than the shape of it. So the trace is kept as local JSONL and everything else is a different reading of the same file. There is no account, no collector, and no default upload, and the run is metadata-only by default. The debugging side reads the execution path: nested steps, tool calls, LLM calls, model and token metadata, durations, errors, and the first causal failure, all without sending the trace to a service:
npx agent-inspect view <run-id> --dir .agent-inspect --summary
npx agent-inspect report <run-id> --dir .agent-inspect
npx agent-inspect explain <run-id> --dir .agent-inspectThree verbs over one artefact, and all of them take the same run id and the same directory. How the trace gets there is deliberately open: manual instrumentation, official adapters, supported standards files, or structured logs you already emit, so you are not forced to rewrite your agent to get value from the tool.
verify-safe is a local judgement, and the page says share-checked is not certification
The sharing path is three commands and one explicit warning. First an assessment, then a bundle, then an offline verification of what was written:
npx agent-inspect verify-safe <run-id> --dir .agent-inspect
npx agent-inspect bundle <run-id> --dir .agent-inspect --profile share --out ./evidence
npx agent-inspect bundle verify ./evidenceThe assessment is described as best-effort and local, and it has four possible answers rather than a boolean: `SAFE`, `SAFE WITH WARNINGS`, `UNSAFE`, or `UNKNOWN`. The bundle step writes a redacted artifact and a manifest, and the verify step rechecks the listed file hashes offline, which means the check does not need a network connection or a service. What the project then says is that share-checked is not compliance certification, and that you should review the artifact before attaching it to a pull request, an incident or a support handoff. That sentence is the most useful thing on the page, because a tool that tells you an artefact is safe to attach is the kind of tool people stop reading. The `UNKNOWN` result exists for the same reason: a tool that can only say yes or no is lying about its own coverage.
Trajectory checks refuse to invent an outcome from a successful tool call
The CI gate is deterministic and provider-free by design: the same trace and the same rules produce the same result, so a failure is reproducible and a pass means something. You start from a preset and then add the expectations that matter to your workflow:
npx agent-inspect check <run-id> --dir .agent-inspect \
--preset trajectory \
--required-tool retrieve_policyThe rule that matters most is a limit on what the checker will do. Use `--fail-on-observation failed` only when the run records explicit OUTCOME events, because trajectory checks do not invent outcomes from tool, LLM or run success. That is a distinction most evaluation tooling blurs: a tool that returned without error is not the same as an observation that succeeded, and treating the first as the second produces a gate that passes while your agent does the wrong thing. When a single CLI check is not enough, the surface widens rather than changing meaning, with TraceContract, suites, cohorts, Vitest and Jest reporters, and an `--evidence-on fail` flag that attaches the evidence only when something actually failed.
Exit codes split a rule failure from a broken configuration
A gate is only useful in a pipeline if a failure means what you think it means, and this project is specific about the distinctions. A passing check exits `0`. A rule failure exits `1`. Invalid configuration and unreadable or unsupported inputs get separate documented exit codes rather than sharing the one. That last part is the one to care about in CI, because a mistyped flag and an unreadable trace file are both errors, and if they share an exit code with a genuine trajectory violation then a red build tells you nothing about which of the three happened. It also means the same command behaves differently depending on whether the file it was handed is a trace in the supported shape, which is a different question from whether the agent behaved correctly. The support levels page names five tiers, Stable, Supported, Beta, Preview and Experimental, and the page states the project's own level as Stable, with network behaviour documented separately, so the two axes, maturity and network access, are kept apart on purpose.
Adapters are separate packages, and plain functions need none at all
Adoption is arranged so that the cheapest case needs no adapter. If your agent is built from custom functions or classes, the base package is enough. Everything else is a sibling package:
npm install agent-inspectLangChain and LangGraph add `@agent-inspect/langchain`, the AI SDK adds `@agent-inspect/ai-sdk`, and OpenAI Agents JS adds `@agent-inspect/openai-agents`. Then there are two paths that need no adapter package at all, which are the interesting ones for an existing codebase: `npx agent-inspect logs` turns structured logs you already emit into a tree, and `npx agent-inspect open` reads OpenInference or OTLP JSON. For test runners there are `@agent-inspect/vitest` and `@agent-inspect/jest` reporters, which is how the trajectory checks end up attached to an existing test suite rather than bolted beside it. The first-run path is also honest about what it does not do: installing the package does not copy the repository examples, `init` writes them into your project, and the generated demo is synthetic, keyless and local.
The programmatic surface is two imports and a contract object
The checks are usable from code rather than only from the command line, and the shape of the API is the shape of the idea: read a trace, declare a contract, evaluate, and use the status. The example in the page is short enough to be the whole story:
import { openTraceFile } from "agent-inspect/readers";
import {
defineTraceContract,
evaluateTraceContractRead,
} from "agent-inspect/checks";
const read = await openTraceFile("./trace.jsonl");
const contract = defineTraceContract({
run: { requireCompleted: true },
tools: { required: ["retrieve_policy"] },
});
const result = evaluateTraceContractRead(read, contract);
if (result.status !== "pass") process.exitCode = 1;The contract declares that the run must have completed and that a named tool must have been used, and the failure path is a process exit code rather than a thrown error, which is what you want inside a test. The full API also exposes TraceFacts for bounded, logical analysis of supported traces, and the word bounded is doing real work: the analysis is defined over a supported subset rather than pretending to reason over arbitrary output. The subpath imports come from the package's export map, which is how the reader and the checks stay separable from the rest of the surface.
The Docker image exists for a directory listing, and pins an older server build
There is a Dockerfile, and it is worth reading because of what it is not. It builds a Glama directory-listing image for the Preview read-only MCP server, on a pinned Alpine Node base, with a dedicated unprivileged user and a single directory at `/traces` exported as `AGENT_INSPECT_TRACE_DIR`. The comment at the top is explicit that the product is local-first and the image only needs to answer initialize and tools/list. So the container exists to make the read-only MCP server appear in a public listing, not to run the product, which is a slightly odd thing to discover and a good example of scope being stated up front. The detail that needs attention is the pin: the image installs `@agent-inspect/[email protected]` while the current package version is 6.31.16. The file says to pin the published package version and to rebuild after each release that should appear in the listing, so the lag is a deliberate maintenance step rather than an accident. It is still a version skew you should confirm before you rely on that image.
Twenty-two build configs, a dual export map, and examples numbered around a gap
The repository root is wide, and the shape of it tells you how many surfaces the project maintains. There are more than twenty `tsup` build configurations, one per package, covering the CLI, core, viewer, studio, TUI, redaction, guardrails, harness, eval, circuit, an SQLite index build, MCP and the MCP server, LangChain, AI SDK, OpenAI Agents, adapter SDK, Vitest, Jest and the VS Code extension, plus separate Vitest and TypeScript configs. Alongside that sit a pnpm workspace with a lockfile, a size-limit configuration for bundle budgets, a `.changeset` directory, Git hooks, an `AGENTS.md`, a roadmap, a security policy, and separate `apps/`, `packages/`, `docs/`, `fixtures/`, `scripts/` and `test/` directories. The published package is ESM first, with `type` set to module and a main, module and types triple, and the export map gives each subpath both an import and a require branch with its own type declarations. The examples directory is numbered and has a gap: 00 quickstart demo, 01 basic, 02 nested steps, 03 parallel steps, 04 error handling, 05 observe wrapper, 06 log to tree, and then 08 for the LangChain adapter, with 07 absent. The release cadence matches the build surface, with three tagged versions published within three hours on the same afternoon.
Editorial conclusion
Use agent-inspect if you are writing TypeScript agents whose failures are spread across a dozen steps and you need one artefact that serves debugging, a CI gate and a review bundle. Do not use it as a hosted observability replacement or an evaluation platform, because the project scopes itself to the laptop-to-pull-request loop and says it complements those rather than replacing them. Before you wire the check into a pipeline, read the exit code table, because a rule failure and a misconfigured rule are different outcomes and only one of them is a bug in your agent. Decide whether your instrumentation records explicit OUTCOME events, since without them the observation-level gate cannot fire. And treat the bundle as a draft for a human to read, not as something safe to attach on autopilot.
Frequently asked questions
What does agent-inspect do with an agent run?
It keeps the run as local JSONL and gives you three readings of it: an execution tree for debugging, deterministic trajectory checks for CI, and a redacted offline evidence bundle. There is no account, no collector and no default upload, and it is metadata-only by default.
Is the agent-inspect safety verdict a compliance certification?
No. verify-safe is a best-effort local assessment that can report SAFE, SAFE WITH WARNINGS, UNSAFE or UNKNOWN, and the project states that share-checked is not compliance certification. Review the artifact before attaching it to a PR, incident or support handoff.
How do I make CI fail when my agent takes the wrong path?
Run the check with a preset and the expectations that matter, for example a trajectory preset with a required tool. A passing check exits 0 and a rule failure exits 1, while invalid configuration and unreadable or unsupported inputs use separate documented exit codes.
Which frameworks does agent-inspect support?
Plain functions and classes need only the base package. LangChain and LangGraph, the AI SDK and OpenAI Agents JS each add a sibling adapter package, and existing structured logs or OpenInference and OTLP JSON can be read without any adapter at all.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/rajudandigam-agent-inspect)