TraceRoot: an open source observability layer that turns agent failures into eval datasets
TraceRoot - open-source observability and self-improving layer for AI agents. YC S25
At a glance
- What is it?
- TraceRoot watches production traces from AI agents, runs LLM-as-judge detectors over them, and connects confirmed failures to a sandbox with your source code so it can open fix PRs. It is aimed at teams already running agents in production, not at people prototyping a first chain.
- Who is it for?
- Adopt TraceRoot if you already run agents in production and want failures to become eval datasets instead of tickets. Skip it if you have no production traffic to screen, if you cannot give the debugging sandbox access to your source and GitHub history, or if you need a stable self-hosting path today: the README marks the Terraform and Helm deployment as experimental.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem TraceRoot targets: trace volume outgrows manual review
The README states the case plainly: traces alone do not scale, and manually sifting through every trace is unsustainable as agent systems grow. That is a real cost for anyone running multi-step agents, where a single user request can produce dozens of spans across model calls, tool invocations and retries. The project's answer is a Detectors feature that acts as an LLM-as-judge evaluator over incoming traces, looking for hallucinations, tool and logic failures, safety violations and intent drift, then surfacing findings and triggering root cause analysis with email and Slack alerts.
The intended user is a team that already has production traffic. Nothing in the README describes a workflow for local experimentation or for a single developer tracing a script on their laptop. The value proposition depends on volume: detectors screen traffic, findings become datasets, datasets verify fixes. With three requests a day, that loop has nothing to chew on.
It is worth being precise about what is being sold here. This is not a tracing library with a dashboard bolted on. The tracing layer is described as OpenTelemetry-compatible and is the input, not the product. The product is the loop that runs on top of the traces.
How the detection and root-cause loop actually fits together
The repository layout tells you the shape of the system before any documentation does. There is a backend/ directory holding Python services (the pyproject.toml packages list backend/db, backend/rest, backend/worker and tmux_tools), a frontend/ directory, a separate ee/ directory for enterprise code, and deploy/ for the Terraform path. The Python package declares FastAPI and uvicorn for the REST API, Celery with Redis for the task queue, clickhouse-connect and boto3 for storage, and opentelemetry-proto alongside opentelemetry-api.
That dependency list is the architecture in miniature. Traces arrive over OTLP and are parsed with the OpenTelemetry protobuf definitions. PostgreSQL holds relational data (the compose file notes that the schema is managed by Prisma migrations under ui/), ClickHouse holds the trace volume, and MinIO or S3 holds objects. Celery workers pick up the heavier jobs, which in this system means running the judge models and the agentic debugging pass. The frontend is Next.js, which is why .env.example warns against using generic HOST or PORT variables: Next.js also reads PORT and would start on the wrong port.
Root cause analysis is the part that reaches outside the observability boundary. The README describes an AI that connects to a sandbox with your production source code, identifies the failing line, and correlates the failure with GitHub commits, pull requests and issues, then opens a fix PR. BYOK support covers OpenAI, Anthropic, Gemini, xAI, DeepSeek, OpenRouter, Kimi and GLM. The design choice worth noting is that the debugging agent gets a copy of your code in a sandbox rather than reading it through an API. That is what makes line-level attribution possible, and it is also the part of the system with the largest security surface, which is presumably why the repository ships both SECURITY.md and THREAT_MODEL.md.
Installing TraceRoot locally and sending a first trace
The README gives three paths. Cloud is described as the fastest way to get started, with storage and LLM tokens for testing and no credit card, through app.traceroot.ai. Self-hosting has two local modes. Developer mode hosts the infrastructure in Docker and runs the app itself locally, and it is the path the README recommends for contributors. Clone the repository and run the dev target:
git clone https://github.com/traceroot-ai/traceroot.git
cd traceroot
make devThe Makefile shows that this target first installs pre-commit hooks, then runs uv run python tmux_tools/launcher.py. That launcher handles dependencies, infrastructure, migrations and a tmux session, and the Makefile notes it is idempotent and will reattach if already running. You should end up with the app reachable at http://localhost:3000. If you are on Windows and have no tmux, the Makefile provides dev-lite, which runs the launcher with --env-only, installs frontend dependencies with pnpm, and then brings up the production compose file.
If you only want to test rather than contribute, local docker mode keeps everything in containers:
git clone https://github.com/traceroot-ai/traceroot.git
cd traceroot
make prodBefore any of this, copy the environment template. The .env.example file states that this single file is used by all services (Python backend, Next.js frontend, TS worker) and that no other .env files are needed. Values marked CHANGEME in the file and in docker-compose.yml are placeholders and should be replaced.
cp .env.example .envInstrumentation is separate from the platform. The README points to traceroot-py and traceroot-ts as the native SDKs, and to per-framework integration pages for Agno, AutoGen, the Claude Agent SDK, CrewAI, DSPy, Google ADK, LangChain and LangGraph, LangChain DeepAgents, LlamaIndex, the Microsoft Agent Framework and Mastra. LangChain and DeepAgents integrate by passing a callback handler; Mastra integrates through the TraceRoot OTLP exporter. The README does not show a full SDK snippet on the page itself, so the integration docs are where the actual constructor arguments live.
Datasets and evals are the part that outlives the incident
Most observability tools stop at the point where a human reads a trace and fixes something. TraceRoot's stated differentiator is that the finding is not the end of the process. The README describes turning production findings into golden datasets with one click, then running offline evals from the TraceRoot CLI or SDK inside coding agents such as Claude Code, Codex and Cursor, so that every fix is verified against the dataset.
The interesting detail is where the eval runs. Putting the eval runner inside a coding agent means the verification step happens in the same session as the fix, rather than in a separate CI job that someone has to wire up. That is a genuine workflow choice, and it also means the quality of your verification depends on the coding agent you happen to use. The README lists three by name and does not describe a generic shell entry point for the CLI, so teams using a different harness will need to check the docs before assuming parity.
There is a second-order effect worth flagging. A golden dataset built from production findings inherits the distribution of your production traffic. If the detectors only fire on the failure classes they were configured for, the dataset will only cover those classes. The README does not describe how datasets are deduplicated or how they age as the agent's behaviour changes, and those are the questions that determine whether the eval suite stays useful six months in.
Where TraceRoot is the wrong tool, and what it costs to run
The clearest boundary is deployment maturity. The README describes the Terraform and Helm path under deploy/ as being for production hosting and still in an experimental stage. A team that needs a supported, stable self-hosted deployment today is looking at two Docker-based local modes and an experimental cloud path, and should treat that as a real constraint rather than a footnote.
Resource footprint is the second constraint. The stack in docker-compose.yml includes PostgreSQL, ClickHouse, MinIO and Redis before you add the application services. That is a lot of infrastructure for a tracing tool, and it reflects the design decision to store trace volume in ClickHouse rather than in the relational database. For a small team, running that locally means a laptop with meaningful headroom. For a team that just wants spans in a dashboard, this is more machinery than the problem requires.
The agentic debugging feature has its own boundary. It needs a sandbox with your production source code and access to GitHub history to correlate failures with commits and pull requests. Teams that cannot or will not give a hosted service that access should either self-host and keep the sandbox inside their own network, or treat the root-cause feature as unavailable. The repository's inclusion of THREAT_MODEL.md suggests the maintainers have thought about this, but the README does not describe the sandbox's isolation model, so that is a question for the security documentation rather than the front page.
Finally, the licence. The repository metadata reports NOASSERTION, and there is an ee/ directory alongside the main source. That combination usually means the licence file does not match a standard SPDX identifier and that some code sits under different terms. The README's badge links to a LICENSE file, but the terms themselves need to be read directly before adoption.
How TraceRoot differs from an OpenTelemetry backend like Langfuse or Phoenix
The obvious comparison is with the open source LLM observability tools that also ingest OpenTelemetry traces, such as Langfuse or Arize Phoenix. The difference in approach is where the pipeline terminates. Those tools are primarily storage, search and visualisation layers for traces, with evaluation features that a team configures and runs. TraceRoot adds a second stage that acts on the traces: judge models screen them, an agent reads your source code to attribute the failure to a line, and the system opens a pull request.
That is a meaningful difference in both directions. It means TraceRoot has an opinion about your repository layout and your GitHub access, which a pure tracing backend does not. It also means the quality of the output depends on the judge models and the debugging agent, which are configurable through BYOK. If you already run Langfuse for trace storage and are happy reading traces yourself, TraceRoot is not a drop-in replacement, it is an additional layer with a different job.
The honest framing is that these tools are complementary more often than they are competitors. TraceRoot's own tracing layer is OpenTelemetry-compatible, which means it can sit alongside or downstream of existing instrumentation. The decision is not which backend stores your spans. It is whether you want a system that proposes fixes, with the access requirements that implies.
Frequently asked questions about adopting TraceRoot
The questions below cover what a new user is most likely to ask before installing. Answers are limited to what the README, the repository files and the release notes state.
Editorial conclusion
Adopt TraceRoot if you already run agents in production and want failures to become eval datasets instead of tickets. Skip it if you have no production traffic to screen, if you cannot give the debugging sandbox access to your source and GitHub history, or if you need a stable self-hosting path today: the README marks the Terraform and Helm deployment as experimental. Before committing, verify that your framework appears in the integrations table, that your Python and Node versions satisfy the repository requirements, and that the licence file matches what your legal review needs, because the repository metadata reports NOASSERTION rather than a named licence.
Frequently asked questions
What is tracing in AI agents?
In TraceRoot's terms, tracing means capturing LLM calls, agent actions and tool usage through an OpenTelemetry-compatible SDK. Those traces become the input that the Detectors screen for hallucinations, tool failures, logic errors and safety violations.
How do I install TraceRoot locally?
Clone the repository and run make dev for developer mode, which handles dependencies, infrastructure, migrations and a tmux session, or make prod to run everything in Docker. The README also lists a cloud option at app.traceroot.ai.
Which model providers does TraceRoot support?
The README states BYOK support for any model provider and names OpenAI, Anthropic, Gemini, xAI, DeepSeek, OpenRouter, Kimi and GLM.
Which agent frameworks can I instrument with TraceRoot?
The README lists Agno, AutoGen, the Claude Agent SDK, CrewAI, DSPy, Google ADK, LangChain and LangGraph, LangChain DeepAgents, LlamaIndex, the Microsoft Agent Framework and Mastra, with Python and JS/TS coverage varying by integration.
Can I self-host TraceRoot in production?
The README points to a Terraform and Helm path under deploy/ for running on k8s, but describes it as still in an experimental stage. The two Docker-based local modes are documented for development and testing rather than as a supported production deployment.
What licence is TraceRoot released under?
The repository metadata reports NOASSERTION rather than a named licence, and the repository contains an ee/ directory alongside the main source. The terms in the LICENSE file need to be read directly, since the metadata alone does not identify them.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/traceroot-ai-traceroot)