Model or dataset
Jwuthri/Tracely-ai avatar
Jwuthri/Tracely-ai

Tracely-ai: Production Failures as Hermetic Regression Tests for AI Agents

Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0.

1,418 stars156 forksPythonMIT

At a glance

What is it?
Tracely-ai is an open-source trace-native CI/CD platform for AI agents. It grades every incoming production trace, clusters failures into issues, promotes failing runs into hermetic regression cases with recorded fixtures, and blocks the pull request that would ship a regression. Replays in CI use no API keys and no model spend.
Who is it for?
Teams running AI agents in production who want automated regression protection without hand-authoring eval datasets will find Tracely-ai's trace-promotion approach a better fit than tools that ask you to write test cases by hand. The full stack is self-hosted, which means you own the data but also own the infrastructure.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Problem Tracely-ai Solves

Most agent evaluation tools ask an engineer to write a dataset: a set of questions, expected answers, and grading criteria assembled before the agent is deployed. That dataset is a prediction about what might fail. Tracely-ai takes the opposite approach: it treats the production trace of a run that already failed as the test case.

The core claim is that the recorded run is the test. A failing trace captured in production carries the exact input, the exact tool calls, and the exact model responses that produced the failure. Tracely-ai freezes that trace into a hermetic case with fixture-recorded tool and LLM outputs, then replays it against every subsequent pull request. No live model calls are made during replay, so the CI step costs nothing in API spend.

The intended audience is engineering teams running AI agents in production who want failures to block deployments automatically, not just appear on a monitoring dashboard.

The Five-Step Pipeline

Tracely-ai divides its work into five stages that correspond to five pages in the application.

Observe: traces arrive over OTLP/HTTP from any language. Agent-specific fields (agent.id, conversation.id, turn, step) are promoted to indexed columns, so runs group into conversation threads rather than a flat list of spans. Evaluators appear as columns on the trace table and stream their verdicts live over SSE as judges finish. A Replay view walks a conversation event by event, and a Fleet view renders the same conversation as a visual room where each agent is a character and tool calls are objects on the wall.

Detect and Triage: online evaluators grade every run as it lands using LLM-as-judge at conversation, run, or span level, plus structural checks that require no model at all. Failures cluster into issues by structural signature and semantic embedding, so thirty similar broken runs become one issue with a count.

Test: one click promotes a failing trace into a hermetic case. The case bundles recorded tool and LLM outputs as fixtures and attaches a fail-to-pass contract: the case must fail on the old code and pass on the fix, or the promotion is rejected as untrustworthy. Multi-turn behavior gets scenarios: a scripted conversation, or a red-team model improvising against an adversarial goal.

Ship: the test suite replays in CI against recorded fixtures, offline and without API keys. The tracely gate command exits non-zero, posts a commit status, and writes a comment on the pull request.

Alert: a rule defines both a trigger condition (a gate failure, a live judge verdict, a new failure cluster, a rate crossing a threshold) and what happens (a Slack message, an email, a webhook, an LLM step whose output the next step can use, or a Python expression for computed values). The alert flow is drawn on a canvas in the application.

Self-Hosting Tracely-ai with Docker Compose

The full stack runs as a single Docker Compose deployment covering the FastAPI backend, the Celery worker, the Next.js frontend, Postgres (registry), ClickHouse (events and scores), Redis (Celery broker), and MinIO (S3-compatible blob storage). The Makefile documents every step:

bash
docker compose up -d --build --wait

This starts all services and waits for health checks. To run on a port other than 8000 for the backend:

bash
TRACELY_BACKEND_PORT=8088 docker compose up -d --build --wait

After the stack is up, apply migrations and seed the default project:

bash
make migrate
make seed

The seed step creates a default project and an ingest key named tracely_dev_key. The demo target populates the full product with sample traces, clusters, cases, and gates in one command:

bash
TRACELY_API=$(TRACELY_API) uv run python scripts/seed_demo.py

The .env.example file documents the required environment variables: DATABASE_URL, CLICKHOUSE_HOST, REDIS_URL, S3_ENDPOINT_URL, and OPENROUTER_API_KEY (for LLM judge calls). Without an OpenRouter key, the pipeline still runs but evaluators degrade to no-ops. The ClickHouse system log (query_log, trace_log) has no TTL by default and accumulates silently; the CH_SYSTEM_LOG_TTL_DAYS variable controls a nightly sweep that caps it, defaulting to 3 days.

Hermetic Cases and the Fail-to-Pass Contract

The design of hermetic cases is the mechanism that separates Tracely-ai from pure observability tools. When a failing trace is promoted to a case, Tracely-ai records the tool and LLM outputs it produced as fixture files. The replay in CI uses those fixtures instead of making live calls, which means the replay is deterministic, fast, and free.

The fail-to-pass contract is a requirement attached to every case: the case must demonstrate a failure on the code that produced it and a pass on the fixed code. This prevents promoting a trace that happens to pass on the current codebase, which would be a test that never catches a regression. The README notes that if a promotion cannot demonstrate fail-to-pass, the promotion is not trusted.

For multi-turn conversations, scenarios replace simple replay. A scenario can script the conversation turn by turn, or it can assign an adversarial goal to a red-team model that improvises its responses during the replay. Assertions and a reference trajectory can be attached to verify the agent's path through the conversation, not just its final output.

Judge Calibration and Trend Analysis

Tracely-ai includes a calibration screen where the team labels judge verdicts against human review and sees each judge's agreement rate. The README notes the purpose directly: to catch an over-flagging judge before it is allowed to gate a release. An evaluator that marks too many passing runs as failures creates alert fatigue and blocks legitimate PRs.

The platform also tracks daily failure and gate pass-rates, latency percentiles, and token spend over time. Per-agent meta-analysis uses Spearman correlations and z-score outlier detection with LLM-synthesized commentary. The evaluator system supports multi-output verdicts (score, boolean, number, text, or JSON schema), advanced prompts with @VARIABLE references and a live preview, deterministic sampling, and advisory verdicts that flag without blocking.

Where Tracely-ai Has Limits

The full stack requires five persistent services: Postgres, ClickHouse, Redis, MinIO, and Celery workers. That is a significant operational surface for a small team. The Celery worker documentation notes that the UMAP and HDBSCAN libraries used for failure clustering are fork-fragile, so the worker pool must be set to solo or 1 in local development. Production requires setting CELERY_POOL to prefork and CELERY_CONCURRENCY to 2 or more, which is documented in the docker-compose.yml but not prominently in the README.

The repository has no GitHub releases, so there is no versioned changelog. The project requires Python 3.12 or later. The LLM judge pipeline routes through LangChain create_agent on OpenRouter; teams on air-gapped networks or with strict API-routing requirements will need to evaluate whether the LLM pipeline can be redirected.

Compared to Langfuse for AI Agent Observability

Langfuse is an open-source observability platform for LLM applications that covers tracing, evaluations, and a prompt management interface. It is one of the more widely deployed tools in the space. The main difference in scope is that Langfuse focuses on observability and analytics: it shows what happened and how scores change over time, but it does not include a mechanism to block a pull request or to promote a failing trace into a CI gate.

Tracely-ai's distinguishing design choice is the CI gate and the hermetic case promotion. For teams that primarily need a dashboard to inspect traces and understand evaluation trends, Langfuse covers that ground with a more established deployment history. For teams who want a production failure to automatically become a blocking regression test without writing the test by hand, Tracely-ai's pipeline is the specific tool.

Editorial conclusion

Teams running AI agents in production who want automated regression protection without hand-authoring eval datasets will find Tracely-ai's trace-promotion approach a better fit than tools that ask you to write test cases by hand. The full stack is self-hosted, which means you own the data but also own the infrastructure. Before deploying, run docker compose up -d --build --wait and verify that Postgres, ClickHouse, Redis, and MinIO all reach a healthy state before sending any traces; an unhealthy ClickHouse is the most common source of silent data loss at ingest.

Frequently asked questions

What is tracing in AI agents, and how does Tracely-ai use it?

In AI agents, a trace is a structured record of one run: the input, the sequence of tool calls, the model responses, and the output. Tracely-ai ingests these traces over OTLP/HTTP, grades them with evaluators, and uses failing traces as the raw material for regression test cases.

What is a reasoning trace in AI, and can Tracely-ai capture it?

A reasoning trace records the intermediate steps an AI agent takes before producing an output: tool calls, sub-agent delegations, and model responses at each step. Tracely-ai captures and displays these as a waterfall with agent, thinking, skill, generation, and hand-off spans, and can promote a failing reasoning trace into a hermetic regression case.

Does Tracely-ai require API keys to replay tests in CI?

No. The CI replay uses recorded fixture files that capture the tool and LLM outputs from the original failing run. No live model calls are made during replay, so no API keys are required and there is no per-test model cost.

Official sources

  1. Issues
  2. Jwuthri/Tracely-ai on GitHub
  3. License: MIT
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/jwuthri-tracely-ai.svg)](https://hysenlabs.com/projects/jwuthri-tracely-ai)