Model or dataset
chirpz-ai/pandaprobe avatar
chirpz-ai/pandaprobe

PandaProbe self-hosts on Docker, and its default compose mounts your Google credentials

open source agent engineering platform: traces, evals, and metrics to debug and improve your AI agents. Integrates with LangGraph, CrewAI, Claude Agent SDK, and more.

784 stars121 forksPythonApache-2.0

At a glance

What is it?
An open source tracing and evaluation platform for AI agents, with a Next.js dashboard, a FastAPI app, Celery workers, PostgreSQL and Redis. The pull-mode compose file most people start from pulls a :latest backend image and mounts ~/.config/gcloud read-only into both the app and the worker, while auth is delegated to Supabase or Firebase and model calls go through LiteLLM.
Who is it for?
PandaProbe is a reasonable self-host target for a team that already runs containers, because the pull-mode path avoids compiling anything and the three compose files separate users, hot reload and tests cleanly. It is a poor fit for a first contact with agent tracing, since the README is short, the integrations live on a separate documentation site, and the hosted cloud with a free tier and no card is one click away for anyone not opposed to sending trace data to a vendor.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Six services, two named ports and one scheduler

The service table is the whole architecture in miniature. The frontend is a Next.js dashboard on port 3000 and the app is a FastAPI application server on 8000. Behind them sit a Celery background worker and a Celery Beat scheduler, neither of which publishes a port, plus PostgreSQL 16 on 5432 and Redis 7 on 6379 doing double duty as broker and cache. After `./start.sh` the two addresses the README names are the dashboard on port 3000 and an API reference served at `http://localhost:8000/scalar`. The language split is visible in the names: the repository is Python, the dashboard is JavaScript, and the background work is driven by Celery rather than by threads in the web process. At the top level of the repository sit `AGENTS.md`, `CLAUDE.md` and `CITATION.cff` next to the three compose files, so the project is set up both for coding agents and for citation, and the root Makefile exists only to delegate to the `backend/` and `frontend/` Makefiles underneath it.

start.sh is the install, and the compose file wants two env files first

Self-hosting is three commands, with Docker as the only stated prerequisite:

bash
git clone https://github.com/chirpz-ai/pandaprobe.git
cd pandaprobe
./start.sh

The compose file in the repository spells out what that script is expected to do, because the header of the file is a set of instructions:

bash
#   1. cp backend/.env.example  backend/.env.development
#   2. cp frontend/.env.example frontend/.env.development
#   3. docker compose up

So the deployment is not a bare `./start.sh`. Two environment files have to exist before the stack starts, one for the backend and one for the frontend, and each is created by copying an example that is not itself in the repository. There are three compose files rather than one: this one is labelled pull mode for open-source users, `docker-compose.dev.yml` is for source hot reload, and `docker-compose.test.yml` backs the integration tests. The hosted alternative is a separate product surface, the cloud with a generous free tier and no credit card required, and the quickstart and integration guides live on the documentation site rather than in the README.

Pull mode means the backend image is :latest, not a version

The compose file states its own policy in the header: it pulls pre-built public images from GHCR and needs no local build. The image reference it uses is:

yaml
    image: ghcr.io/chirpz-ai/pandaprobe-backend:latest

The tag is `latest`, and the same tag is used for the app and for the worker, which means the two always match each other and neither is pinned to a release. Nothing in the repository ties that tag to v0.6.1, the newest release, so a fresh `docker compose up` a week apart can run different code with an identical compose file. That is the trade the file makes: no compiler on your machine, and no way to say which version you are running without reading the image. The alternative is the dev file, which builds from source and mounts the working tree, which is a different deployment shape rather than a different flag.

The default compose mounts ~/.config/gcloud into app and worker

Both backend services carry the same volume, and the app service names the file it expects to find there:

yaml
    volumes:
      - ~/.config/gcloud:/home/appuser/.config/gcloud:ro
    env_file:
      - ./backend/.env.development
    environment:
      - APP_ENV=development
      - GOOGLE_APPLICATION_CREDENTIALS=/home/appuser/.config/gcloud/application_default_credentials.json

The mount is read-only, which limits what a compromised container can change, and it still hands the container the application default credentials of the host user. The worker mounts the same directory, so background jobs have the same view. This is the line to read twice before running the stack on a laptop that also holds production cloud access, since it is present in the default file rather than behind a comment. Two smaller things sit in the same block: the environment is `APP_ENV=development` in a file described as the open-source user path, and the credentials path is an absolute path inside the container that only makes sense because of the mount above it.

Startup ordering hangs off a /health endpoint checked with curl

The app service defines its own health check and the rest of the stack waits on it:

yaml
    healthcheck:
      test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
      interval: 30s
      timeout: 10s
      retries: 3
      start_period: 10s
    restart: on-failure
    depends_on:
      postgres:
        condition: service_healthy
      redis:
        condition: service_healthy

Two things follow. The application waits for Postgres and Redis to report healthy, and anything that waits for the application inherits the same condition, so a slow first boot propagates. The check is a `curl` binary inside the image rather than a probe defined by the runtime, which means the health of the service depends on curl being present in a pre-built image you do not control. The numbers are tight: 30 second intervals, a 10 second timeout, three retries and a 10 second grace period, so four failed probes inside roughly two minutes and the restart policy starts cycling. The worker has no health check at all, it only has `--autoscale=10,2`, which lets Celery scale between 2 and 10 worker processes.

Auth is delegated to Supabase or Firebase, models go through LiteLLM

The architecture diagram is a sequence diagram with eleven participants, and the cast is the design decision. An SDK or HTTP client talks to the FastAPI app, which talks to an auth service, which talks to an identity provider that is either Supabase or Firebase, then an identity service, then trace and eval services, with PostgreSQL, Redis, a Celery worker and an LLM engine carrying LiteLLM. The management plane is annotated as a bearer token, and the first message drawn is the client sending `Authorization: Bearer <idp_token>` to the API. The diagram stops there, on a line that begins with a single letter, so the call sequence past authentication is not spelled out on the page. Two consequences are worth stating plainly. There is no local user store to administer, since accounts live with an external identity provider. And model access is not the application's own client code, it is LiteLLM's, which is why the same image serves the cloud and a self-hosted install with different provider keys.

Tests are split by the container boundary, on different ports

The root Makefile documents its own strategy in a comment block: backend unit tests run on the host, backend integration tests use Docker with Postgres and Redis, and every frontend test runs on the host with yarn, with Docker reserved for the dev server. The integration target shows the price of that split:

make
backend-test-integration:  ## Run backend integration tests (starts test infra)
	docker compose -f docker-compose.test.yml up -d --wait
	cd backend && POSTGRES_PORT=5433 POSTGRES_DB=pandaprobe_test_db REDIS_PORT=638

The test stack is brought up with `--wait` so it is usable before the tests run, and it listens on different ports from the development stack, Postgres on 5433 rather than 5432, with its own database name, so a test run and a running dev stack can coexist. The last visible line of that command is cut off mid-value, at the Redis port. The rest of the Makefile is delegation: about forty phony targets from `backend-install` through `frontend-test-e2e` to `up`, `down`, `ps` and `logs-app`, each one a thin wrapper around the same target in the subdirectory Makefile. No target builds a Docker image, which is consistent with a repository that expects users to pull them.

Editorial conclusion

PandaProbe is a reasonable self-host target for a team that already runs containers, because the pull-mode path avoids compiling anything and the three compose files separate users, hot reload and tests cleanly. It is a poor fit for a first contact with agent tracing, since the README is short, the integrations live on a separate documentation site, and the hosted cloud with a free tier and no card is one click away for anyone not opposed to sending trace data to a vendor. Four things to check before `docker compose up`: whether you want `~/.config/gcloud` visible inside the backend containers, because both the app and the worker mount it read-only and the app sets GOOGLE_APPLICATION_CREDENTIALS from it; that you pin the backend image yourself, since the file pulls `:latest` and a change to that tag changes your deployment with no edit in the repository; that the health check has `curl` available in the image, because both startup ordering and the restart policy hang off `/health`; and which Postgres and Redis you point integration tests at, since the test file runs them on port 5433 with a separate database. The license is Apache-2.0, the newest release is v0.6.1 from August 26, 2026, and the last commit is dated September 23, 2026.

Frequently asked questions

What is PandaProbe?

An open source agent engineering platform for tracing, evaluating, monitoring and debugging AI agents, usable as the hosted cloud service or self-hosted. The documentation site covers integrations, including clients built on LangGraph, CrewAI and the Claude Agent SDK.

How do I self-host PandaProbe?

With Docker installed and running, clone the repository, `cd pandaprobe` and run `./start.sh`. The dashboard is then on `http://localhost:3000` and the API reference on `http://localhost:8000/scalar`.

What does PandaProbe need before the stack starts?

Two environment files, created by copying `backend/.env.example` to `backend/.env.development` and `frontend/.env.example` to `frontend/.env.development`. The compose file also mounts `~/.config/gcloud` read-only into both backend containers.

Which databases does PandaProbe use?

PostgreSQL 16 on port 5432 and Redis 7 on port 6379 as both broker and cache, fronted by a Celery worker and a Celery Beat scheduler. Integration tests run against the same services on separate ports, with Postgres on 5433 and its own database.

Does PandaProbe build locally or pull images?

The default compose file is pull mode and states that no local build is required, using pre-built images from GHCR under the tag `latest`. Source hot reload lives in a separate file, `docker-compose.dev.yml`, and the tests use `docker-compose.test.yml`.

Official sources

  1. chirpz-ai/pandaprobe on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/chirpz-ai-pandaprobe.svg)](https://hysenlabs.com/projects/chirpz-ai-pandaprobe)