# Pezzo: a self-hosted LLMOps platform for prompt versioning and observability

> Pezzo is an Apache-2.0 LLMOps stack that stores prompts, serves them to your app through a proxy, and records every call in ClickHouse. It is aimed at teams already running PostgreSQL and Redis who want prompt changes to ship without a redeploy.

**pezzolabs/pezzo** — 🕹️ Open-source, developer-first LLMOps platform designed to streamline prompt design, version management, instant delivery, collaboration, troubleshooting, observability and more.

- Repository: https://github.com/pezzolabs/pezzo
- Website: https://pezzo.ai
- Stars: 3,276 · Forks: 278
- Language: TypeScript
- License: Apache-2.0
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/pezzolabs-pezzo

## What problem Pezzo addresses, and for whom

Most teams start by pasting a prompt string into application code. That works until someone wants to change a wording, test a variant, or explain why a response was bad last Tuesday. Pezzo targets that gap. It is a platform for prompt management, instant delivery, collaboration and observability, and the README frames the value as saving on cost and latency while letting you "instantly deliver AI changes".

The intended user is a developer or platform engineer, not a prompt writer working alone. The repository is a TypeScript monorepo built with Nx, containing a server, a console and a proxy under apps/, with shared code in libs/. The README lists Node.js 18+ and Docker as prerequisites, and the client table covers Node.js, Python and LangChain integrations for prompt management, observability and caching.

Where it fits is the middle of the stack. Your application keeps calling an LLM, but the prompt text and the request metadata flow through Pezzo instead of living in a constant. That is a real architectural commitment, not a library you import and forget.

## Architecture: GraphQL server, ClickHouse analytics, and a proxy in front of your LLM calls

The README describes Pezzo as cloud-native and says it relies on PostgreSQL, ClickHouse, Redis and Supertokens. Those four dependencies map onto distinct jobs. PostgreSQL holds the relational data, and Prisma migrations are deployed against apps/server/prisma/schema.prisma. ClickHouse stores the high-volume request records, reached through @clickhouse/client and a Knex dialect package. Redis handles caching. Supertokens handles authentication.

The server is a NestJS application exposing a GraphQL API through Apollo, and the README points at http://localhost:3000/api/healthz as the health check. A separate proxy service sits in front of your LLM traffic: docker-compose.yaml gives it PEZZO_SERVER_URL pointing at the server and maps container port 3000 to host port 3001. That is the endpoint your application is meant to call.

The console is a Next.js front end, mapped from container port 8080 to host port 4200 in the compose file. The compose stack also uses local-kms on port 9981 for key management, and the env example lists KAFKA_BROKERS and OPENSEARCH_URL alongside the ClickHouse and Redis settings, so the dependency surface is wider than the four names in the README suggest.

## Running Pezzo with Docker Compose or in development mode

The README offers two paths. For a local full stack it points at the "Running With Docker Compose" page in the documentation. For development mode it lists the steps below, starting with dependencies.

Install the npm packages first:

```bash
npm install
```

Pezzo reads a .env file, and when using Docker you also need a .env.docker file, with .env.example as the reference. Then bring up the infrastructure services:

```bash
docker-compose -f docker-compose.infra.yaml up
```

Deploy the Prisma migrations, then start the server:

```bash
npx dotenv-cli -e apps/server/.env -- npx prisma migrate deploy --schema apps/server/prisma/schema.prisma
npx nx serve server
```

According to the README, you can confirm the server is up by navigating to http://localhost:3000/api/healthz. In a separate terminal, run GraphQL codegen in watch mode so types regenerate when the schema changes, then start the console:

```bash
npm run graphql:codegen:watch
npx nx serve console
```

The README states the console is then reachable at http://localhost:4200. Note that this development path is four long-running processes plus the Docker infrastructure, and the first command in the sequence is the one people forget.

## Where Pezzo is the wrong tool

The heaviest cost is operational. A working deployment needs PostgreSQL, ClickHouse, Redis, Supertokens and a key management service, and the compose file gates the server on health checks for postgres, supertokens, clickhouse and redis-stack-server, plus successful completion of two migration jobs. If you are a solo developer shipping a single prompt to one endpoint, that is more moving parts than the problem justifies. A prompt stored in a config file and reloaded on deploy gets you most of the way.

ClickHouse is the sharpest constraint. It is excellent at the append-heavy workload Pezzo generates, but it is not a database most application teams already run. Adding it means a new backup routine, a new upgrade cadence and a new failure mode. The README does not discuss sizing, retention or what happens when ClickHouse is unavailable, so the operational envelope is something you would have to establish yourself.

There is also a maturity question. The most recent release listed is v0.9.2 from 2024-05-15, while the last push to the repository was on 2026-08-21. That gap between release tags and commit activity is worth understanding before you depend on a specific version. The README also documents no rollback procedure for a prompt change, which matters because instant delivery is the headline feature.

## How Pezzo differs from Langfuse and from plain prompt files

Langfuse is the closest comparison in the observability space, and the difference is where each puts its centre of gravity. Langfuse is built around tracing and evaluation: you send it spans, and it gives you analysis over them. Pezzo pairs observability with a delivery mechanism. The proxy service is the tell. In the compose file it is a first-class container with its own host port, which means Pezzo expects to sit in the request path and hand back a prompt, not just record what happened after the fact.

Against a prompt file in your repository, the difference is the deploy. A file change needs a build and a release. Pezzo's stated goal is instant delivery, so a prompt edit in the console is meant to take effect on the next call. The trade-off is that prompt text now lives in a database you must back up and secure, and the audit trail for who changed what lives in Pezzo rather than in git history.

A third option is the LLM provider's own prompt tooling. That keeps you inside one vendor and avoids running ClickHouse, but it does not cover multi-provider setups, and the README's client table suggests Pezzo expects you to mix providers rather than commit to one.

## Licence, upgrades and the maintenance picture

The repository licence is Apache-2.0, and the README states the source code is available under that licence. Apache-2.0 includes an explicit patent grant and permits commercial use. One detail to check rather than assume: the root package.json declares "license": "MIT" while the repository metadata and the LICENSE file say Apache-2.0. The root package is marked private, so the practical effect is limited, but if your process scans package manifests rather than the LICENSE file, that mismatch will surface. This is not legal advice; ask your own counsel if the distinction matters to you.

Upgrades follow the release tags, and the published ones stop at v0.9.2 from 2024-05-15. The last push to the default branch was on 2026-08-21, so commits have continued past the last tagged release. That means tracking main and tracking a release are meaningfully different choices here. The compose file pins images to the latest tag for the server, console and proxy, which is convenient for a first run and a poor default for anything you care about.

Maintenance cost is dominated by the migration jobs. The compose stack depends on pezzo-prisma-migrate and pezzo-clickhouse-migrate completing successfully before the server starts, so a failed migration blocks the whole stack rather than degrading it.

## Conclusion

Adopt Pezzo if you already operate PostgreSQL and Redis and want prompt edits to reach production without a redeploy, and if you are willing to run ClickHouse for the observability side. Do not adopt it if you need a managed service with a published support path, or if your team has no capacity to keep a multi-service stack healthy. Before committing, verify the npm client version you intend to install, confirm that the proxy port 3001 is what your application will call, and check whether the docs cover the rollback story you need, because the README does not document rollback.

## FAQ

### What is Pezzo?

Pezzo is an open-source LLMOps platform for prompt design, version management, instant delivery, collaboration and observability. The repository is a TypeScript monorepo containing a NestJS GraphQL server, a Next.js console and a proxy that sits in front of your LLM calls.

### Which infrastructure services does Pezzo require?

The README states Pezzo relies on PostgreSQL, ClickHouse, Redis and Supertokens, and the env example also lists Kafka and OpenSearch settings. The Docker Compose file additionally uses a local key management service on port 9981.

### Which LLM clients does Pezzo support?

The README's client table lists Node.js, Python and LangChain, each supporting prompt management, observability and caching. The README invites you to open an issue if you need a client that is not listed.

## Sources

- [License: Apache-2.0](https://github.com/pezzolabs/pezzo/blob/main/LICENSE)
- [pezzolabs/pezzo on GitHub](https://github.com/pezzolabs/pezzo)
- [Project website](https://pezzo.ai)
- [README](https://github.com/pezzolabs/pezzo/blob/main/README.md)
- [Releases](https://github.com/pezzolabs/pezzo/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/pezzolabs-pezzo
