Model or dataset
langwatch/langwatch avatar
langwatch/langwatch

LangWatch: self-hosted LLM evaluation, tracing and agent simulation in one loop

The platform for LLM evaluations and AI agent testing

4,886 stars399 forksTypeScriptApache-2.0

At a glance

What is it?
LangWatch bundles tracing, datasets, offline evaluation, prompt optimisation and an OpenAI-compatible AI gateway into a single Apache-2.0 codebase. It is aimed at teams that want regression testing and production observability without stitching five tools together, and the one-command local install is the part worth judging it on.
Who is it for?
Adopt LangWatch if you already trace LLM calls and want evaluation, datasets and prompt versions to live in the same place as the traces, and if you can run Postgres, Redis and ClickHouse yourself. Do not adopt it if you only need a hosted dashboard with no operational surface, or if your stack cannot emit OpenTelemetry spans.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

The gap LangWatch fills between tracing and regression testing

Most teams end up with a trace viewer in one tab and a spreadsheet of eval runs in another. The trace tells you what an agent did; the eval tells you whether it was any good; nothing connects the two. LangWatch's pitch is that the loop closes inside one system: trace a run, promote the interesting traces into a dataset, evaluate against that dataset, change the prompt or model, then re-test. The README describes this as "Trace → dataset → evaluate → optimize prompts/models → re-test" with "no glue code".

The audience is narrower than the tagline suggests. This is built for teams shipping LLM-powered agents that have already moved past a prototype, where someone needs to answer "did last week's prompt change break the refund flow?" LangWatch also bundles agent simulation, which runs scenarios against the full stack including tools, state, a user simulator and a judge, so the failure can be traced to a specific decision rather than an aggregate score. If you are still choosing a model for a single-shot summariser, the platform is heavier than the problem.

How the pieces fit: OpenTelemetry in, ClickHouse and Postgres behind

The integration story is OpenTelemetry/OTLP-native, and the README claims the platform is framework- and LLM-provider agnostic by design. In practice that means your application emits spans, and LangWatch stores and indexes them rather than requiring a vendor SDK in the request path. The repository layout backs this up: there are separate `sdks/python`, `sdks/typescript` and `sdks/go` directories, plus `packages/observability/` and `packages/clickhouse-client/` in the workspace.

The storage split matters for sizing. Postgres and Redis appear in the local development instructions alongside `opensearch`, and the one-command installer scaffolds `uv`, `postgres`, `redis` and `clickhouse` into `~/.langwatch/`. A Go service tree (`cmd/`, `pkg/`, `services/`) sits next to the TypeScript platform, and the AI gateway is explicitly a separate Go binary under `services/aigateway/` with its own Helm sub-chart at `charts/gateway/`. So this is not a single Node process you can drop on a small VM and forget. The gateway is also where governance lives: the README describes an OpenAI/Anthropic-compatible proxy with virtual keys, hierarchical budgets, inline guardrails, automatic fallback across providers and Anthropic `cache_control` passthrough, and states roughly 700 ns hot-path overhead. That number is the project's own claim, not something verified here.

One design detail worth noting: evaluation runs as a separate service directory, `services/langevals/`, rather than inside the main application. Evaluators are therefore something you can reason about and scale independently, which also means the local install has to decide which of them to fetch.

Installing LangWatch locally with npx @langwatch/server

The README calls this the fastest local path and says only Node.js is required. The package declares `"node": ">=20"` in its engines field, so check that first. The CLI downloads and installs `uv`, `postgres`, `redis`, `clickhouse`, the AI gateway binary and the Langy assistant runtime into `~/.langwatch/`, generates a `.env` with local secrets, then starts every service in parallel and opens the app.

bash
npx @langwatch/server

When it finishes you should see the UI at `http://localhost:5560`, where the README says you create your first project and API key. Everything the installer writes lives under `~/.langwatch/`, and the README notes that `rm -rf ~/.langwatch` is a clean reset. That is a genuinely useful property for a stack with four backing services.

Three environment variables in `~/.langwatch/.env` control optional components, and the README gives their defaults and costs explicitly. `LANGWATCH_ENABLE_LANGY` defaults to `true` and adds about 45MB for the assistant runtime, which the README warns runs unsandboxed as your user on your own machine. `LANGWATCH_ENABLE_PRESIDIO` defaults to `false` and adds roughly 670MB of language model for PII detection, larger than the rest of the evaluator environment combined. `LANGWATCH_ENABLE_LINGUA` defaults to `false` and adds about 95MB for language detection. Every other evaluator installs either way. Edit the file and restart the server to change them.

If you would rather use containers, the README gives a Compose path from a clone of the repository:

bash
git clone https://github.com/langwatch/langwatch.git
cd platform/app
cp platform/app/.env.example platform/app/.env
docker compose up -d --wait --build

The same `http://localhost:5560` address is where you land afterwards. Note the path: the Compose file lives under `platform/app`, not at the repository root, so a naive `docker compose up` from the clone root will not find it.

The evaluator downloads are the real install cost

The headline is one command, but the honest part of the README is the table of optional components. Presidio's PII model alone is about 670MB, which is more than everything else in the evaluator environment put together. On a laptop that is an inconvenience; in a CI image or a locked-down build environment it is a procurement conversation.

The Langy assistant deserves the same scrutiny. It is on by default, and the README states plainly that its workers run unsandboxed as your user on your own machine. That is a reasonable default for a local dev tool and a questionable one anywhere shared. Turning it off is a one-line change in `~/.langwatch/.env`, and nothing in the README suggests the rest of the platform depends on it.

There is a subtler limitation in the architecture. Because LangWatch is OTLP-native and framework-agnostic, the quality of what you see is bounded by what your instrumentation emits. If your agent's tool calls are not spanned, the simulation and evaluation layers have nothing granular to score, and you are back to judging final outputs. The README positions OpenTelemetry support as a benefit, and it is, but it is also a prerequisite you have to satisfy before the platform earns its keep.

How LangWatch compares with Langfuse and LangSmith

The comparison people search for is LangWatch versus Langfuse and LangSmith, and the meaningful difference is where evaluation lives. Langfuse is primarily an observability and tracing backend with evaluation features layered on top; you bring your own harness or use its SDKs to score traces. LangSmith is tied to the LangChain ecosystem, which is convenient if you are already there and a reason to look elsewhere if you are not.

LangWatch's distinguishing claim is that offline evaluation and prompt optimisation are first-class parts of the same product rather than add-ons, and that it stays framework-agnostic through OpenTelemetry rather than a framework-specific callback system. It also ships the AI gateway as a separate Go binary, so cost governance and provider fallback sit inside the same deployment rather than in a second vendor. Whether that consolidation is worth it depends on how much you dislike running two systems. If your evaluation needs are simple and your tracing is already solved, Langfuse's narrower scope is easier to operate. If you are deep in LangChain and want managed everything, LangSmith removes decisions LangWatch makes you own.

Licence, self-hosting and the cost of upgrades

The licence is Apache-2.0, and the README's badge describes the model as "open-core: Apache 2.0 floor + Enterprise extension". That framing is worth reading carefully. The Apache-2.0 grant covers the repository as published, but the badge signals that some capability sits behind a commercial edition. Nothing in the repository description enumerates which features are in which tier, so if a specific capability is load-bearing for you, confirm its tier before you build a migration plan around it. This is a description of the licence situation, not legal advice.

Upgrade cost is shaped by the deployment surface. The self-hosting documentation covers Docker Compose, Kubernetes via Helm, and cloud-specific on-prem setups for AWS, Google Cloud and Azure, plus a hybrid mode for data-residency requirements. Each of those is a different amount of work to keep current. The repository uses release-please with per-package tags, and the recent releases are versioned independently: `sdks/go/v1.0.0`, `[email protected]`, `[email protected]`. Independent versioning is normal for a monorepo but means "upgrade LangWatch" is not one version number. The last push to the default branch was on 2026-09-09.

The reset story helps here. Because the local install is self-contained under `~/.langwatch/`, you can throw away a broken local environment entirely rather than debugging a half-migrated database. That does not apply to a Helm deployment, where the Postgres and ClickHouse migrations are yours to run and roll back.

Editorial conclusion

Adopt LangWatch if you already trace LLM calls and want evaluation, datasets and prompt versions to live in the same place as the traces, and if you can run Postgres, Redis and ClickHouse yourself. Do not adopt it if you only need a hosted dashboard with no operational surface, or if your stack cannot emit OpenTelemetry spans. Before committing, run npx @langwatch/server, confirm which evaluators you actually need, and check the size of the LANGWATCH_ENABLE_PRESIDIO and LANGWATCH_ENABLE_LINGUA model downloads against your disk budget.

Frequently asked questions

What is LangWatch?

LangWatch is described as the platform for LLM evaluations and AI agent testing, covering tracing, datasets, offline evaluation, prompt optimisation and an OpenAI-compatible AI gateway in one Apache-2.0 codebase. It is built for teams that need regression testing, simulations and production observability without assembling custom tooling.

What are the alternatives to LangWatch?

The repository does not name alternatives, but searches around this project pair it with Langfuse and LangSmith. Langfuse is primarily a tracing backend with evaluation layered on, and LangSmith is tied to the LangChain ecosystem, while LangWatch keeps evaluation and prompt optimisation inside the same product and integrates through OpenTelemetry.

What does LLM stand for?

The repository does not define the term. LangWatch's README and package description refer to LLM-powered agents and LLM evaluations without expanding the abbreviation.

How do large language models work?

The repository does not cover how language models work. LangWatch treats models as interchangeable providers behind an OpenAI/Anthropic-compatible gateway and does not document model internals.

Official sources

  1. langwatch/langwatch on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/langwatch-langwatch.svg)](https://hysenlabs.com/projects/langwatch-langwatch)
Community notes

Community notes