# judgeval's open-source part is the client, the judging runs on someone else's servers

> An Apache-licensed agent observability SDK whose manifest declares version 0.0.0, whose quickstart imports a package its runtime dependencies do not include, and whose most detailed feature is a read-only SQL surface over telemetry held on a vendor's side.

**JudgmentLabs/judgeval** — The Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and monitoring.

- Repository: https://github.com/JudgmentLabs/judgeval
- Website: https://judgmentlabs.ai/
- Stars: 1,063 · Forks: 101
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/judgmentlabs-judgeval

## Tracing is local, judging is described as server-side

Split the product along the line the readme itself draws. The tracing layer is built on OpenTelemetry, instruments any function with a decorator, and captures inputs, outputs and token usage through a standard exporter. That part is a client, and it can point at whatever collector you already run.

Everything else in the value proposition runs somewhere else. Setup exports an organisation id and an API key. Online monitoring is described as automatically scoring live production traffic server-side with no latency impact. Judges run against live traffic or replay historical traces. And the SQL feature queries what the file calls Judgment's virtual schema, which abstracts the underlying storage, with the server validating queries and deriving organisation and project scope.

So the honest description of the architecture is: an open-source client, an Apache-licensed one, in front of a hosted service that holds the traces and does the scoring. That is a normal shape for this category of tool and the file is not evasive about it. What it does not do is say what happens if you stop paying or stop sending, which is the question to answer before you route production conversations through it.

The MCP server and the terminal tooling extend the same assumption: both exist to read traces and invoke judges, which are server-side objects.

## The manifest says version 0.0.0 and the tags say 1.3.3

The project manifest carries a placeholder version, the string zero point zero point zero, as a literal. Meanwhile the release tags are 1.3.1, 1.3.2 and 1.3.3, the newest published within seconds of the last push.

There is a script at the repository root whose name explains the placeholder, alongside a custom build hook and a packaging manifest that points at a source layout under a src directory. The version is therefore written into the artefact by tooling rather than committed to the tree.

That is a common pattern and it has one consequence worth planning for. Nothing in the source tree tells you what you have. A library that checks its own version against a server, and this package depends on a packaging library and ships an automatic update script, cannot read its version from the manifest it shipped with. If you pin judgeval, pin the artefact, not the checkout, and if you check versions at runtime, expect the number to come from somewhere other than the metadata.

The rest of the manifest is unusually thin in one direction and unusually careful in another. Two classifiers, one for the language and one for the operating system, with no development status, no framework marker and no typing marker, even though the wheel configuration explicitly ships a typing marker file from the package.

## The quickstart imports a package the runtime dependencies do not install

The install line is one command, and the credentials are two exports. Then the example begins:

```python
from judgeval import Tracer, wrap
from openai import OpenAI

Tracer.init(project_name="my-project")
client = wrap(OpenAI())
```

The runtime dependency list is eight libraries: an HTTP client, two OpenTelemetry packages, a fast JSON serialiser, a command line framework, a dotenv loader, a path specification matcher and a packaging library. None of them is a model client.

Every provider is in the development group instead: one major model's client, another, a generative AI client, a fourth provider, plus type stubs for a couple of them. That is the right place for them, since the package wraps clients you already have rather than shipping them. The consequence is that copying the quickstart after a plain install of the package fails on its second import.

The rest of the snippet has the same flavour. It decorates a function that calls a vector database object that is never created, and it decorates an agent function that calls a chat completion on a model name. The file then ends mid-string on the line that would run the agent, so the printed output is not shown either. Two providers named in the example and in the development group are absent from the integrations paragraph, which names four auto-instrumented services and three frameworks.

## Two command line tools, one of them in a different repository

The package manifest declares a console script pointing at a module inside this distribution, so a command is installed with the library whether or not you want it. That command needs a terminal UI framework and a dotenv loader, and both are runtime dependencies of the tracing SDK.

The readme's CLI section then points somewhere else entirely, at a separate repository under the same organisation, and describes what that tool does: managing agents, traces, judges, behaviours and evaluations from the shell, with a documentation page of its own.

So a reader who installs the package gets a command, and a reader who reads the readme is sent to a different project for the command the readme is describing. Which of the two is current, whether the bundled one is a stub, and whether the two share code, none of the visible files say.

There is also an MCP server, so the same capabilities are reachable from an assistant or an editor. Three access paths, two repositories, and one sentence in the readme telling you which one to use.

The optional dependency list has exactly one entry, an object storage client, which suggests bulk export of logs is the one feature that stays entirely under your control.

## The SQL surface allows one statement, a thousand rows, and no writes

The most precisely specified part of the readme is the SQL feature, and its limits are the point. It is read-only. The server validates incoming queries, rejects writes, and enforces organisation and project scope from the credential rather than from anything the caller sends. Physical tables are unsupported, as are multiple statements and caller-specified execution limits. Results are capped at 1,000 rows and 5 MiB, over-limit results return an error rather than truncating, and query text must contain a non-whitespace character and stay under 50,000 characters.

The example query is a clue about what sits behind it. A count with empty parentheses and a catalogue namespace for telemetry is not portable SQL, and it points at an analytical columnar store behind the virtual schema rather than at the database your agent actually uses.

There is a second sharp detail. Integers outside the range a JavaScript number can represent exactly arrive as decimal strings, including inside nested structures, with a worked example of a specific large integer arriving quoted. Small integers and floats stay numbers and column types keep their SQL types. That is a careful contract, and it exists because the response crosses a JSON boundary where large integers would otherwise be silently rounded. It is a good sign about the engineering and a reminder that the interface is an HTTP API with a serialiser in the middle.

## The SQL feature is behind an opt-in the quickstart never mentions

Two access levels exist and they are described in the same paragraph. The schema discovery call returns a Markdown reference with tables, column types, row semantics, examples and query limits, and it needs organisation viewer access but neither a resolved project nor the public query opt-in. The execution path needs the existing API key, organisation membership and a resolved project. The file also says plainly that viewer access and the public query rate limit apply.

Which means the SQL feature has a gate at the organisation level, and the readme's quickstart says nothing about it. A developer who installs the package, exports two credentials and follows the documented call gets an authorisation failure, and the reason is one sentence buried in a long paragraph about response fields.

The schema endpoint has an HTTP equivalent too, so the whole surface is usable without the SDK, which is either a deliberate integration path or a leak of the internal API. It returns a single object with a schema key holding the Markdown reference, and the execution endpoint takes a bearer token, an organisation header and a JSON body with one field.

Whether the discovery call also needs the public query opt-in in practice is exactly the kind of thing the split description leaves ambiguous, and it is worth asking the vendor before you build tooling on top of it.

## One framework appears in the examples and not in the integrations list

The integrations sentence names four auto-instrumented model services and three frameworks: two tracing or agent frameworks and one instrumentation project. The examples directory contains six projects, and one of them is for a Google agent development kit that sentence does not mention.

The same gap shows up in the development dependencies, which add a fifth model provider that appears in neither the integrations sentence nor the examples. So there are three lists in this repository that describe the same surface and none of them agrees with the other two.

The examples themselves are worth reading as a curriculum. There is a basic tracing example, a basic evaluation example, and then two that name the thing they are about: distributed tracing and linked traces. Linked tracing is the interesting one for anybody instrumenting a multi-agent system, since it is the difference between a span per function call and a span per conversation.

Two of the six examples match the framework integrations by name, which suggests those two are the maintained paths and the rest are illustrative.

## A linter frozen two minor series back, and no documentation directory

The dependency groups are where a project's real habits show. The type checker floats on a recent floor. The linter is held between two versions in an old minor series, a range narrow enough that it will not pick up a fix without a manifest edit, while the type checker next to it moves freely. Everything else in the group is unpinned at the top and current.

Test configuration lives in a separate file at the repository root rather than in the manifest, alongside a pre-commit configuration, a lock file from a different package manager than the one the readme tells you to use, and a custom build hook.

The wheel configuration lists the package directory and then an include list that repeats the same paths, plus the typing marker. Redundant, harmless, and a sign of a build section that grew by addition rather than by design.

The documentation is the outlier. Every documentation link in the readme points at an external site, and there is no documentation directory in the repository at all. The readme itself is a landing page: an overview, a why section, a quickstart, an integrations sentence, and then a long SQL paragraph that reads as reference documentation pasted into a readme because there was nowhere else for it to live.

## Conclusion

judgeval is worth a look if you already run an OpenTelemetry collector and want agent traces to carry semantic scoring rather than just spans, because the judges produce scored, labelled records you can replay against historical traces before shipping a fix, which is a better workflow than reading dashboards. Four things to know before you commit. The tracing client is open source and the judging is not local: setup requires an API key and an organisation id, monitoring is described as running server-side, and the SQL view is a virtual schema on the vendor's storage. So the data your agents saw goes to a third party by design, and you should decide that deliberately rather than discover it in a privacy review. The SQL surface is narrow on purpose, one statement, an allowlisted catalogue, a thousand rows, no writes, so treat it as a convenience rather than an analytics database. The published version is not the one in the manifest, which declares a placeholder and is rewritten by a script at build time. And pin your own versions, because the quickstart needs a model client that is not a runtime dependency of the package.

## FAQ

### What does judgeval do?

It is an open-source Python SDK for agent improvement built on OpenTelemetry. A decorator on any function captures inputs, outputs and token usage, prompt-based agent judges produce scored and labelled records of how the agent behaved, and those accumulate into a searchable history you can replay against to validate a fix before shipping.

### Does judgeval send my agent data to Judgment Labs?

The documented setup exports an API key and an organisation id, online monitoring is described as scoring live traffic server-side with no latency impact, and the SQL feature queries a virtual schema on their side with the server deriving organisation and project scope. The tracing layer is built on OpenTelemetry, but the judging and monitoring described here are hosted.

### What can Judgeval.sql do, and what can it not?

It runs read-only queries against a virtual schema. The server rejects writes, single statements only, no physical tables, no caller-specified execution limits, results capped at 1,000 rows and 5 MiB, and query text capped at 50,000 characters. Schema discovery returns a Markdown reference and needs organisation viewer access but no resolved project.

### How do I install judgeval and start tracing an agent?

Install the package, export an API key and an organisation id, call the tracer initialisation with a project name, wrap a model client, and decorate functions with the observe decorator and a span type. Python 3.10 or later is required, and the runtime dependencies are an HTTP client, two OpenTelemetry packages, orjson, typer, dotenv, pathspec and packaging.

## Sources

- [JudgmentLabs/judgeval on GitHub](https://github.com/JudgmentLabs/judgeval)
- [License: Apache-2.0](https://github.com/JudgmentLabs/judgeval/blob/main/LICENSE)
- [Project website](https://judgmentlabs.ai/)
- [README](https://github.com/JudgmentLabs/judgeval/blob/main/README.md)
- [Releases](https://github.com/JudgmentLabs/judgeval/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/judgmentlabs-judgeval
