Model or dataset
langchain-ai/openevals avatar
langchain-ai/openevals

openevals: readymade evaluators for LLM apps, and where they stop

Readymade evaluators for your LLM apps

1,206 stars123 forksPythonMIT

At a glance

What is it?
LangChain's openevals ships prebuilt LLM-as-judge prompts, code checks and trajectory matchers for Python and TypeScript. It is a starting point, not an evaluation platform, and the README says so.
Who is it for?
Adopt openevals if you already have a test harness or tracing backend and want prebuilt judge prompts, code evaluators and trajectory matchers you can edit, in Python or TypeScript. Do not adopt it expecting a hosted dashboard, dataset management or result storage: the repository has no such component, and the README points to agentevals for agent-specific evals.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 11 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 22, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What openevals actually solves for LLM application teams

Traditional software has assertions. LLM applications mostly do not, because the output is free text and the pass condition is a judgement call. openevals occupies that gap with a library of evaluators you call from your own test code. The README frames the intent plainly: the package is "a starting point for you to write evals for your LLM applications, from which you can write more custom evals specific to your application."

The audience is narrow and identifiable. You are building an LLM feature, you want to check its output before shipping, and you would rather not write a grading prompt from scratch. The package covers quality, safety, security, RAG, extraction and tool calls, code, sandboxed code, agent trajectory, plus exact match, Levenshtein distance and embedding similarity. It is not an evaluation platform. There is no server in the repository layout, no database, no results UI. The top-level entries are .github/, js/, python/, sandbox/ and scripts/, which tells you this is a library with two language implementations and a sandbox helper, not a service.

How the LLM-as-judge mechanism works

The core primitive is create_llm_as_judge, and its behaviour is simple enough to reason about without running it. You hand it a prompt and a model. When you call the returned evaluator, the keyword arguments you pass are formatted directly into that prompt, the prompt goes to the model, and the model's response is parsed into a result object with a score and a comment. The README states that parameters are formatted into the prompt, and that CONCISENESS_PROMPT is "just an f-string". That detail matters more than it looks: customisation is string interpolation, not a schema. If you want the judge to consider a reference answer, you pass a reference; if you want a different criterion, you edit the prompt text.

The output shape is configurable in both directions. You can change the score values so the judge returns floats instead of True/False, and you can change the output schema, including structured prompts and logging feedback with custom schemas. The model is configurable too, and multimodal inputs are supported either through an attachments parameter or through a LangChain prompt template.

The evaluator family is wider than the judge. Exact match and Levenshtein distance are string comparisons with no model call. Embedding similarity sits between the two. For code, there are Pyright and Mypy checks on the Python side and a TypeScript type-checking evaluator on the JavaScript side, plus a sandbox directory for running Pyright or TypeScript checks in isolation. For agents, trajectory match supports strict, unordered, subset and superset modes, with configurable tool-args match modes. The README also directs readers who want agent-specific evals to a separate repository, agentevals.

Installing openevals and running a first eval

The Python package installs from PyPI. The README gives this command, and the published release openevals==0.2.1 matches the distribution name.

bash
pip install openevals

The TypeScript package installs from npm and declares @langchain/core as a peer dependency in the README's install line, so install both.

bash
npm install openevals @langchain/core

The quickstart judges outputs with an OpenAI model, so an API key has to be present in the environment under the name the README uses.

bash
export OPENAI_API_KEY="your_openai_api_key"

The first real eval is short. This Python example builds a conciseness evaluator from a prebuilt prompt, then scores a fake answer to a weather question.

python
from openevals.llm import create_llm_as_judge
from openevals.prompts import CONCISENESS_PROMPT

conciseness_evaluator = create_llm_as_judge(
    prompt=CONCISENESS_PROMPT,
    model="openai:gpt-5.6-sol",
)

inputs = "How is the weather in San Francisco?"
outputs = "Thanks for asking! The current weather in San Francisco is sunny and 90 degrees."

eval_result = conciseness_evaluator(inputs=inputs, outputs=outputs)
print(eval_result)

The README shows the expected result as a dictionary with a key of 'score', a boolean False, and a comment explaining that the output opens with an unnecessary greeting. The judge is not checking facts here. It is checking whether the answer is concise, and the sample answer is not. The TypeScript version is the same shape: createLLMAsJudge takes the prompt and model, and the returned function is awaited with inputs and outputs.

Where openevals is the wrong tool

The README is explicit that the prebuilt prompts are a starting point, and that framing is also the main limitation. A judge prompt encodes a definition of quality. CONCISENESS_PROMPT decides that a greeting is a defect. That may be correct for a support bot and wrong for a product that depends on warmth. You are expected to read and edit the prompt, and the README's customisation sections exist for exactly that reason. Teams that install the package, wire up the prebuilt prompts and treat the booleans as ground truth will get a number that measures the prompt author's taste, not their product.

The second boundary is infrastructure. Nothing in the repository layout stores eval results, versions datasets or renders a dashboard. If you need to compare runs over time, track regressions across releases, or share results with people who do not read Python, openevals does not do that part. You supply the harness. The README's RAG and trajectory sections assume you already have retrieval outputs and agent traces to feed in; the package grades them, it does not collect them.

Third, the LLM-as-judge path costs money and adds latency per call, because each evaluation is a model request. The string evaluators (exact match, Levenshtein, embedding similarity) avoid the model call, but exact match on free-form text is brittle, and Levenshtein distance tells you about character overlap rather than meaning. Pick the evaluator that matches the failure you are trying to catch, not the one that is cheapest to run.

openevals compared with DeepEval and agentevals

DeepEval appears in the search phrases people use around this project, and the two libraries make different bets. DeepEval is a testing framework: it wraps evaluation in a pytest-style structure with its own runner and reporting, so the framework owns the test lifecycle. openevals is a set of evaluator functions. You keep your existing test runner, your existing CI, and your existing tracing, and you call an evaluator inside it. The practical consequence is that openevals composes with whatever you already have and gives you nothing to adopt beyond the functions, while DeepEval gives you more structure and more to learn. Neither is universally better; the choice depends on whether you want a framework or a library.

The comparison inside the same organisation is cleaner. The README points readers who need evals specific to LLM agents to agentevals, a separate repository. openevals does include agent trajectory evaluators, with strict, unordered, subset and superset matching plus tool-args match modes, so the split is not absolute. The distinction the README draws is that agentevals is the destination for agent-specific evaluation work. If your system is a single model call with a prompt, openevals is the right entry point. If your system plans, calls tools and runs for many steps, read the agentevals README before you settle on trajectory match alone.

Maintenance, packaging and licence

The repository is not archived, and the last push was on 2026-09-01. The most recent releases listed are openevals==0.2.1 and openevals-js==0.2.2, both published on 2026-08-18. Python and TypeScript are versioned and released separately, which is worth knowing before you assume a feature documented in the README exists in both at the same version. The README presents Python and TypeScript side by side throughout, but the version numbers are independent, so check the release you are installing against the section you are reading.

The package is MIT licensed. That is permissive enough for commercial use, and it imposes no copyleft obligation on your application. It says nothing about the models you route evaluations through: your judge prompts and your application outputs leave your environment and go to whichever provider the model string names, and that provider's terms, not the MIT licence, govern what happens to them. The MIT grant also carries no warranty, so the accuracy of a judge's verdict is your problem, not the maintainers'.

Upgrade cost is low in the ordinary case, because the public surface is a set of factory functions and prompt constants. The risk sits in the prompts. If a release changes a prebuilt prompt's wording, your pass rates can move without any change to your application, and nothing in the version number will tell you that. Pin the version, and diff the prompt constants you depend on when you bump it.

Editorial conclusion

Adopt openevals if you already have a test harness or tracing backend and want prebuilt judge prompts, code evaluators and trajectory matchers you can edit, in Python or TypeScript. Do not adopt it expecting a hosted dashboard, dataset management or result storage: the repository has no such component, and the README points to agentevals for agent-specific evals. Before committing, run the quickstart with a real OPENAI_API_KEY and read the prebuilt prompt you intend to use, since the judge criteria live in those prompt strings rather than in the evaluator code.

Frequently asked questions

What is openevals and who is it for?

openevals is a Python and TypeScript library of readymade evaluators for LLM applications, from LangChain. It is aimed at teams that want a starting point for writing evals and then customising them, rather than a hosted evaluation platform.

Does openevals work without LangChain?

The Python quickstart imports from openevals.llm and openevals.prompts and passes a model string such as openai:gpt-5.6-sol, with no LangChain import shown. The TypeScript install line does require @langchain/core alongside openevals, and the README documents customising prompts with LangChain prompt templates as one option.

Can openevals judge a RAG pipeline or an agent trajectory?

The README lists prebuilt RAG prompts for correctness, helpfulness, groundedness and retrieval relevance, and agent trajectory evaluators with strict, unordered, subset and superset match modes. It also points readers who need evals specific to LLM agents to the separate agentevals repository.

Official sources

  1. Issues
  2. langchain-ai/openevals on GitHub
  3. License: MIT
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/langchain-ai-openevals.svg)](https://hysenlabs.com/projects/langchain-ai-openevals)