# openai/evals: a registry and framework for LLM evaluation

> OpenAI's evals repository pairs a benchmark registry with a YAML-driven framework for running your own tests. It is a strong fit if you already call OpenAI models and want reproducible comparisons; it is not a general-purpose evaluation platform.

**openai/evals** — Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.

- Repository: https://github.com/openai/evals
- Stars: 19,513 · Forks: 3,093
- Language: Python
- License: NOASSERTION
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/openai-evals

## What openai/evals is for, and who it is not for

The problem this repository addresses is comparison. The README frames it plainly: without evals, understanding how different model versions affect your use case is difficult and time intensive. The project ships two things that are easy to conflate. The first is a registry of existing benchmarks covering different dimensions of OpenAI models. The second is a framework for writing your own evals, including private ones that use your own data without publishing it.

The intended user is someone building on LLMs who wants a repeatable way to score model behaviour. The README is explicit that you do not need to write evaluation code to contribute: if you follow an existing eval template, you supply data as JSON and parameters as YAML. That lowers the entry barrier considerably, and it also defines the ceiling. Anything that does not fit a template needs code, and the README notes that evals with custom code are not currently being accepted into the registry. You can still write them locally, but the contribution path is closed for that category.

A second boundary is provider scope. The dependency list in pyproject.toml includes anthropic and google-generativeai alongside openai, so the package is not limited to one vendor at the dependency level. The README, however, is written around the OpenAI API key and OpenAI model dimensions. Treat the non-OpenAI dependencies as plumbing that exists, not as a documented multi-provider evaluation story.

## How an eval runs: YAML definitions, JSON data, completion functions

The architecture visible in the repository is a three-part split. Eval definitions live as YAML under evals/registry/evals, and the README points to evals/registry/evals/coqa.yaml as an example of the same dataset implemented across several templates. Data lives under evals/registry/data and is stored with Git-LFS, which is why the setup instructions start with fetching and pulling LFS objects rather than cloning alone. Execution is driven by Python, with two console entry points declared in pyproject.toml: oaieval and oaievalset.

For advanced cases the README names the Completion Function Protocol, documented in docs/completion-fns.md. This is the extension point for prompt chains and tool-using agents, meaning the framework does not assume a single request-response call. The examples folder reinforces that: examples/retrieval-completionfn.ipynb sits alongside notebook walkthroughs for lambada, mmlu and lafand-mt.

The registry is the part with the longest half-life. Because definitions are declarative, the same dataset can be scored under different templates, and coqa.yaml is the repository's own demonstration of that. The trade-off is indirection: to understand what a given eval measures you have to read the YAML, follow it to the data file, and then find the template that interprets the output. There is no single file that tells you the whole story of one eval.

## Installing evals and running a first eval

The README states a minimum required version of Python 3.9, and pyproject.toml sets requires-python to >=3.9. If you only want to run existing evals rather than author new ones, the README gives a pip install and nothing more.

```bash
pip install evals
```

If you intend to write or modify evals, the README instead suggests cloning the repository and installing in editable mode, so your changes take effect without reinstalling.

```bash
pip install -e .
```

The README also documents an optional install for the pre-commit formatters.

```bash
pip install -e .[formatters]
```

Registry data is stored with Git-LFS. From inside your local copy of the repository, the README gives this sequence, which populates the pointer files under evals/registry/data.

```bash
cd evals
git lfs fetch --all
git lfs pull
```

Fetching everything can be heavy. The README offers a scoped alternative using the include flag, where ${your eval} is the eval you want.

```bash
git lfs fetch --include=evals/registry/data/${your eval}
git lfs pull
```

Running an eval requires an OpenAI API key passed through the OPENAI_API_KEY environment variable. The README also warns that API usage costs money, so a first run should be a small eval rather than a full benchmark sweep. The README points to docs/run-evals.md for full instructions on running existing evals and docs/eval-templates.md for the available templates. If you want results written to Snowflake, four more variables are required: SNOWFLAKE_ACCOUNT, SNOWFLAKE_DATABASE, SNOWFLAKE_USERNAME and SNOWFLAKE_PASSWORD. One known behaviour worth knowing before you start: the README documents that an eval sometimes hangs after the final report, and that interrupting it is safe and the eval finishes immediately after.

## The contribution rules constrain what the registry can hold

The most consequential limitation is stated in the README rather than buried in a policy page: evals with custom code are not currently being accepted. Model-graded evals with custom YAML files are still welcome. The practical effect is that the registry skews toward declarative, template-shaped evaluations, and the most interesting bespoke logic stays in private forks.

That is a defensible choice for a registry that OpenAI staff review when considering model improvements, since reviewing arbitrary code is a different job from reviewing a YAML file and a JSON dataset. It is also a real constraint on anyone hoping to upstream an evaluation that needs custom scoring logic. The README's own advice for that audience is to follow an existing template, provide data in JSON and parameters in YAML, and lean on the Jupyter notebooks in examples for orientation.

Two smaller friction points are documented rather than hidden. The end-of-run hang is acknowledged as a known issue. And the licence situation deserves a direct look: pyproject.toml declares no license field, the repository carries LICENSE.md, and the README's disclaimer states that contributing means agreeing to make your evaluation logic and data available under the same MIT license as the repository. If you plan to contribute data you did not create, the README requires that you have adequate rights to upload it, and notes that OpenAI reserves the right to use contributed data in future service improvements. That last sentence is the one to read twice before submitting anything sensitive.

## How openai/evals compares with lm-evaluation-harness

The obvious alternative in this space is EleutherAI's lm-evaluation-harness, and the difference is structural rather than cosmetic. lm-evaluation-harness is built around a broad set of academic benchmarks run against many model backends through a single harness, with the goal of comparable numbers across models. openai/evals is built around a registry of eval definitions plus a template system, with the goal of letting you define and run evaluations that match your own workflow, including private ones on your own data.

That shows up in the extension model. In openai/evals, the primary artifact you author is a YAML file, and the README's pitch to non-coders rests on exactly that: follow a template and you need no evaluation code at all. In lm-evaluation-harness, adding a task generally means writing task code in Python. The second difference is the completion function protocol. openai/evals documents a path for prompt chains and tool-using agents, which matters if what you are evaluating is a system rather than a single model call. A harness oriented around static benchmark prompts has less to say about that case.

The trade-off runs the other way too. If your goal is a published, cross-vendor leaderboard-style comparison, openai/evals is the wrong shape: its README is written around OpenAI models and the OpenAI API key, and its registry is a contribution channel into OpenAI's own model improvement process. Choose based on whether you are measuring your system or measuring the field.

## Maintenance, upgrades and licence cost

The repository is not archived, and the last push was on 2026-04-14. That is roughly five months before today, which puts it inside the six-month window, but it is not a fast-moving project and the README gives no release cadence to rely on. There are two console entry points and a single package name, evals, so a pip upgrade is the whole upgrade story for consumers. Version 3.0.1.post1 is what pyproject.toml declares.

The dependency list is long and that is the real upgrade cost. It includes openai>=1.0.0, anthropic, google-generativeai, langchain, datasets, evaluate, spacy-universal-sentence-encoder, snowflake-connector-python, playwright, docker and torch as an optional extra. Several of these move quickly and independently. An editable install (pip install -e .) pins nothing, so a fresh install months later can resolve to a very different dependency set than the one the eval was written against. If you are running evals for comparison over time, that is a reproducibility hazard you have to manage yourself; the repository does not ship a lock file in the top-level entries.

On licensing: the README states that contributions are made under the same MIT license as the repository, and the disclaimer covers usage policies and data rights. pyproject.toml carries no license field, so the authoritative statement is LICENSE.md. This is not legal advice, and if you are contributing third-party data or running evals on regulated data, the rights question is yours to settle.

## Conclusion

Adopt openai/evals if your evaluation targets are OpenAI models or systems you can express as a completion function, and if you want a registry of existing benchmarks rather than a bespoke harness. Do not adopt it if you need a vendor-neutral runner across many providers with identical semantics, or if you expect custom-code evals to be accepted upstream, since the README states those are not being taken. Before committing, verify the Python version on your machine meets the 3.9 minimum, confirm whether Git-LFS is available for pulling registry data, and decide whether results need to land in Snowflake, because that path requires four separate environment variables.

## FAQ

### What is the meaning of evals in openai/evals?

In this project, evals are evaluations: a framework for testing large language models or systems built on them, plus a registry of existing benchmarks. The README describes them as the way to understand how different model versions might affect your use case.

### What are evals in LLM work?

The README defines evals as a framework for evaluating large language models or systems built using LLMs. It ships a registry to test different dimensions of OpenAI models and lets you write custom evals for your own use cases.

### How do you use openai/evals?

Install the package with pip install evals, set the OPENAI_API_KEY environment variable, then run existing evals following docs/run-evals.md. To author your own, the README suggests cloning the repository, installing with pip install -e ., and following docs/build-eval.md.

### How do you set up evals for agents in openai/evals?

The README points to the Completion Function Protocol in docs/completion-fns.md for more advanced use cases such as prompt chains or tool-using agents. The examples folder includes examples/retrieval-completionfn.ipynb as a worked case.

### What is openai/evals?

It is a framework for evaluating large language models and systems built with them, together with an open-source registry of benchmarks. The registry is stored with Git-LFS, and the README points to docs/run-evals.md for running existing evals.

## Sources

- [Issues](https://github.com/openai/evals/issues)
- [openai/evals on GitHub](https://github.com/openai/evals)
- [README](https://github.com/openai/evals/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/openai-evals
