# LiveBench's own manifest records the coding category losing 28 points in a bare venv

> A benchmark built to avoid contamination and to avoid an LLM judge, shipping four vendor SDKs, a default release that is not the one fully published, and one measured scoring trap written into its dependency list.

**LiveBench/LiveBench** — LiveBench: A Challenging, Contamination-Free LLM Benchmark

- Repository: https://github.com/LiveBench/LiveBench
- Stars: 1,336 · Forks: 125
- Language: Python
- License: NOASSERTION
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/livebench-livebench

## The manifest records a 28 point coding false zero

The dependency list ends with three packages and a comment explaining why they cannot be optional. The coding grader executes model generated solutions inside the same interpreter you installed LiveBench into, so anything a solution imports has to be present:

```toml
# grading-sandbox deps: the coding grader executes model solutions with THIS
# interpreter (reliability_guard pre-imports matplotlib; solutions commonly import
# scipy/sortedcontainers). A venv without these false-zeros the coding category
# (measured 2026-08-22: coding graded 39.6 in a minimal venv vs its true 68.2).
"matplotlib", "scipy", "sortedcontainers"
```

That last line is the most useful sentence in the repository. The same model, the same questions, scored 39.6 in a minimal environment and 68.2 in a complete one, a gap of 28.6 points caused by missing imports rather than by model behaviour. Anything that reads a coding number without knowing which environment produced it is reading a number that may be an artifact.

## The default release is not the release that is fully public

LiveBench names a current release of 2025-04-25 and then warns that not all of its questions are on Hugging Face. To evaluate every category you are told to pass `--livebench-release-option 2024-11-25` to the scripts instead, which is what the basic usage example already does:

```bash
python run_livebench.py --model gpt-4o --bench-name live_bench/coding --livebench-release-option 2024-11-25
```

So the newest release is partly withheld, and the older one is the reproducible choice. That is a defensible way to keep answers private while the set is being written, but it means a run without the flag and a run with it are different benchmarks with different totals. Version 2024-11-25 and version 2025-04-25 should never be compared side by side, and the README does not offer a way to tell which one a saved result file came from.

## The agentic coding category can occupy 150GB of images

Evaluating the agentic coding questions needs Docker on the machine, checked before those tasks run with `docker --version`. The README then states the storage cost plainly: building task specific images may take up to 150GB, and the images are needed for both inference and evaluation, with optimisation of that requirement left as future work.

This is the one place where the benchmark's self contained claim breaks down. Inference for text tasks is an HTTP call to an endpoint, but the coding categories run generated code, and running untrusted generated code is what the images are for. Plan disk before you start, and remember the images are per task, so the number grows with the subset you select. Scoring the two non agentic coding tasks avoids this entirely.

## Four model SDKs ship as direct dependencies of a benchmark with no judge

The dependency list runs to roughly forty seven entries. Five of them exist to reach model providers: `anthropic>=0.3`, `openai`, `google-genai`, `together`, and `litellm`, the last acting as a router in front of the others. Alongside sit `docker`, `gitpython`, `PyGithub`, `litellm`, `spacy`, `sympy>=1.13` and a single exact pin, `latex2sympy2==1.9.1`.

That dependency shape is worth stating plainly, because the project's central claim is that every question has verifiable ground truth and can be scored without an LLM judge. That claim is about scoring, and it holds. It is not a claim about the harness, which still talks to several commercial providers directly. Anyone worried about model vendors seeing their prompts is worrying about the wrong half of this repository.

## --api-key takes a raw secret while --api-key-name does not

Two options cover credentials, and the difference matters on a shared machine. `--api-key-name` takes the name of an environment variable, defaulting to OPENAI_API_KEY for OpenAI models. `--api-key` takes the key value itself, on the command line. A value passed that way is written into shell history and is readable by anything that can list processes on the box.

The endpoint is configurable too, through `--api-base` for OpenAI compatible servers, which is the supported path for self hosted models. The same section recommends serving a local model with vllm behind such an endpoint and states plainly that local model inference is unmaintained. So the supported story is remote endpoint plus local server, and the unsupported one is in-process inference.

Two more options exist for long runs: `--resume` continues an interrupted run and `--retry-failures` retries questions that failed in earlier ones.

## --mode parallel spawns one tmux session per category

Parallelism is split across two flags that do different jobs. `--parallel-requests` sets how many questions are in flight inside a single task instance. `--mode parallel` is coarser: it creates separate tmux sessions per category or per task, and parallelises the ground truth evaluation phase as well. Together they are the throughput setting for a high rate limit:

```bash
python run_livebench.py --model gpt-4o --bench-name live_bench --mode parallel --parallel-requests 10
```

For a low rate limit the README's advice is to keep `--mode parallel` and drop the request count, on the grounds that session overhead is wasted on one or two tasks. The catch is that `--mode parallel` requires tmux installed, which is the only external program besides Docker that the evaluation path depends on. A container without tmux in it can still run with `--parallel-requests` alone.

## Version 0.0.4, no releases, and a leaderboard frozen at September 2024

The manifest declares version 0.0.4 and there are no published releases at all. Release history lives in changelog.md, which the README points at for details about each LiveBench release, so version tracking is a file to read rather than a tag to compare. The last push to main was on 2026-09-29.

The README's own leaderboard excerpt is headed Top models as of 30th September 2024 and defers to the hosted site for anything current, while the newest question set it names is dated 2025-04-25. The pinned snapshot in the file is therefore older than the benchmark it documents.

On licensing, three signals conflict. The platform record reports no license, the classifier claims the Apache Software License, and a LICENSE file sits at the root. The repository also states that it contains code from LiveCodeBench and IFEval, which is a provenance note a reader should weigh when deciding what the licence covers.

## Conclusion

The evaluation design is the strongest part of LiveBench: monthly question drops, verifiable ground truth, and no judge model to argue with. Two operational facts should decide whether it suits you. The agentic coding category builds Docker images that can reach 150GB, so budget storage before you promise a coding number. And the release you get by default is not the release that is fully public, so pass --livebench-release-option 2024-11-25 explicitly or your category totals will not match anyone else's. Install from the repository rather than a hand-built environment, and read the licensing position yourself: the platform record reports no license while the classifier claims Apache and a LICENSE file sits at the root.

## FAQ

### What is LiveBench used for?

It is a benchmark for evaluating language models. It releases new questions monthly to limit contamination, draws questions from recently released datasets, arXiv papers, news articles and IMDb movie synopses, and scores against verifiable ground truth so that no LLM judge is needed. It currently covers 18 tasks across 6 categories.

### What is the latest version of LiveBench?

The manifest declares 0.0.4, and the repository has no published GitHub releases. Release history is kept in changelog.md, and the newest question set named in the README is 2025-04-25, of which not all questions are public on Hugging Face.

### Is LiveBench good for beginners?

Nothing in the repository addresses beginners. Running it means a Python 3.10 environment, an OpenAI compatible endpoint or a self hosted model served with vllm, Docker for the agentic coding category, and tmux if you want category level parallelism. The quickest starting point is examples-free basic usage with run_livebench.py on a single task subset.

### How much disk does LiveBench need for coding tasks?

The README states that building the task specific Docker images for the agentic coding questions may take up to 150GB, and that those images are needed for both inference and evaluation. Docker must be installed and available before those tasks run.

### Can a LiveBench coding score be wrong because of the environment?

Yes, and the project measures it. The manifest records that coding graded 39.6 in a minimal venv against a true 68.2, because the grader executes model solutions with the same interpreter LiveBench is installed into, so matplotlib, scipy and sortedcontainers are declared as dependencies rather than extras.

## Sources

- [Issues](https://github.com/LiveBench/LiveBench/issues)
- [LiveBench/LiveBench on GitHub](https://github.com/LiveBench/LiveBench)
- [README](https://github.com/LiveBench/LiveBench/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/livebench-livebench
