# LM Evaluation Harness: Reproducible Benchmarking for Language Models

> LM Evaluation Harness is an open source Python framework from EleutherAI for running standardized few-shot evaluations across language models and dozens of academic benchmarks. It powers the Hugging Face Open LLM Leaderboard and supports backends from local HuggingFace models to commercial APIs.

**EleutherAI/lm-evaluation-harness** — A framework for few-shot evaluation of language models.

- Repository: https://github.com/EleutherAI/lm-evaluation-harness
- Website: https://www.eleuther.ai
- Stars: 14,087 · Forks: 3,611
- Language: Python
- License: MIT
- Published: 2026-09-21 · Updated: 2026-09-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/eleutherai-lm-evaluation-harness

## What LM Evaluation Harness Solves and Who Uses It

Comparing language models across research papers has historically been difficult because different teams use different prompt formats, different few-shot example counts, and different scoring methods for the same nominal benchmark. LM Evaluation Harness addresses this by standardizing the entire evaluation pipeline: a shared set of publicly available prompt definitions, a consistent few-shot sampling mechanism, and a reproducible output format that lets researchers compare results directly.

The framework runs over 60 standard academic benchmarks with hundreds of subtasks and variants. According to the README, it is the backend for Hugging Face's Open LLM Leaderboard and has been used in hundreds of papers by organizations including NVIDIA, Cohere, BigScience, BigCode, Nous Research, and Mosaic ML. That position as leaderboard infrastructure means its task implementations and scoring rules are as close to a community standard as exists in open source LLM evaluation.

The intended users are ML researchers benchmarking new models before publication, teams comparing candidate models before deployment, and anyone who needs to reproduce evaluation results from a paper. The framework is not a monitoring tool and does not track model behavior in production.

## How Tasks Are Defined and Configured

Each benchmark in the harness is defined by a YAML configuration file inside the `lm_eval/tasks/` directory. These configuration files describe the dataset source, the prompt template in Jinja2 format, the few-shot settings, output post-processing rules, answer extraction logic, and the metric or aggregation to apply. The v0.4.0 release made config-based task creation the primary workflow, replacing earlier Python-only task definitions.

The Jinja2 template system allows editing prompts without changing Python code. Teams can import prompts from Promptsource, write their own templates, and configure how many few-shot examples to include per call. Multiple LLM generations per document are also configurable, which matters for tasks where a single generation is not a reliable signal.

Task groups allow organising related subtasks under a shared name. The `leaderboard` task group, added after the v0.4.0 release, contains the tasks that feed the Hugging Face Open LLM Leaderboard. Running that group requires the tasks to be installed as part of the package: they live under `lm_eval/tasks/leaderboard/`.

The CLI has three subcommands: `run` to execute an evaluation, `ls` to list available tasks or models, and `validate` to check a task configuration. A YAML config file passed with `--config` can hold all evaluation parameters, making it possible to reproduce a run from a single file rather than a long command line.

## Installing the Harness and Running a First Evaluation

The base install provides the evaluation framework without any model backend. Install it from the repository:

```bash
git clone --depth 1 https://github.com/EleutherAI/lm-evaluation-harness
cd lm-evaluation-harness
pip install -e .
```

Model backends are installed separately with optional extras. For HuggingFace transformers:

```bash
pip install "lm_eval[hf]"
```

For vLLM fast inference:

```bash
pip install "lm_eval[vllm]"
```

For commercial API providers such as OpenAI:

```bash
pip install "lm_eval[api]"
```

Multiple backends install together:

```bash
pip install "lm_eval[hf,vllm,api]"
```

The project requires Python 3.10 or later, as specified in `pyproject.toml`. To see which tasks are available after installation:

```bash
lm-eval ls tasks
```

To run an evaluation against a HuggingFace model, use the `hf` model type with the `--model_args` flag to specify the model checkpoint. The README gives hellaswag as a starter benchmark example and notes that a CUDA-compatible GPU is the assumed target for HuggingFace backend runs. Results are printed to standard output in a table format; the framework also supports logging to Weights and Biases or Zeno, as shown in the examples directory.

The current development version in `pyproject.toml` is `0.4.14.dev0`. The most recent stable release was v0.4.13, published on 2026-08-31.

## Supported Model Backends and Inference Modes

The HuggingFace backend, installed with `lm_eval[hf]`, loads models via the `transformers` library and requires `torch` and `accelerate`. It supports data-parallel evaluation across multiple GPUs. Quantized models are supported through GPTQModel and AutoGPTQ as optional dependencies. LoRA adapters and other PEFT-compatible adapters load through HuggingFace's PEFT library, specified as part of `model_args`.

The vLLM backend offers faster, more memory-efficient inference for large models and installs via `lm_eval[vllm]`, which requires vLLM 0.18 or later. SGLang support was added in early 2025 and is available as a separate backend type. For very large models such as Llama 405B, the README recommends using vLLM's OpenAI-compatible API server and evaluating through the `local-completions` model type rather than loading the model directly.

The API backend (`lm_eval[api]`) supports OpenAI and TextSynth, with support for batched and async requests. The API backend was refactored in mid-2024 to make it easier to add new API providers without modifying the core library.

The GPT-NeoX and Megatron-DeepSpeed backends target large-scale distributed training setups and are listed in the README as supported options for teams running those frameworks.

For models that produce chain-of-thought reasoning traces before giving a final answer, the `think_end_token` argument strips the reasoning prefix before scoring. It accepts a token string and is supported on the `hf`, `vllm`, and `sglang` backends.

## Extending the Harness with Plugins

As of September 2026, the plugin system allows registering custom model backends, filters, metrics, and aggregations from an external Python package without forking the repository. Two methods are available. The first is declaring an `lm_eval.*` entry point in the external package's `pyproject.toml`, which the harness discovers automatically with no additional configuration. The second is passing `--plugins` on the command line pointing at a local module path.

This matters in practice for teams building evaluations on models or deployment setups not covered by the built-in backends. Before the plugin system, the only path was forking the repository and maintaining a custom fork across upstream updates.

The framework also supports a prototype multimodal mode with the `hf-multimodal` and `vllm-vlm` model types for text-plus-image inputs with text outputs. The README describes this as a prototype feature and suggests teams requiring broader multimodal task coverage look at lmms-eval, a project that originally forked from LM Evaluation Harness and has extended it for that purpose.

Hugging Face model steering is supported through the `hf` backend, meaning evaluation can target steered or redirected model variants without separate configuration.

## Where the Framework Falls Short

LM Evaluation Harness does not evaluate models on open-ended generation quality directly. It scores outputs against expected answers or uses classification-style metrics, so tasks that require human judgment or that measure writing quality, coherence, or factual accuracy of free-form text fall outside the framework's native capabilities. Adding such metrics is possible via the plugin or YAML config system, but the README does not document built-in support for human evaluation loops.

The framework assumes access to a CUDA GPU for most realistic workloads. The README explicitly notes that HuggingFace backend runs assume a CUDA-compatible GPU. Running evaluations on CPU is possible but not practical for models of meaningful size.

Task configuration via YAML is flexible but not self-evident. Writing a new task requires understanding how the Jinja2 template system interacts with the scoring pipeline, how few-shot sampling works for the specific dataset format, and how output post-processing should be configured. The docs directory contains guides, but teams writing custom tasks should expect to study existing YAML examples before authoring their own.

The base install no longer bundles `transformers` or `torch`, which simplifies installation for teams that use only API backends. It also means that a bare `pip install lm_eval` produces a framework that cannot run any model evaluation without an additional step.

## LM Evaluation Harness vs LightEval

The most directly comparable alternative is LightEval, an open source evaluation library developed by Hugging Face. LightEval is built around the Hugging Face ecosystem and integrates tightly with the HF Hub and Transformers library for model loading and result tracking.

LM Evaluation Harness differs in scope. The harness supports backends beyond HuggingFace Transformers, including vLLM, SGLang, GPT-NeoX, Megatron-DeepSpeed, and commercial APIs. Its position as the backend for the Open LLM Leaderboard means its task implementations and few-shot evaluation logic are the reference for a large portion of published model comparisons. Teams that need to reproduce results cited in papers that used the harness have a straightforward path: install the same version and use the same task names.

LightEval's tighter HF Hub integration may be preferable for teams whose workflow is entirely within the HF ecosystem and who prioritise simpler setup over backend flexibility. For teams working with models deployed on vLLM or evaluated through commercial APIs, the harness is the more practical choice.

Both projects use MIT-compatible licenses. LM Evaluation Harness is MIT-licensed, per the `LICENSE.md` in the repository and the `pyproject.toml` license field.

## Conclusion

Researchers and engineering teams running systematic LLM comparisons should use LM Evaluation Harness when they need reproducible, published-prompt evaluations across standardized benchmarks. The framework is the most practical choice for any team already in the HuggingFace or vLLM ecosystem. Teams that need only a single quick benchmark check, or that work exclusively with proprietary non-API models, will find the setup overhead harder to justify. Before running any evaluation, verify that Python 3.10 or later is installed, install the correct backend extras for the target model, and confirm that available GPU memory matches the model's requirements.

## FAQ

### What is LM Evaluation Harness?

LM Evaluation Harness is an open source Python framework from EleutherAI for standardized few-shot evaluation of language models across over 60 academic benchmarks. It is the backend for Hugging Face's Open LLM Leaderboard.

### What are the benchmarks for LM Harness?

The harness includes over 60 standard academic benchmarks with hundreds of subtasks, covering tasks used in the Open LLM Leaderboard. Available tasks can be listed with `lm-eval ls tasks` after installation.

### How do I install LM Evaluation Harness?

Clone the repository and run `pip install -e .` for the base framework. Model backends are installed separately using optional extras such as `pip install "lm_eval[hf]"` for HuggingFace or `pip install "lm_eval[vllm]"` for vLLM. Python 3.10 or later is required.

### How does LM Evaluation Harness compare to LightEval?

LM Evaluation Harness supports a wider range of inference backends, including vLLM, SGLang, and commercial APIs such as OpenAI, and is the reference implementation behind the Hugging Face Open LLM Leaderboard. LightEval is a Hugging Face project focused on HF ecosystem integration. The README does not directly compare the two.

### How do I use LM Evaluation Harness with vLLM?

Install the vLLM backend with `pip install "lm_eval[vllm]"`, which requires vLLM 0.18 or later. For very large models, the README recommends using vLLM's OpenAI-compatible API server and evaluating through the `local-completions` model type.

## Sources

- [EleutherAI/lm-evaluation-harness on GitHub](https://github.com/EleutherAI/lm-evaluation-harness)
- [License: MIT](https://github.com/EleutherAI/lm-evaluation-harness/blob/main/LICENSE)
- [Project website](https://www.eleuther.ai)
- [README](https://github.com/EleutherAI/lm-evaluation-harness/blob/main/README.md)
- [Releases](https://github.com/EleutherAI/lm-evaluation-harness/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/eleutherai-lm-evaluation-harness
