Model or dataset
huggingface/lighteval avatar
huggingface/lighteval

Lighteval: a Hugging Face toolkit for running LLM benchmarks across vLLM, SGLang and hosted endpoints

Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends

2,549 stars561 forksPythonMIT

At a glance

What is it?
Lighteval is an MIT-licensed Python package that runs 1000+ evaluation tasks against models served remotely or loaded in memory. It is the evaluation stack behind Hugging Face's own leaderboard work, and its main constraint is that it is untested on Windows.
Who is it for?
Adopt Lighteval if you need to run a named benchmark such as gsm8k, mmlu or gpqa:diamond against a model you either serve yourself or hold in memory, and you are working on Linux or macOS. Skip it if your evaluation must run on Windows, since the README states the package is completely untested there, or if you need a stable released API rather than a package whose pyproject.toml still declares version 0.13.1.dev0.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Lighteval is for, and who actually needs it

The problem Lighteval addresses is the gap between a model checkpoint and a defensible number. Running MMLU or GSM8K by hand means writing prompt templates, parsing generations, handling few-shot examples, and then repeating all of it when you swap the serving stack. Lighteval packages that work as named tasks and separates it from the model, so the same benchmark string can be pointed at a local GPU, a vLLM server, or a hosted inference provider.

The intended audience is engineers who already have a model somewhere and need comparable scores. The README describes two situations explicitly: a model "being served somewhere" or one "already loaded in memory." That split drives the entire command surface. If you are a researcher comparing two fine-tunes of the same base model, you want the in-memory path. If you are shipping a product and want to know whether the model behind your API regressed after a provider update, you want the endpoint path.

The project comes from Hugging Face's Leaderboard and Evals Team, which matters less as an endorsement than as an explanation for the task coverage: 1000+ tasks spanning knowledge, math, code, multilingual and long-context evaluation. The README lists specific benchmarks rather than categories alone, including MMLU, MMLU-Pro, GPQA, GSM8K, MATH, AIME24, AIME25, LiveCodeBench, IFEval, MT-Bench, RULER, HELM and BIG-Bench. Multilingual coverage is unusually broad for a toolkit of this size, with named sets for Arabic, Filipino, French, German, Serbian, Turkic languages, Chinese, Russian and Kyrgyz.

The backend split: one task string, six execution paths

Lighteval's architecture is a set of entry points that share a task and metric layer but differ in how tokens get generated. The README lists six: lighteval eval (which uses inspect-ai as a backend and is marked as preferred), lighteval accelerate for CPU or multi-GPU via Hugging Face Accelerate, lighteval nanotron for distributed runs, lighteval vllm, lighteval sglang, and lighteval endpoint with four subcommands covering Hugging Face Inference Endpoints, a local Text Generation Inference server, LiteLLM, and Hugging Face inference providers.

That is a wider backend menu than most evaluation tools offer, and the practical consequence is that the task definition is decoupled from the generation mechanism. A benchmark identifier like gpqa:diamond or gsm8k resolves to a dataset, a prompt template and a metric; the backend decides only how completions are produced. The colon syntax in gpqa:diamond selects a subset of a task, which is how the same benchmark can be run at different difficulty or split levels without a separate task name.

The design has a cost. Because each backend is a separate integration with its own dependencies, the install is not one-size-fits-all. The README notes that Lighteval allows for "many extras" at install time and points to a separate installation page for the complete list. In practice this means the failure mode you are most likely to hit first is a missing extra, not a broken benchmark. The repository layout reinforces the separation: examples/model_configs/ and examples/nanotron/ exist as distinct directories, and the src/ tree is organized by model family rather than by task.

Results are handled by an EvaluationTracker that writes to an output directory, and the README's framing is that sample-by-sample results are saved so you can debug individual cases rather than only reading an aggregate score. For benchmark work where a single parsing bug can shift a number by several points, per-sample output is the more useful artifact.

Installing Lighteval and running your first benchmark

The README gives a single install command and a platform warning in the same breath: the package is "currently completely untested on Windows," and should be fully functional on Mac and Linux. Take that literally. If your CI runs on Windows, this is not a tool you can adopt without doing the porting yourself.

The base install is one line.

bash
pip install lighteval

That gets you the core package. Backends such as vLLM, SGLang, Nanotron and Accelerate are pulled in through extras, and the README defers the full list to the installation page rather than enumerating them. If you intend to push results to the Hugging Face Hub, authenticate first; the README uses the hf CLI for this.

shell
hf auth login

The quickest real evaluation uses a remote inference service, so nothing has to be downloaded or loaded locally. The README's own example evaluates an OpenAI gpt-oss-20b model through Hugging Face inference providers on the diamond subset of GPQA.

shell
lighteval eval "hf-inference-providers/openai/gpt-oss-20b" gpqa:diamond

The argument order is model first, then task. The model string encodes the backend (hf-inference-providers) and the provider model path, which is how a single command can target a hosted model without a config file. Expect the run to stream progress and write per-sample results into the output directory configured on the tracker.

If your model is already loaded in Python, the README shows the programmatic route: build an EvaluationTracker with an output_dir, construct PipelineParameters with launcher_type set to ParallelismManager.NONE and a max_samples cap, wrap a Transformers model through TransformersModel.from_model with a TransformersModelConfig, and hand both to Pipeline. The max_samples=2 in that example is a smoke-test value, useful for confirming the wiring before committing GPU hours.

Before running anything large, check the task name. The README points to the Open Benchmark Index space for browsing what exists, and a task that does not resolve will fail before generation starts.

Where Lighteval gets in your way

The Windows gap is the clearest limitation, and it is stated as a support boundary rather than a bug. There is no partial-support language in the README, no "mostly works" caveat. That is honest, but it also means any team standardized on Windows workstations has to run Lighteval in a container or on a remote Linux host.

The second constraint is the version signal. The pyproject.toml in the repository declares version "0.13.1.dev0" and the classifiers include "Development Status :: 3 - Alpha". The most recent tagged release is v0.13.0 from 2025-11-24, with v0.12.2 and v0.12.1 before it in the same month. Alpha status plus a dev version in the main branch means the API you write against today may move. The README's Python example imports from lighteval.logging.evaluation_tracker, lighteval.models.transformers.transformers_model and lighteval.pipeline, which are internal module paths rather than a curated public API. Pin your version if you script against them.

The third issue is the one most evaluation tools share: task definitions change. A benchmark like MATH or AIME24 is only comparable across runs if the prompt template and parsing rules are identical. Lighteval lets you write custom tasks and custom metrics, which is the right feature to have, but it also means two teams running "the same" benchmark through different task definitions can produce numbers that should not be compared. The README offers no guidance on versioning task definitions, and the documentation does not describe a mechanism for pinning them.

Finally, the backend menu is a maintenance surface. Six entry points, four endpoint subcommands, and separate model config directories each carry their own dependency and their own failure modes. If you only ever evaluate one model family on one GPU type, most of that surface is dead weight you still have to keep installed.

Lighteval against lm-evaluation-harness

The obvious comparison is EleutherAI's lm-evaluation-harness, and the difference is architectural rather than a matter of task count. Both run named benchmarks against language models. The harness grew up around local Hugging Face model loading with a CLI that takes a model and a task list, and its task registry is the reference implementation many papers cite. Lighteval's distinguishing choice is the breadth of its serving backends. Where the harness centers on loading weights, Lighteval treats "served somewhere" as a first-class case, with dedicated subcommands for Inference Endpoints, a local TGI server, LiteLLM and inference providers.

That matters if your model is not yours to load. Evaluating a hosted model through LiteLLM in Lighteval is a supported path with a named subcommand; reproducing that in a harness-oriented workflow typically means writing a custom model wrapper. The reverse is also true: if you want a single well-trodden path with the widest third-party task ecosystem, the harness is the safer default, and Lighteval's own README points users toward writing a custom model API when the built-in backends do not fit.

The second difference is inspect-ai. Lighteval's preferred entry point, lighteval eval, delegates to inspect-ai as a backend rather than implementing generation itself. That is a deliberate layering choice, and it means the preferred path inherits inspect-ai's behavior and its release cadence alongside Lighteval's own.

A third option worth naming is the Hugging Face Open Benchmark Index space, which the README links as the way to browse supported tasks. It is not a competitor to either tool, but it is where you should look before assuming a benchmark is missing.

Licence, releases and what upgrades cost you

Lighteval is MIT licensed, both in the LICENSE file and in the pyproject.toml license field, and the setup.py header carries the same MIT text with the copyright attributed to The HuggingFace Team for 2024. MIT is permissive: you can use it commercially, modify it and redistribute it, provided the copyright notice and permission notice travel with the code. That is a summary of the licence text, not legal advice; if you are redistributing Lighteval inside a product, have your own counsel read the LICENSE file rather than this paragraph.

The practical licence question is not the code but the data. Benchmarks such as GPQA, MMLU or LiveCodeBench are datasets with their own terms, and Lighteval downloads them at run time. The README does not discuss dataset licensing, so verify the terms of each benchmark you run before publishing scores.

Upgrade cost is driven by the release cadence. Three tagged releases landed within November 2025 (v0.12.1 on 2025-11-06, v0.12.2 on 2025-11-12, v0.13.0 on 2025-11-24), and the main branch has since moved to 0.13.1.dev0. The last push to the repository was on 2026-09-09. Frequent releases are good for fixes and bad for pinned scripts. The code style tooling is also pinned in-repo: pyproject.toml configures ruff with a line length of 119 and a select list including C, E, F, I, W, CPY, D417 and DOC, and the Makefile exposes make style and make quality. If you contribute a custom task, expect to satisfy those rules.

There is no migration guide in the README, and the README does not document rollback or version-pinning procedure. Treat the version number in your requirements file as the contract.

Editorial conclusion

Adopt Lighteval if you need to run a named benchmark such as gsm8k, mmlu or gpqa:diamond against a model you either serve yourself or hold in memory, and you are working on Linux or macOS. Skip it if your evaluation must run on Windows, since the README states the package is completely untested there, or if you need a stable released API rather than a package whose pyproject.toml still declares version 0.13.1.dev0. Before committing, verify three things: that the backend you intend to use (vllm, sglang, accelerate, nanotron, or an endpoint) is installed with the matching extra, that your task name resolves in the Open Benchmark Index space, and that you have run hf auth login if you plan to push results to the Hub.

Frequently asked questions

How do I use Lighteval to evaluate a model?

Install it with pip install lighteval, then run the CLI with a model string and a task name, for example lighteval eval "hf-inference-providers/openai/gpt-oss-20b" gpqa:diamond. The README also documents a Python API for models already loaded in memory, built around EvaluationTracker, PipelineParameters and Pipeline.

How does Lighteval compare with lm-evaluation-harness?

The README does not compare them directly. The architectural difference visible in the documentation is that Lighteval exposes several serving backends as separate subcommands (inspect-ai, accelerate, nanotron, vllm, sglang, and endpoint variants for Inference Endpoints, TGI, LiteLLM and inference providers), so evaluating a model that is served remotely is a supported path rather than a custom wrapper.

What techniques does Lighteval use for evaluating a language model?

Lighteval resolves a task name to a dataset, prompt template and metric, then generates completions through whichever backend you selected. The README states that sample-by-sample results are saved so you can inspect individual cases, and that custom tasks and custom metrics can be defined when the built-in ones do not fit.

How do you evaluate LLM model performance with Lighteval?

Pick a backend subcommand that matches where your model lives, then pass a model identifier and a task name. The README's example uses lighteval eval with a hosted model and gpqa:diamond, and results are written to the output_dir set on the EvaluationTracker.

What metrics are used to evaluate LLMs as judges?

The README does not describe LLM-as-judge metrics. It documents a metrics list page and a guide for adding a new metric, and lists task families such as instruction following and dialogue, but it does not specify which metrics apply to judge-style evaluation.

Official sources

  1. huggingface/lighteval on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/huggingface-lighteval.svg)](https://hysenlabs.com/projects/huggingface-lighteval)