# lmms-eval: A Multimodal Evaluation Toolkit for Text, Image, Video and Audio Models

> lmms-eval is a Python framework that runs one evaluation pipeline across text, image, video and audio benchmarks. It installs from PyPI or Git, ships 100+ tasks and 30+ model backends, and its licence metadata is inconsistent between the README and the package manifest.

**EvolvingLMMs-Lab/lmms-eval** — One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks

- Repository: https://github.com/EvolvingLMMs-Lab/lmms-eval
- Website: https://www.lmms-lab.com
- Stars: 4,437 · Forks: 662
- Language: Python
- License: NOASSERTION
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/evolvinglmms-lab-lmms-eval

## The fragmentation problem lmms-eval was built to remove

Multimodal evaluation has a specific failure mode that text-only evaluation mostly avoids. A vision benchmark needs images decoded, resized and sometimes sampled from video; an audio benchmark needs waveforms at the right sample rate; a reasoning benchmark needs answer extraction from free-form text. Each of those steps is a place where two teams can diverge. The README states the problem directly: scattered datasets, inconsistent post-processing, and single-number accuracy scores that hide whether a gain is real or random. It also claims that two teams evaluating the same model on the same benchmark routinely report different results.

That claim is the whole justification for the project. lmms-eval is aimed at model teams who need to compare checkpoints against published baselines, and at researchers who need to reproduce someone else's reported number. It is not aimed at someone who wants a quick quality score for a single image captioning model. The scope is text, image, video and audio in one harness, which is unusual: most evaluation toolkits pick one modality and do it well. The README frames the goal as mapping the border of what models can do, and the three stated principles are reproducible, efficient and trustworthy, with the third backed by confidence intervals, clustered standard errors and paired comparisons rather than accuracy alone.

## How the evaluation pipeline works

The README points to two documentation sections, How the Evaluation Pipeline Works and Why it's Efficient and Trustworthy, but the README text itself only describes the shape of the system. What is visible from the repository layout is a task-oriented design: configs/ holds task configuration, lmms_eval/ holds the package including a models/ directory, and docs/advanced/current_tasks.md lists the tasks. The quickstart command selects a model by name (--model qwen2_5_vl), passes model arguments (--model_args pretrained=Qwen/Qwen2.5-VL-3B-Instruct), and selects one or more tasks (--tasks mme). That means tasks and models are separate registries combined at runtime.

The efficiency claims are specific enough to check. The README states that v0.6 delivered roughly 7.5x throughput over v0.5 through a standalone HTTP eval server, and that v0.7 added a video I/O overhaul with TorchCodec described as up to 3.58x faster. It also names async serving and adaptive batching as the mechanisms that keep GPUs saturated. Those are release-note claims, not independently verified numbers. The v0.6 release also introduced statistically grounded results, which is the part that changes how you read output: a single accuracy figure becomes an estimate with an interval, and paired comparisons become possible between two models on the same task set.

## Installing lmms-eval and running a first evaluation

The README recommends uv for package management so that every developer resolves the same versions. The first step is installing uv itself, which the README does with a shell installer.

```bash
curl -LsSf https://astral.sh/uv/install.sh | sh
```

After that, clone the repository and install the package in editable mode with the all extra, which pulls the full dependency set including torch, torchvision, transformers, datasets and av.

```bash
git clone https://github.com/EvolvingLMMs-Lab/lmms-eval
cd lmms-eval
uv pip install -e ".[all]"
```

The README notes that uv sync is an alternative, creating or updating the environment from uv.lock. To run anything, prefix it with uv run, for example uv run python -m lmms_eval --help.

The first real evaluation in the README runs Qwen2.5-VL on the MME task with a limit of 8 samples. The limit keeps the run short; batch_size of 1 keeps memory low.

```bash
python -m lmms_eval \
  --model qwen2_5_vl \
  --model_args pretrained=Qwen/Qwen2.5-VL-3B-Instruct \
  --tasks mme \
  --batch_size 1 \
  --limit 8
```

The README says that if it prints metrics, your environment is ready. That is a low bar and a useful one: a failure here is almost always a model loading or dependency problem, not a task logic problem. There is also a Git-only install path using uv venv and uv pip install git+https://github.com/EvolvingLMMs-Lab/lmms-eval.git, with a warning that you might need to add and include your own task yaml if you install that way. The README does not document rollback for either install path.

## Where lmms-eval breaks down or is the wrong tool

The first limitation is the licence. The repository metadata reports NOASSERTION, pyproject.toml declares license = "MIT" with license-files = ["LICENSE"], and the README does not state a licence at all. For a toolkit whose output feeds published comparisons, that inconsistency is something to resolve with the maintainers before you build a pipeline on it. This is a factual gap, not a legal opinion.

The second is Python version coupling. pyproject.toml sets requires-python = ">=3.10" and pins av>=17.1.0,<18.0.0 with a comment that PyAV 18 requires Python 3.11 or newer. That pin exists to keep 3.10 working, and it means you are held on an older av release until the project raises its floor. The README's Git install example uses uv venv --python 3.12, so 3.12 is the tested path in that example, not 3.10.

The third is scope. If you evaluate text-only models, the multimodal machinery (video decoding, image preprocessing, audio handling) is overhead you pay for and never use. If you need a fixed, stable grading contract that does not move between releases, this is the wrong tool: v0.7 changed log output to flattened JSONL and added a new generation path called generate_until_agentic, and v0.6 added an HTTP eval server. Each of those changes alters how results are produced or stored. The README also notes that torch and CUDA version differences caused small variations when reproducing LLaVA-1.5 paper results, which is a candid admission that bit-level reproducibility across environments is not guaranteed.

## How lmms-eval differs from VLMEvalKit

VLMEvalKit is the obvious comparison and the one people search for. The difference in approach is breadth of modality. VLMEvalKit is built around vision-language benchmarks. lmms-eval's README frames the project as covering text, image, video and audio in one pipeline, and the release history backs that: v0.3 added audio evaluation with Qwen2-Audio and Gemini-Audio, v0.5 expanded audio with 50+ benchmark variants across audio, vision and reasoning, and v0.7 added 25+ tasks across 8 domains plus safety and red-teaming baselines.

The second difference is the statistical posture. lmms-eval's README states that v0.6 introduced confidence intervals and paired t-tests as first-class output, and that the project is doing ongoing research into evaluation methodology. That matters if you are comparing two checkpoints that differ by a fraction of a point, because without an interval you cannot tell a real gain from noise. If you only need a leaderboard submission number, that machinery is extra work with no payoff.

The third difference is operational shape. lmms-eval ships a standalone HTTP eval server as of v0.6, which turns evaluation into a service you can call rather than a script you run. That is a real architectural choice, and it is the reason the project can claim throughput improvements between releases.

## Maintenance, upgrade cost and licence status

The repository is not archived and the last push was on 2026-09-19. Releases are frequent: v0.7.3 on 2026-08-28, v0.7.2 on 2026-06-24, v0.7.1 on 2026-03-15. That cadence is a maintenance cost as much as a benefit. A toolkit that ships minor releases every few months and changes log format and adds generation paths will move under you. If you pin to a version and stay there, you avoid churn but lose new tasks and the video I/O improvements. If you track main, you should expect to re-run baselines after upgrades, because the README itself notes that torch and CUDA version differences cause small variations in results.

Dependencies are heavy and pinned with care in places: torch>=2.1.0, transformers>=4.39.2, av>=17.1.0,<18.0.0, wandb==0.25.0 (an exact pin, unusual among the rest), and chess>=1.11.2,<2 for MET-Bench board-state scoring. The exact wandb pin is worth noting if you already run a different wandb version in the same environment.

On licence: pyproject.toml says MIT and lists a LICENSE file, the README does not mention a licence, and repository metadata reports NOASSERTION. Those three sources disagree, and that is the thing to verify with the maintainers before you redistribute the package or ship it inside a product. The README does not document rollback, and it does not document what happens to cached results when task definitions change between versions.

## Conclusion

Adopt lmms-eval if you need one harness that covers image, video and audio benchmarks and you want the same numbers every run, which is the reproducibility claim the project makes in its README. Do not adopt it if you only evaluate text-only LLMs, or if you need a stable, documented grading contract that will not shift between releases, since v0.7 explicitly reworked log output and added an agentic generation path. Before committing, verify the licence status: the README does not state a licence, pyproject.toml declares MIT with a LICENSE file, and the repository metadata reports NOASSERTION. Then run the quickstart command with --limit 8 and confirm that your model backend actually loads, because the model list and the task list are maintained separately and a task can exist without a working backend for the model you care about.

## FAQ

### How do I install lmms-eval?

The README recommends installing uv first, then cloning the repository and running uv pip install -e ".[all]" inside it. A Git-only alternative is uv venv followed by uv pip install git+https://github.com/EvolvingLMMs-Lab/lmms-eval.git, though the README warns you might need to add your own task yaml with that path.

### What is lmms-eval used for?

It is a multimodal evaluation toolkit that runs benchmarks across text, image, video and audio tasks through one pipeline. The README describes the goal as reproducible, efficient and trustworthy evaluation, with confidence intervals and paired comparisons rather than accuracy alone.

### How does lmms-eval compare with VLMEvalKit?

VLMEvalKit is built around vision-language benchmarks, while the lmms-eval README frames its scope as text, image, video and audio in one harness, with audio support added in v0.3 and expanded in v0.5. lmms-eval also treats confidence intervals and paired t-tests as part of its output, and ships a standalone HTTP eval server as of v0.6.

### Which models and tasks does lmms-eval support?

The README links to 100+ tasks in docs/advanced/current_tasks.md and 30+ model backends in the lmms_eval/models directory. The quickstart example uses the qwen2_5_vl backend with pretrained=Qwen/Qwen2.5-VL-3B-Instruct on the mme task.

### Does lmms-eval give the same numbers every time?

The README states reproducibility as a principle: same model, same benchmark, same numbers, every time. It also notes that torch and CUDA version differences caused small variations when reproducing LLaVA-1.5 paper results, so identical results across different environments are not guaranteed.

### What licence does lmms-eval use?

The sources disagree: pyproject.toml declares MIT and lists a LICENSE file, the README does not state a licence, and the repository metadata reports NOASSERTION. Verify this with the maintainers before redistributing the package.

## Sources

- [EvolvingLMMs-Lab/lmms-eval on GitHub](https://github.com/EvolvingLMMs-Lab/lmms-eval)
- [Issues](https://github.com/EvolvingLMMs-Lab/lmms-eval/issues)
- [Project website](https://www.lmms-lab.com)
- [README](https://github.com/EvolvingLMMs-Lab/lmms-eval/blob/main/README.md)
- [Releases](https://github.com/EvolvingLMMs-Lab/lmms-eval/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/evolvinglmms-lab-lmms-eval
