Model or dataset
open-compass/VLMEvalKit avatar
open-compass/VLMEvalKit

VLMEvalKit: 220 models, 80 benchmarks, one command

Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks

4,418 stars771 forksPythonApache-2.0

At a glance

What is it?
VLMEvalKit, published on PyPI as vlmeval, is OpenCompass's Apache-2.0 evaluation toolkit for large vision-language models, enabling one-command evaluation across 220 plus models and 80 plus benchmarks with generation-based scoring by both exact matching and LLM-based answer extraction. Recent updates add thinking-mode response splitting, TSV prediction files for long outputs, and multi-node distributed inference through LMDeploy or vLLM.
Who is it for?
Use VLMEvalKit when a vision-language model needs benchmark numbers comparable to the published leaderboards, since its benchmark coverage, scoring methods and OpenVLM Leaderboard records form the de facto reference implementation. Prefer task-specific evaluation harnesses when a single custom benchmark dominates your needs.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

One command instead of twenty repositories

The problem VLMEvalKit solves is stated in one sentence, it enables one-command evaluation of large vision-language models on various benchmarks without the heavy workload of data preparation under multiple repositories. Every benchmark suite historically shipped its own evaluation code, its own data pipeline and its own quirks, so evaluating one model on ten benchmarks meant assembling ten toolchains. VLMEvalKit centralizes that, and its scoring is generation-based for all models, producing results under both exact matching and LLM-based answer extraction, the two scoring regimes whose differences matter for multiple choice benchmarks especially. The package is vlmeval on the Python side, Apache-2.0 licensed, trilingual in documentation with English, Chinese and Japanese readmes, and development is intense, with the repository pushed 2026-09-29.

Scale: 220 models, 80 benchmarks, video included

The support numbers are the headline, 220 plus LMMs and 80 plus benchmarks, and the news entries show both sides growing continuously. The 2025-05-24 update names the model wave it absorbed, InternVL3, Gemini-2.5-Pro, Kimi-VL, LLaMA4, NVILA, Qwen2.5-Omni, Phi4, SmolVLM2, Grok and more, alongside a benchmark wave including HLE-Bench, MMVP, OmniDocBench, OCR-Reasoning, MedXpertQA, Video-MMLU and the Spatial457 benchmark. Video understanding is a first class dimension with its own leaderboard, the OpenVLM Video Leaderboard, and the 2026-04-08 entry adds Video-MME-v2, described as an authoritative benchmark toward the next stage of video understanding evaluation. Domain-specific benchmarks arrive through contributors, physics reasoning through PhyX and SeePhys, medicine through MedXpertQA, documents through OmniDocBench and wildDoc.

Thinking mode: split before you score

The September 2025 update for models with thinking mode reflects how reasoning models broke evaluation assumptions. A new split_thinking function separates the model's reasoning from its answer, and the recommendation is emphatic, strongly recommended for models with thinking mode to ensure the accuracy of evaluation, activated with the environment variable SPLIT_THINK=True. By default the function parses content within think tags and stores it under a thinking key in the output, and for advanced customization a model can define its own split_think function, with the InternVL implementation named as the example. Without splitting, a chain of thought contaminates answer extraction, and the scoring libraries grade reasoning text as if it were the answer, which is precisely the failure this feature prevents.

Long responses: TSV over xlsx

The same update wave addressed long outputs with an infrastructure observation, individual cells in an xlsx file are limited to 32,767 characters, so prediction files for models generating long responses, exceeding 16k or 32k tokens, get silently truncated in spreadsheet format. The remedy is TSV prediction saving, enabled with PRED_FORMAT=tsv, and the recommendation mirrors the thinking-mode one in strength. The pairing of the two updates is coherent, reasoning models produce long chains of thought, long chains of thought overflow spreadsheet cells, and both fixes shipped in the same pull request numbered 1229, which also refined the can_infer_option and can_infer_text routing to LLM choice extractors, empirically improving multiple choice benchmark scores slightly.

Distributed inference through LMDeploy and vLLM

For large scale evaluation, the toolkit supports multi-node distributed inference using LMDeploy for the InternVL series, QwenVL series and LLaMA4, or vLLM for the QwenVL series and LLaMA4, activated by adding the use_lmdeploy or use_vllm flag to a custom model configuration in config.py. The framing is honest about when it matters, faster evaluations for large scale or thinking models, since thinking models generate orders of magnitude more tokens per question and dominate wall clock time on standard benchmarks. With the distributed path, the toolkit stops being a single-GPU research script and becomes a cluster workload, which is the operating mode behind the leaderboard numbers published for the biggest models.

Leaderboards as shared infrastructure

The evaluation results themselves are published as artifacts rather than only tables. The OpenVLM Leaderboard on Hugging Face Spaces hosts the interactive rankings, a companion video leaderboard extends the same to video models, the full detailed results download as a JSON file, and the evaluation records live as a Hugging Face dataset named OpenVLMRecords, so a paper or a model card can link to the exact underlying numbers. An OpenCompass leaderboard on the project's own site mirrors the rankings, and the toolkit's report is on arXiv for citation. This closes the loop most benchmarks leave open, the code that ran the numbers, the numbers themselves, and the leaderboard displaying them are all public and mutually consistent.

A dependency list that is the feature list

The requirements file reads as a map of what evaluation actually touches. torch, transformers and accelerate run the open source models, litellm pinned between 1.55 and 1.85 plus openai and google-genai drive the API models, decord handles video decoding, latex2sympy2-extended and math-verify grade mathematical answers, jieba supports Chinese text scoring, rdkit brings chemistry structure comparison, sentence_transformers and bert_score power semantic similarity metrics, and json_repair parses the malformed JSON that LLM-based extractors inevitably produce. The run.py entry point drives evaluation from the shell, scripts and tests surround it, and releases run on a slower track than commits, v0.2 in March 2025 and v0.3rc1 in June 2025, with the main branch carrying the active work.

Editorial conclusion

Use VLMEvalKit when a vision-language model needs benchmark numbers comparable to the published leaderboards, since its benchmark coverage, scoring methods and OpenVLM Leaderboard records form the de facto reference implementation. Prefer task-specific evaluation harnesses when a single custom benchmark dominates your needs. Before running, read the Quickstart for the run.py invocation patterns, enable SPLIT_THINK for models with thinking mode and PRED_FORMAT=tsv for models generating very long responses, both strongly recommended in the project's own updates, and consider the use_lmdeploy or use_vllm flags for large scale or thinking models where inference speed dominates wall clock time.

Frequently asked questions

What is VLMEvalKit?

VLMEvalKit is an open-source Apache-2.0 evaluation toolkit for large vision-language models, published as the vlmeval Python package, enabling one-command evaluation across more than 220 models and 80 benchmarks. It uses generation-based evaluation with both exact matching and LLM-based answer extraction, and backs the OpenVLM Leaderboard on Hugging Face.

How do you run VLMEvalKit?

Follow the Quickstart documentation for the run.py invocation patterns, installing the package from the repository and selecting your model and benchmarks. For thinking-mode models set SPLIT_THINK=True, for models with very long outputs set PRED_FORMAT=tsv, and for large scale evaluation add use_lmdeploy or use_vllm to your model configuration.

Does VLMEvalKit support video models?

Yes, video understanding is a supported dimension with a dedicated OpenVLM Video Leaderboard, video benchmarks downloaded optionally through ModelScope via the VLMEVALKIT_USE_MODELSCOPE flag, and benchmarks like Video-MME-v2, Video-MMLU and QBench-Video in the supported set, with decord handling video decoding.

Official sources

  1. License: Apache-2.0
  2. open-compass/VLMEvalKit on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/open-compass-vlmevalkit.svg)](https://hysenlabs.com/projects/open-compass-vlmevalkit)