VLM2Vec and MMEB-V3: a benchmark and training codebase for omni-modality embeddings
This repo contains the code for "VLM2Vec / MMEB" [ICLR 2025], "VLM2Vec-V2 / MMEB-V2" [TMLR 2026], and "MMEB-V3" [COLM 2026]
At a glance
- What is it?
- VLM2Vec is the TIGER-AI-Lab repository behind the MMEB benchmark line, now at MMEB-V3 with 190 tasks across text, image, video, audio, visual documents and agent scenarios. It is evaluation infrastructure for people training or selecting multimodal retrievers, not a drop-in embedding service.
- Who is it for?
- Adopt VLM2Vec if you are training or selecting a multimodal embedding model and need a task suite that covers audio, agent and text retrieval, not just image-text pairs. Do not adopt it if you want a hosted embedding endpoint or a small pip-installable library; this is a research harness with a heavy dependency set and a multi-hundred-gigabyte data preparation step.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 9 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What VLM2Vec is for, and who actually needs it
The repository hosts code and a data interface for MMEB-V3, described in the README as a benchmark for evaluating omni-modality embedding models across text, image, video, audio, visual document and agent-centric retrieval. It also carries the training entrypoints for the VLM2Vec models themselves, plus the MMEB-V1 and MMEB-V2 evaluation data folded into one root.
The audience is narrow and specific. If you are training a multimodal retriever and want to know whether it follows an instruction like "retrieve the audio clip matching this description" rather than just ranking images against captions, this is the harness that measures that. If you are choosing between published embedding models for a RAG pipeline that mixes documents, frames and audio, the leaderboard and task breakdown give you a comparison axis that single-modality MTEB-style suites do not.
What it is not: an embedding server, a Python library you import to embed a string, or a managed API. There is no inference daemon in the repository layout. You bring a model checkpoint, point eval.py at it, and read task scores.
How MMEB-V3 extends the earlier benchmark versions
MMEB-V3 adds 111 new tasks on top of MMEB-V2, reaching 190 tasks in total. The README groups the additions into three categories: audio tasks (audio classification, cross-modal audio retrieval, audio temporal grounding), text retrieval (instruction-following, reasoning, long-context, multi-condition and general text retrieval), and agent tasks (tool retrieval, GUI control, agent memory retrieval).
The agent category is the one that changes the character of the benchmark. Tool retrieval and GUI control are not similarity search over natural media; they are retrieval over structured action spaces where the correct answer depends on an interface definition. That makes MMEB-V3 a test of instruction following under explicit task constraints, which is the framing the README uses: diagnosing whether models can reliably follow modality-specific instructions.
OmniSET, short for Omni-modality Semantic Equivalence Tuples, is a separate diagnostic component. It groups semantically equivalent instances across text, image, video and audio so that modality effects can be studied in a controlled way. That is a different kind of measurement from a task score. A model can rank well on individual retrieval tasks and still fail to place an equivalent image, video and audio instance at the same point in the space. OmniSET is the part of the release aimed at that question.
Installing the evaluation harness and running your first task
Dependencies live in requirements.txt at the repository root and include torch, transformers pinned to 4.52.3, numpy pinned to 1.26.4, qwen-vl-utils[decord] pinned to 0.0.8, flash-attn, decord, ray and pytrec-eval. The pins matter: transformers and numpy are exact versions, so install into a fresh environment rather than upgrading an existing one. The flash-attn line carries an inline comment pointing at the upstream flash-attention README for accelerated installation, which is the repository telling you that this dependency needs its own build step.
Data comes first. The MMEB-V3 Hugging Face dataset ships as compressed assets plus lightweight metadata, and the README's flow is to download it with the hf CLI into a single root:
export MMEB_V3_ROOT=/path/to/MMEB-V3
hf download VLM2Vec/MMEB-V3 \
--repo-type dataset \
--local-dir $MMEB_V3_ROOTAfter that, a setup script materializes the evaluation-ready layout. It is described as idempotent, skipping directories marked with .done and safely extracting tar and zip files:
python experiments/public/data/dataset_setup_v3.py --root $MMEB_V3_ROOT
python experiments/public/data/dataset_setup_v3.py --root $MMEB_V3_ROOT --check-onlyThe --check-only pass is the one to run before anything else. It verifies the expected final layout instead of extracting. If your download omitted image-query, the README says to pass an existing local copy with --image-query-source /path/to/image-query; otherwise no extra argument is needed.
Evaluation then goes through eval.py, with the model selected by --model_backbone and --model_name. The README's example evaluates image tasks with Omni-Embed-Nemotron:
CUDA_VISIBLE_DEVICES=0 python eval.py \
--pooling mean \
--normalize true \
--per_device_eval_batch_size 8 \
--dataloader_num_workers 1 \
--model_backbone nvomniembed \
--model_name /path/to/omni-embed-nemotron-3b \
--dataset_config experiments/public/eval/image.yaml \
--encode_output_path exps/vlm2vec/omni-embed-nemotron-3b/image \
--data_basedir $MMEB_V3_ROOTExpect encoded outputs under the --encode_output_path directory and scores computed from them. The README notes the same entrypoint handles adapted omni-modality baselines such as e5_omni and lco_omni by changing --model_backbone, --model_name, pooling and related flags. The README does not document a resume path for an interrupted evaluation run.
The dataset layout is the real installation cost
The setup script is not the hard part. The hard part is that MMEB-V3 is a large multi-archive corpus, and the README spends most of its dataset section on the directory contract because getting it wrong is the most common failure.
Archives arrive under _tasks directories holding raw compressed assets. The setup script materializes them into -tasks directories, which is what the evaluation code consumes. So image_tasks/mmeb_v1.tar.gz becomes image-tasks/MMEB/, and the video, visual-document, GUI, memory, text and tool archives follow the same pattern. The README is explicit that you should not upload already materialized directories when the corresponding archives are present, naming video-tasks/frames/video_cls/, video-tasks/frames/video_ret/, video-tasks/frames/video_mret/, video-tasks/frames/video_qa/, visdoc-tasks/data/ and visdoc-tasks/images/ as examples. Duplicating both forms wastes space and, more importantly, makes it ambiguous which copy the harness reads.
OmniSET has its own archive. If the release contains omniset.tar.gz, the setup script extracts it into MMEB-V3/omniset/, and manual extraction is documented as equally valid:
tar -xzf $MMEB_V3_ROOT/omniset.tar.gz -C $MMEB_V3_ROOTThe expected omniset/ directory holds omniset.jsonl, catalog.jsonl and media subdirectories for val2014 images, videos, audios and frames_omni. If those files are absent, the diagnostic component has nothing to group, and the failure will look like an empty or missing task rather than a broken installation.
Where VLM2Vec is the wrong tool
The clearest limitation is scope mismatch. If you need to embed a few thousand product images and search them by text, MMEB-V3 is enormous overkill. The dependency list alone, with flash-attn requiring a build and ray pulled in for distributed execution, is a heavier environment than a single-modality retrieval job justifies. A CLIP-style model with a sentence-transformers wrapper will get you there with a fraction of the setup.
The second limitation is that the repository is an evaluation and training harness, not a serving stack. Nothing in the layout suggests an HTTP endpoint, batching server or index management layer. If your requirement is latency and throughput in production, this codebase does not address it, and the README does not claim to.
The third is data preparation. The instructions assume you can host a root directory containing image, video, audio, visual-document, GUI, memory, text and tool assets, with video frames extracted into per-task directories. That is a storage and I/O commitment, and the setup script's idempotent .done markers mean a partially extracted archive can be skipped on a re-run. If you suspect a bad extraction, the .done marker is the thing to inspect, and the README does not document a reset flag for it.
How it differs from single-modality benchmark suites
The obvious comparison is MTEB and its multimodal extensions, which concentrate on text and image-text retrieval. MMEB-V3's distinguishing choice is breadth of modality with instruction conditioning attached to each task, plus the agent category that no text-only suite covers. Tool retrieval, GUI control and agent memory retrieval are retrieval problems where the query encodes a constraint and the candidate set is an interface, not a document.
A second comparison is against per-modality evaluation practice, where you measure image retrieval with one harness, video retrieval with another and audio with a third, then compare numbers that were never calibrated against each other. MMEB-V3's single --data_basedir argument and shared eval.py entrypoint are the design answer to that fragmentation. The cost is that you inherit one layout contract and one dependency set for all of it.
OmniSET is the part with no direct equivalent in the suites above. Grouping semantically equivalent instances across text, image, video and audio turns a leaderboard question into a representation question. That is a genuinely different measurement, and it is the reason to look at this repository even if you already run MTEB.
Maintenance, licence and what to verify before adopting
The repository is not archived, and the last push was on 2026-08-23, roughly a month before this writing. Two releases are listed: v1.0 on 2025-06-10 and v2.0.1 on 2025-06-30. The README describes MMEB-V3 as accompanying a COLM 2026 paper, MMEB-V2 as a TMLR 2026 paper, and the original VLM2Vec and MMEB work as ICLR 2025, so the codebase tracks a research publication cadence rather than a product release cycle. Pin the commit you evaluate against, because benchmark task definitions and the expected directory layout are exactly the kind of thing that shifts between paper versions.
The licence is Apache-2.0, which permits commercial use and modification with the usual notice and patent-grant conditions. That covers the repository code. It does not automatically cover the datasets, model checkpoints or third-party components referenced by the README, which live on Hugging Face and in third_party/ and may carry their own terms. Check those separately; this is not legal advice.
Upgrade cost is dominated by the version pins. transformers==4.52.3 and numpy==1.26.4 in requirements.txt mean that moving to a newer transformers for an unrelated reason can break the harness. Treat the environment as frozen per evaluation run. The changelog at CHANGELOG.md is the place to check what moved between v1.0 and v2.0.1 before you decide whether to re-run an existing evaluation.
Editorial conclusion
Adopt VLM2Vec if you are training or selecting a multimodal embedding model and need a task suite that covers audio, agent and text retrieval, not just image-text pairs. Do not adopt it if you want a hosted embedding endpoint or a small pip-installable library; this is a research harness with a heavy dependency set and a multi-hundred-gigabyte data preparation step. Before committing, run the dataset setup script with --check-only against your extracted root and confirm the layout matches the expected image-tasks, video-tasks, visdoc-tasks, gui-tasks and omniset directories, because a missing archive surfaces as a task failure rather than an install error.
Frequently asked questions
What is the difference between VLM2Vec and MMEB?
VLM2Vec is the model line and MMEB is the benchmark it is evaluated on, and both live in this repository. The README states the repo contains the code for VLM2Vec and MMEB, with MMEB-V3 now covering 190 tasks across text, image, video, audio, visual document and agent retrieval.
How do I install VLM2Vec and run an evaluation?
Install the packages from requirements.txt, download the MMEB-V3 dataset with the hf CLI into a root directory, then run experiments/public/data/dataset_setup_v3.py against that root. Evaluation runs through eval.py with --data_basedir pointing at the same root and a model selected via --model_backbone and --model_name.
Can I use an LLM as an embedding model with VLM2Vec?
The repository evaluates embedding backbones rather than chat models, and the README's example uses --model_backbone nvomniembed with a local Omni-Embed-Nemotron checkpoint. It also notes adapted omni-modality baselines such as e5_omni and lco_omni, selected by changing --model_backbone and --model_name.
What is OmniSET in MMEB-V3?
OmniSET stands for Omni-modality Semantic Equivalence Tuples, a diagnostic component that groups semantically equivalent instances across text, image, video and audio. The README describes it as designed for controlled analysis of modality effects and instruction-conditioned cross-modal retrieval behavior.
Does the MMEB-V3 dataset need manual extraction?
The setup script handles extraction, and the README says it is idempotent, skipping directories marked with .done and safely extracting tar and zip files. Manual extraction is also documented as valid for the OmniSET archive via tar -xzf into the MMEB-V3 root.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/tiger-ai-lab-vlm2vec)