Library / SDK
embeddings-benchmark/mteb avatar
embeddings-benchmark/mteb

MTEB: Running the Embedding Benchmark Locally

MTEB: State-of-the-art evaluation of embeddings across languages and modalities

3,439 stars710 forksPythonApache-2.0

At a glance

What is it?
MTEB is a Python package that evaluates text and multimodal embedding models against named tasks and benchmarks. It is aimed at engineers choosing a retrieval or embedding model, and it ships both a Python API and a CLI.
Who is it for?
Adopt MTEB when you need comparable numbers across several candidate embedding models on tasks you can name, and you are prepared to download datasets and spend GPU time. Do not adopt it as a leaderboard shortcut: the hosted leaderboard ranks models, but your corpus and query distribution are not in it.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

The problem MTEB solves: choosing an embedding model without a leaderboard

Embedding models are interchangeable at the API level and very different in behaviour. Swapping one for another changes retrieval quality, index size, inference latency and language coverage, and none of that is visible from a model card. MTEB's answer is to fix the task definitions, the splits and the metric, then let any model be scored against them. The repository describes itself as a "Multimodal toolbox for evaluating embeddings and retrieval systems", and its topic list covers retrieval, reranking, clustering, classification, semantic search and bitext mining. That breadth is the point: a model that wins on semantic similarity may lose on classification, and MTEB reports the per-task numbers rather than a single figure you have to trust. The audience is narrow but real. If you are picking an encoder for a search index, a deduplication pipeline or a RAG store, MTEB gives you a reproducible way to compare candidates on tasks that at least resemble yours. If you already know your model, it gives you nothing.

How evaluation actually runs: models, tasks, and a results folder

The mechanism is a registry plus a runner. Models are resolved through mteb.get_model, which the README notes falls back to SentenceTransformer(model_name) when the model is not implemented in MTEB. Tasks are resolved by name through mteb.get_tasks, and task names carry a version suffix, as in Banking77Classification.v2. Calling mteb.evaluate(model, tasks=tasks) encodes the task data with the model, computes the task's metrics and returns results. The CLI wraps the same path and writes to disk. Results are stored in a results folder, and the pyproject dependencies include polars, which is the format the project uses for tabular result data and for the leaderboard cache. Under the hood the package depends on datasets, sentence_transformers, transformers, torch, scikit-learn, scipy and pytrec_eval_terrier, so retrieval metrics are computed by the Terrier evaluation library rather than reimplemented. The Dockerfile in the repository shows the same pipeline used to serve the leaderboard: a stage pre-warms the mteb/results parquet cache and persists a per-benchmark split to ~/.cache/mteb/leaderboard/ so the first request skips roughly 30 seconds of cold work. That cache path is worth knowing about even if you never run the server, because it is where repeated evaluations accumulate.

Installing MTEB and running your first evaluation

The README gives two install paths, pip and uv. Both install the mteb package; nothing else is required to start. Note the Python constraint in pyproject.toml: requires-python is >=3.10,<3.15, so a very new interpreter will be rejected at install time.

bash
pip install mteb

If you already use uv, the README shows the equivalent as a project dependency rather than a global install.

bash
uv add mteb

The shortest working evaluation is a few lines. The README uses sentence-transformers/all-MiniLM-L6-v2 against a single classification task, which is a reasonable first run because the model is small and the task is narrow.

python
import mteb
from sentence_transformers import SentenceTransformer

model_name = "sentence-transformers/all-MiniLM-L6-v2"
model = mteb.get_model(model_name)
tasks = mteb.get_tasks(tasks=["Banking77Classification.v2"])
results = mteb.evaluate(model, tasks=tasks)

Expect the first run to spend most of its time downloading the task dataset and the model weights, not encoding. If you prefer a shell, the CLI does the same thing and writes output to a folder you name.

bash
mteb run \
    -m sentence-transformers/all-MiniLM-L6-v2 \
    -t "Banking77Classification.v2" \
    --output-folder results

After it finishes, the results folder holds the scores for that task, and you can point the same command at other task names to build up a comparison. The CLI documentation at docs.mteb.org covers flags beyond -m, -t and --output-folder; the README only demonstrates those three.

Where MTEB gets in the way

The dependency set is heavy. torch, transformers, sentence_transformers, datasets, scikit-learn, scipy and polars are all required, and the install extras used by the project's own Makefile are heavier still: bm25s, image, audio, leaderboard, faiss-cpu, github and api are separate extras, and the test target runs with all of them. If you want a lightweight scoring script in a small container, MTEB is the wrong shape. Dataset downloads are the other cost. Tasks are pulled through the datasets library, and a broad benchmark means many datasets, each with its own licence and size. The README does not document an offline mode or a way to run against a private corpus without packaging it as a task, so if your data cannot leave your network, you are looking at the contributing path for adding a dataset rather than a config flag. Versioning deserves attention too. Task names carry suffixes like .v2, which means a score is only comparable to another score computed on the same task version. Reporting "MTEB score" without the task list and version is not a reproducible claim, and the project does not prevent you from doing it.

MTEB against BEIR and hand-rolled evaluation sets

BEIR is the obvious comparison point, and the difference is scope and packaging. BEIR is a retrieval benchmark: a fixed set of zero-shot IR datasets with nDCG as the headline metric. MTEB started from the same retrieval lineage but adds classification, clustering, reranking, pair classification, semantic textual similarity and bitext mining, and it now covers multilingual and multimodal evaluation. Practically, that means MTEB is a framework you install and extend, with a task registry and a model registry, whereas BEIR is closer to a dataset collection you evaluate against. The trade-off runs the other way too. Because MTEB spans many task types, its aggregate figures mix metrics that are not directly comparable, and a single ranking number hides which task drove it. A team that only cares about retrieval may find BEIR's narrower framing easier to reason about, and a team that wants a bespoke metric on a private corpus will end up writing its own harness either way. MTEB's advantage is that when you do add a task, the runner, caching and result format are already there.

Maintenance, licence and the cost of keeping up

The project is not archived, and the last push was on 2026-09-20. Releases are frequent: 2.21.0 landed on 2026-09-14, with 2.20.13 earlier the same day and 2.20.12 on 2026-09-09. That cadence is good for correctness and bad for pinning. If you report numbers in a paper or an internal decision doc, pin the mteb version alongside the task versions, because a task definition can change between releases and the same model can score differently. For development, the Makefile shows the intended workflow: uv sync --extra image --group dev for a working environment, make test for the fast suite, and make pr to run lint plus tests before submitting. Note that the default test target excludes test_datasets, leaderboard_stability and test_reference_models, so the fast suite is not the whole story. The licence is Apache-2.0, declared in pyproject.toml with license-files pointing at LICENSE. That is a permissive licence, but it covers the MTEB code and not the datasets it downloads, each of which carries its own terms. Check the dataset licences separately before you ship anything derived from them; this is not legal advice.

Editorial conclusion

Adopt MTEB when you need comparable numbers across several candidate embedding models on tasks you can name, and you are prepared to download datasets and spend GPU time. Do not adopt it as a leaderboard shortcut: the hosted leaderboard ranks models, but your corpus and query distribution are not in it. Before committing, verify three things: that the task you want exists under the exact version suffix you plan to report, that the model is either implemented in MTEB or loadable through SentenceTransformer, and that the machine running the evaluation has disk and time for the datasets and the encoder. If those hold, a single mteb run command gives you a results folder you can keep under version control.

Frequently asked questions

What is the MTEB benchmark?

MTEB is a benchmark and a Python package for evaluating embedding and retrieval models. The repository describes it as a multimodal toolbox covering tasks such as retrieval, reranking, clustering, classification, semantic search and bitext mining, and it is distributed on PyPI as mteb.

What does the MTEB score mean?

The package returns per-task results rather than one number, and task names carry version suffixes such as Banking77Classification.v2. A score is therefore tied to the task and its version, so a figure quoted without the task list is not comparable to another run.

What are the best benchmark embedding models?

The package itself does not rank models. The project publishes an interactive leaderboard on Hugging Face, and locally the ranking you get depends on which tasks you pass to mteb.get_tasks and which model you load with mteb.get_model.

How to use mteb?

Install it with pip install mteb, load a model through mteb.get_model, select tasks with mteb.get_tasks and call mteb.evaluate. The same evaluation is available from the command line as mteb run with -m, -t and --output-folder.

Official sources

  1. embeddings-benchmark/mteb on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/embeddings-benchmark-mteb.svg)](https://hysenlabs.com/projects/embeddings-benchmark-mteb)
Community notes

Community notes