# MinishLab/model2vec: static embeddings distilled from a sentence transformer

> Model2Vec turns any sentence transformer into a static embedding model, cutting size by up to 50x and CPU inference cost by up to 500x. The trade-off is a small drop in retrieval quality, and the documentation is honest about that.

**MinishLab/model2vec** — Fast State-of-the-Art Static Embeddings

- Repository: https://github.com/MinishLab/model2vec
- Website: https://minish.ai/packages/model2vec/introduction
- Stars: 2,215 · Forks: 127
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/minishlab-model2vec

## The problem model2vec solves: sentence transformers are too heavy for CPU-bound serving

A sentence transformer produces one embedding per input by running a full transformer forward pass over every token. That is fine on a GPU and painful on a CPU, where per-request latency scales with sequence length and model depth. Model2Vec's answer is to remove the transformer from the inference path entirely. The README states that Model2Vec reduces model size by a factor of up to 50 and makes models up to 500 times faster on CPU, with "a small drop in performance". The intended audience is developers who need embeddings for classification, retrieval, clustering or RAG but cannot pay transformer inference cost per document. The project ships its own pre-trained models under the minishlab/potion-* names, and the smallest of these is described as roughly 8 MB on disk, which the README calls the smallest model on MTEB at the time of writing. That size claim is what makes the library interesting for browser and edge deployments, not just server-side batch jobs.

## How distillation works: a vocabulary forward pass, then post-processing

The mechanism is unusual because it needs no training data. The README states the core idea plainly: forward pass a vocabulary through a sentence transformer model, creating static embeddings for the individual tokens. A vocabulary is a list of tokens, not a corpus, so the expensive part is one pass over the tokenizer's vocabulary rather than a training loop over sentences. After that pass, the README mentions "a number of post-processing steps" that produce the flagship models, plus an optional pre-training step to boost performance further. The deepdive is deferred to the official documentation rather than the README, which is a gap worth noting if you want to reproduce the post-processing yourself. Inference is then a lookup: each token maps to a vector, and the model combines token vectors into a sentence vector. The README exposes this directly through encode_as_sequence, which returns the per-token vectors instead of the pooled sentence vector. That method is the clearest evidence of the static design: there is no attention over context, so a token's vector does not change with the sentence it appears in. Subword handling exists, and the README compares it to BPEmb, claiming better performance at the same kind of subword granularity.

## Installing model2vec and making a first embedding

The base package is deliberately thin. The README says the only major dependency is numpy, and pyproject.toml lists jinja2, joblib, numpy, safetensors, tokenizers>=0.20, tqdm and huggingface-hub>=1.0.0. Python 3.10 or newer is required. Install it with pip:

```bash
pip install model2vec
```

Loading a pre-trained model from the HuggingFace hub and encoding two strings is the shortest useful program. The README uses the potion-base-32M model for this example:

```python
from model2vec import StaticModel

model = StaticModel.from_pretrained("minishlab/potion-base-32M")

embeddings = model.encode(["It's dangerous to go alone!", "It's a secret to everybody."])

token_embeddings = model.encode_as_sequence(["It's dangerous to go alone!", "It's a secret to everybody."])
```

encode returns one vector per input string. encode_as_sequence returns the per-token vectors, which is what you want if you plan to pool differently or inspect subword behaviour. If you would rather build your own model than use a potion checkpoint, install the distillation extras and run distill against a sentence transformer. The README states this takes about 30 seconds on a CPU and needs no dataset:

```bash
pip install model2vec[distill]
```

```python
from model2vec.distill import distill

m2v_model = distill(model_name="BAAI/bge-base-en-v1.5")

m2v_model.save_pretrained("m2v_model")
```

The result is a directory you can load with StaticModel.from_pretrained("m2v_model") and, per the README, push to the hub with the familiar push_to_hub call. The distillation extras pull in torch and transformers<5.4.0 alongside skeletoken, so the install is heavier than the base package even though inference is not.

## Fine-tuning a classifier on top of a static model

The training extras add torch and expose StaticModelForClassification, which wraps a static model with a trainable head. The README's example loads a dataset with the datasets library, initialises the classifier from potion-base-32M, calls fit on the training split, then evaluate on the test split. Both single-label and multi-label datasets are supported according to the README. This is the most practical entry point for teams whose actual task is classification rather than retrieval, because it avoids the question of whether the static embeddings are good enough for nearest-neighbour search: the head learns what the frozen vectors cannot express on their own. The cost is that you are back to a training loop, and the static backbone will not adapt to your domain the way a fine-tuned transformer would. The README points to separate training documentation for advanced usage, and does not describe early stopping, class weighting or how the head handles long inputs in the README itself.

## Where model2vec is the wrong tool

The README is direct about the central limitation: there is a small drop in performance relative to the sentence transformer you distil from. If your application is already at the edge of acceptable retrieval quality, that drop can be the difference between working and not working, and no amount of speed compensates. The second limitation is structural. Because token vectors are static, the model cannot disambiguate a word by its context. Polysemous terms, negation and word order all collapse into the same bag of vectors. Tasks that depend on those distinctions, such as natural language inference over minimal pairs or fine-grained sentiment, are poor candidates. Third, the distillation path is data-free by design, which means you cannot steer the resulting model toward your domain by feeding it your corpus; the only domain adaptation offered is fine-tuning a classifier head. Fourth, the repository's pyproject.toml still carries the classifier "Development Status :: 4 - Beta", so the API surface should be treated as movable. The README does not document rollback or how to pin a distilled model to a specific upstream revision, which matters if you distil from a hub model that later changes.

## How it compares with sentence-transformers and GLoVe

The most direct alternative is the source model itself: sentence-transformers, which model2vec distils from. The difference in approach is that sentence-transformers runs a transformer per input and therefore produces contextual embeddings, while model2vec runs a lookup and produces static ones. That single change accounts for the size and speed figures the README quotes, and it also accounts for the quality drop. If you have a GPU and your throughput is adequate, there is no reason to move. The second comparison the README draws is against older static embeddings, GLoVe and BPEmb. GLoVe is trained on co-occurrence statistics over a corpus; BPEmb learns subword embeddings from a corpus as well. Model2Vec's distillation needs no data at all, only a vocabulary and a model, which is a genuinely different production path: you can generate a static model from a checkpoint you already trust in about half a minute on a CPU. Whether the resulting vectors beat GLoVe on your task is an empirical question, and the README points to a results directory rather than making a task-specific claim. For subword granularity specifically, the README claims better performance than BPEmb at comparable granularity.

## Licence, packaging and maintenance cost

The repository is MIT licensed, and pyproject.toml declares the MIT classifier with the licence file referenced under [project]. MIT is permissive, so redistribution inside a commercial product is not the obstacle here. The obstacle is upstream: the base model you distil from carries its own licence, and model2vec does not change that. Distilling BAAI/bge-base-en-v1.5, as the README's example does, produces a derived artefact whose terms follow the base model, not model2vec's. Check the base model card before shipping. On maintenance, the last push was on 2026-09-10 and the most recent release listed is v0.9.0 from 2026-08-12, so the project is being changed on a roughly monthly cadence. The beta classifier in pyproject.toml means upgrades can move APIs; the Makefile shows the project's own test target runs pytest with coverage and ignores tests/integration, with separate integration targets for distillation baselines. That layout suggests distilled outputs are regression-tested against stored baselines, which is a useful signal if you plan to re-distil on every upstream model update.

## Conclusion

Adopt model2vec when you need embeddings that run on CPU at high throughput, when model size matters (edge, browser, container images), or when you want a classifier head on top of frozen embeddings. Skip it when your retrieval quality budget is tight and you cannot accept the documented small drop in performance versus the source sentence transformer, or when you need contextual token representations. Verify first: the licence of the base model you distil from is separate from model2vec's MIT licence, and the README does not document rollback or version pinning for distilled artefacts.

## FAQ

### Which is the best text embedding model?

The README claims that Model2Vec models outperform other static embeddings such as GLoVe and BPEmb by a large margin, and that its best model is the most performant static embedding model, but it points to a results directory rather than making a claim against contextual models.

### What are model embeddings?

In model2vec, embeddings are vectors produced by looking up static vectors for each token and combining them into a sentence vector. The README also exposes encode_as_sequence, which returns the per-token vectors directly.

### Is GPT an embedding model?

The README does not discuss GPT. It describes model2vec as a technique that turns a sentence transformer into a static embedding model, and the models it ships are the minishlab potion checkpoints on the HuggingFace hub.

### How do I train an embedding model?

Model2Vec does not train embeddings from a corpus. It distils them from an existing sentence transformer with no data, using distill(model_name=...) after installing the distillation extras, and the README says this takes about 30 seconds on a CPU. Fine-tuning is limited to a classification head via StaticModelForClassification.

## Sources

- [License: MIT](https://github.com/MinishLab/model2vec/blob/main/LICENSE)
- [MinishLab/model2vec on GitHub](https://github.com/MinishLab/model2vec)
- [Project website](https://minish.ai/packages/model2vec/introduction)
- [README](https://github.com/MinishLab/model2vec/blob/main/README.md)
- [Releases](https://github.com/MinishLab/model2vec/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/minishlab-model2vec
