# FastEmbed: ONNX embeddings for Python without the PyTorch install

> FastEmbed is a Python embedding library from Qdrant that runs ONNX Runtime instead of PyTorch. It covers dense, sparse, late-interaction, image and reranking models, and the trade-offs sit in its dependency pins and its model catalogue.

**qdrant/fastembed** — Fast, Accurate, Lightweight Python library to make State of the Art Embedding

- Repository: https://github.com/qdrant/fastembed
- Website: https://qdrant.github.io/fastembed/
- Stars: 3,221 · Forks: 253
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/qdrant-fastembed

## The problem FastEmbed targets: embedding generation without a deep learning stack

Most Python embedding code arrives attached to a training framework. Sentence-Transformers pulls in PyTorch, and the README states FastEmbed avoids that by using ONNX Runtime, so it does not require a GPU and does not download gigabytes of PyTorch dependencies. That single packaging decision is the project's identity. The intended user is someone who wants vectors inside an application, not someone training a model: a retrieval service, a RAG pipeline, a batch job that has to run in a constrained container.

The README names serverless runtimes like AWS Lambda as a fit, which follows from the dependency set rather than from a benchmark. The project is maintained by Qdrant, and the README says fastembed is supported by and maintained by Qdrant, so the library is aligned with a vector database vendor rather than being a general research toolkit. That alignment explains the model mix: retrieval-oriented text models, sparse models for hybrid search, and ColBERT-style late interaction.

## How FastEmbed works: ONNX Runtime, a model registry, and generators

The mechanism is consistent across modalities. A class such as TextEmbedding is constructed with a model name, which triggers a download through huggingface-hub, and the resulting object exposes an embed method that returns a generator. The README is explicit that embed returns a generator and that you convert it with list() when you want materialised vectors. This matters for memory: if you stream a large corpus through embed and write each vector out, you never hold the whole set in RAM.

Each modality has its own entry point. TextEmbedding covers dense vectors, SparseTextEmbedding returns SparseEmbedding objects with indices and values, LateInteractionTextEmbedding returns per-token matrices for ColBERT-style scoring, ImageEmbedding takes file paths, and LateInteractionMultimodalEmbedding handles ColPali-style document images with a separate embed_text path for the query. Reranking is a different package path, fastembed.rerank.cross_encoder, with TextCrossEncoder and a rerank method that scores a query against documents.

The extension point is the interesting part of the design. TextEmbedding.add_custom_model and TextCrossEncoder.add_custom_model let you register a model that is not in the built-in list by declaring its pooling type, normalization, dimension, model file and a ModelSource. The README notes the source can point at a Hugging Face repository or at a URL for private storage, and that model_file can select a different ONNX optimisation or quantisation of an already supported model. In practice this turns FastEmbed into a thin runner over ONNX artefacts you supply, which is a narrower promise than a training library but a much smaller surface to operate.

## Installing FastEmbed with pip and running a first embedding

Installation is a single pip command, with a separate package for GPU builds. The README gives both forms.

```bash
pip install fastembed

# or with GPU support

pip install fastembed-gpu
```

The default install pulls onnxruntime, tokenizers, huggingface-hub, numpy and a handful of smaller packages. The GPU variant is a distinct distribution name, so dependency resolution differs between them; do not mix the two in one environment.

The quickstart constructs a TextEmbedding with no arguments, which selects the default model. The README says this triggers the model download and initialization, and that the default is BAAI/bge-small-en-v1.5, producing 384-dimensional vectors.

```python
from fastembed import TextEmbedding

documents: list[str] = [
    "This is built to be faster and lighter than other embedding libraries e.g. Transformers, Sentence-Transformers, etc.",
    "fastembed is supported by and maintained by Qdrant.",
]

embedding_model = TextEmbedding()
embeddings_generator = embedding_model.embed(documents)
embeddings_list = list(embedding_model.embed(documents))
```

Expect the first call to hit the network and take noticeably longer than later ones. To pin the model explicitly, pass model_name, as in the dense example: TextEmbedding(model_name="BAAI/bge-small-en-v1.5"). Sparse and reranking work use SparseTextEmbedding with prithivida/Splade_PP_en_v1 and TextCrossEncoder with Xenova/ms-marco-MiniLM-L-6-v2 respectively, both shown in the README.

## Where FastEmbed is the wrong tool

FastEmbed does not train or fine-tune. There is no mention of a training loop, a loss function or a dataset loader in the README or in the repository layout, which lists fastembed/, tests/, docs/ and experiments/ but no training entry point. If your retrieval quality depends on adapting an encoder to your domain, FastEmbed is downstream of that work, not a substitute for it.

The second limit is the model catalogue. The README points to a supported models page and asks users to open a GitHub issue to request additions, which tells you the list is curated rather than open. A model absent from that list is only usable if you can produce an ONNX file and register it through add_custom_model with the right pooling, normalisation and dimension. If you cannot, the library cannot help you.

The third limit is the dependency pinning. The onnxruntime constraint in pyproject.toml excludes specific versions (1.20.0, and for some Python versions 1.24.0 and 1.24.1) and varies by interpreter version, and numpy is pinned per Python version as well, with Python 3.10 capped below 2.3.0. On a new Python release you may be waiting for an onnxruntime wheel that satisfies those ranges before FastEmbed installs at all. The README does not document rollback behaviour for downloaded models, and there is no mention of an offline or air-gapped install path beyond pointing ModelSource at a URL.

## FastEmbed compared with Sentence-Transformers and the OpenAI embeddings API

Sentence-Transformers is the closest comparison and the difference is runtime, not API shape. Both expose a model object and an encode-style call. Sentence-Transformers runs on PyTorch and carries the training lineage, so the same package can fine-tune a model on your data and then serve it. FastEmbed runs ONNX Runtime, which the README describes as faster than PyTorch, and it omits the training path entirely. Choose Sentence-Transformers when you need to adapt a model; choose FastEmbed when the model is fixed and the install size is the constraint.

The OpenAI embeddings API removes the local runtime altogether. You send text over HTTP and get vectors back, so there is no model download and no onnxruntime pin to satisfy, but every embedding is a network call with a per-token cost and your text leaves your infrastructure. FastEmbed runs locally and the README states it does not require a GPU. The README also claims FastEmbed is better than OpenAI Ada-002, which is a claim from the project and is not something this article can verify; treat it as a starting point for your own evaluation on your own data rather than a settled result.

Ollama occupies a third position: a local model server rather than a library. FastEmbed is imported into your process and returns generators; Ollama is a separate service you call over a socket. If you already run Ollama for generation, adding FastEmbed for retrieval means one more runtime in the stack, but it also means embeddings do not depend on the server being up.

## Maintenance, release cadence and the Apache-2.0 licence

The repository is not archived, and the last push was on 2026-09-09, which is recent relative to the release history. Releases are not frequent: v0.7.2 in August 2025, v0.7.4 in December 2025, and v0.8.0 in March 2026. That cadence suggests a project that ships when the model catalogue or the runtime pins need to move rather than on a fixed schedule.

Upgrade cost is dominated by two things. First, the onnxruntime and numpy pins in pyproject.toml are per Python version, so a Python upgrade can force a FastEmbed upgrade or block one. Second, the model catalogue changes between releases; a model you depend on is fetched by name at runtime, so pinning the FastEmbed version in your own lockfile is the practical way to keep a working combination. The README does not describe a migration guide between versions, and the repository has a RELEASE.md at the top level, so release process documentation exists there rather than in the README.

The library is Apache-2.0, and the repository carries a NOTICE file alongside the LICENSE. That covers FastEmbed's own code. It does not cover the model weights, which come from Hugging Face and carry their own licences; the README does not state a licence per model, so check each model you ship. This is a description of the project's licensing, not legal advice.

## Conclusion

Adopt FastEmbed if you need embeddings inside a Python service or a serverless runtime where a PyTorch install is the blocker, and if your model is on the supported list. Do not adopt it if you need to fine-tune the encoder, or if your model is not in the catalogue and you cannot supply an ONNX file. Before committing, check three things: that your target model appears on the supported models page, that onnxruntime has a wheel for your Python version under the pins in pyproject.toml, and that Apache-2.0 plus the individual model licences fit how you ship.

## FAQ

### How do I install FastEmbed?

Install it with pip install fastembed, or pip install fastembed-gpu for the GPU build. The README gives both commands and notes that the default install uses ONNX Runtime rather than PyTorch.

### Which models are supported by FastEmbed?

The README links to a supported models page and says the set covers dense text, sparse text such as SPLADE++, late interaction models like colbert-ir/colbertv2.0, image models, multimodal ColPali models and cross-encoder rerankers. Models outside that list can be registered with add_custom_model if you supply an ONNX file.

### Is FastEmbed free?

The FastEmbed library itself is Apache-2.0, and the repository includes a NOTICE file next to the LICENSE. Model weights are downloaded from Hugging Face and carry their own licences, which the README does not enumerate per model.

### What is FastEmbed?

FastEmbed is a Python library from Qdrant for embedding generation. It runs models through ONNX Runtime instead of PyTorch, does not require a GPU, and exposes separate classes for dense text, sparse text, late interaction, image and multimodal embeddings plus cross-encoder reranking.

### How does FastEmbed compare with Sentence-Transformers?

Sentence-Transformers runs on PyTorch and can fine-tune a model on your data, while FastEmbed runs ONNX Runtime and offers no training path. The README frames FastEmbed as lighter and faster, without the PyTorch dependency download.

## Sources

- [License: Apache-2.0](https://github.com/qdrant/fastembed/blob/main/LICENSE)
- [Project website](https://qdrant.github.io/fastembed/)
- [qdrant/fastembed on GitHub](https://github.com/qdrant/fastembed)
- [README](https://github.com/qdrant/fastembed/blob/main/README.md)
- [Releases](https://github.com/qdrant/fastembed/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/qdrant-fastembed
