fastembed-rs: Local ONNX Embeddings and Reranking in Rust
Rust library for generating vector embeddings and reranking locally!
At a glance
- What is it?
- fastembed-rs wraps ONNX Runtime and Hugging Face tokenizers into a synchronous Rust API for text, sparse, image and reranker models. It fits Rust services that want local inference without a Python sidecar; it fits poorly when you need GPU batching or a model outside its fixed enum.
- Who is it for?
- Adopt fastembed-rs when your application is already Rust, you want embeddings and reranking in-process, and you are happy to pick from the model enum rather than bring arbitrary checkpoints. Do not adopt it if you need GPU throughput, dynamic model loading, or a model that is not on the supported list.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 8 days ago.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Who fastembed-rs is for
The library exists to remove a specific piece of infrastructure: a Python sidecar whose only job is to turn text into vectors. The README describes it as a Rust library for generating vector embeddings and reranking locally, and the Cargo.toml description says the same thing in one line. The audience is a Rust service that already has a retrieval or search path and does not want a second process, a second runtime, and a second deployment artifact just to call a transformer.
That framing matters because the alternatives in this space are mostly Python-first. The README explicitly points readers elsewhere if they are not writing Rust, listing the Python fastembed, fastembed-go, and fastembed-js projects. So this crate is not trying to be the universal client. It is the Rust member of a family, and it inherits the family's model list rather than accepting arbitrary checkpoints.
The second design decision is sync. The README lists "Supports synchronous usage. No dependency on Tokio." as a feature. If your service is async, you still get the library, but you decide where the blocking call sits, typically behind a worker thread or a bounded pool. That is a deliberate trade: the crate does not force a runtime on you, and in exchange it does not hide the fact that inference is blocking work.
What sits underneath: ort, tokenizers and a fixed model enum
Inference runs through ort, the Rust wrapper for the ONNX runtime, and tokenization through the Hugging Face tokenizers crate. Both are named in the README's feature list and both appear in Cargo.toml, where ort is pinned to an exact release candidate and tokenizers is on the 0.23 line. Pinning ort exactly is a reasonable choice for a crate that exposes tensor shapes to users, but it also means ort upgrades arrive as fastembed-rs releases, not as a patch you apply yourself.
The model catalog is an enum, not a string. Callers name something like EmbeddingModel::AllMiniLML6V2, and the crate resolves that to a specific repository on Hugging Face plus the preprocessing that model expects. The README's list covers BGE small, base and large in English and Chinese, the MiniLM and mpnet sentence-transformers family, multilingual E5, mxbai-embed-large-v1, GTE, ModernBERT-embed-large, Jina v2 for code and English, Snowflake Arctic in five sizes, and EmbeddingGemma. Sparse embeddings have their own short list, image embeddings another, and reranking a fourth.
Quantized builds are exposed by suffix rather than by a separate API: append Q to the variant, so EmbeddingModel::BGESmallENV15Q. EmbeddingGemma also ships a 4-bit build as EmbeddingModel::EmbeddingGemma300MQ4. This is a clean convention, and it is also a constraint. If a model you want is not in the enum, there is no documented path in the README to register your own ONNX file and tokenizer pair.
Some models are gated behind Cargo features rather than being available by default. The README marks nomic-ai/nomic-embed-text-v2-moe as requiring the nomic-v2-moe feature and the Qwen3 embedding models, including Qwen/Qwen3-VL-Embedding-2B, as requiring the qwen3 feature. Both run on a candle backend, which is a separate inference path from ort. That split is worth noticing: the default build and the feature-gated builds do not share an execution engine.
Installing fastembed-rs and running a first embedding
The README gives two equivalent installation routes. Either add the crate with cargo, or write the dependency into the manifest yourself. The README shows the manifest line as version 5, while the repository's Cargo.toml is at 6.1.0, so check crates.io for the current version rather than copying the snippet literally.
Run this in your project directory:
cargo add fastembedOr edit Cargo.toml directly:
[dependencies]
fastembed = "5"The README's text embedding example shows both the default constructor and a customized one. The customized path is where the useful knobs live: you choose the model variant, opt into download progress output, and set the intra-op thread count.
use fastembed::{TextEmbedding, TextInitOptions, EmbeddingModel};
let mut model = TextEmbedding::try_new(Default::default())?;
let mut model = TextEmbedding::try_new(
TextInitOptions::new(EmbeddingModel::AllMiniLML6V2)
.with_show_download_progress(true)
.with_intra_threads(4),
)?;The first call to try_new is where the model files are fetched, which is why the progress flag exists. Expect a pause and network access on a cold cache. The README's document example also shows the prefix convention in the input strings themselves, with "passage: " and "query: " written into the text, and a comment noting that you can leave the prefix out but that it is recommended. The crate does not add those prefixes for you.
What you should see after a successful run is one vector per input string. The README's example passes four strings, three of them prefixed and one not, which is a small hint that mixed inputs are legal but not necessarily a good idea for retrieval quality.
Reranking and the other model families
Reranking is a separate entry point with its own default. The README lists BAAI/bge-reranker-base as the default, alongside bge-reranker-v2-m3, jinaai/jina-reranker-v1-turbo-en and jinaai/jina-reranker-v2-base-multiligual (spelled that way in the README). A reranker takes a query and a candidate passage together and produces a relevance score, which is a different shape of work from embedding: you cannot precompute it, and cost scales with the number of candidates you rescore. That makes the reranker a second-stage tool in a retrieval pipeline, not a replacement for the embedding step.
Sparse embeddings are the third family, defaulting to prithivida/Splade_PP_en_v1 with BAAI/bge-m3 also listed. Sparse vectors are useful when you want lexical matching behavior alongside dense similarity, and having both in one crate avoids a second dependency for hybrid search.
Image embeddings default to Qdrant/clip-ViT-B-32-vision, with resnet50-onnx, two Unicom-ViT variants and nomic-embed-vision-v1.5 also listed. The README notes two cross-modal pairings: nomic-embed-text-v1.5 pairs with nomic-embed-vision-v1.5, and Qdrant/clip-ViT-B-32-text pairs with clip-ViT-B-32-vision. Those pairings are the part people get wrong. Choosing a vision model without its matching text model breaks image-to-text search, and the README states the pairing rather than enforcing it in the type system.
Where fastembed-rs is the wrong choice
The model enum is the sharpest limitation. Everything the crate can run is on a list in the README. There is no documented mechanism for pointing it at an arbitrary ONNX file with a custom tokenizer, which means a model released after the crate's last update is unavailable until someone adds a variant. For teams that treat model choice as an ongoing experiment, that is a real cost.
The feature-gated models split the dependency tree. nomic-embed-text-v2-moe and the Qwen3 family need the nomic-v2-moe and qwen3 features respectively, and both use a candle backend rather than ort. A build that enables them pulls candle-core and candle-nn in addition to ort, which is a meaningful increase in compile time and binary size for a project that may only use one of the two paths.
The README does not document rollback, cache eviction, or offline model provisioning. If your deployment environment has no outbound network access, the first run cannot fetch weights, and the README does not describe a supported way to pre-seed the cache or point the loader at a local directory. That is the gap to check before designing around this crate. The README also does not state throughput or latency figures for any model, so sizing decisions have to come from your own measurements.
Finally, the exact pin on ort means a bug or a platform gap in that release candidate is not something you can patch by bumping a version in your own manifest. You wait for the crate.
How it compares with the Python fastembed
The README's own answer to "not looking for Rust?" is the Python fastembed from Qdrant, and the two share a lineage and a model list rather than being independent reimplementations. The practical difference is the runtime boundary. Python fastembed runs in the Python process, which means it composes with the rest of that ecosystem and is the natural choice if your retrieval pipeline already lives there. fastembed-rs runs in your Rust process, with no interpreter, no GIL, and no separate service to deploy or monitor.
That difference cuts both ways. In Python you can reach for a different backend or a custom model through that ecosystem's conventions. In Rust you get the enum. If your workload is a Rust API server that needs to embed incoming text at request time, the in-process option removes a network hop and a failure mode. If your workload is a batch job where you want to swap models weekly, the Python side is more forgiving. The Go and JavaScript ports listed in the README follow the same pattern for their respective runtimes, so the choice is really about which language your service is already written in.
Editorial conclusion
Adopt fastembed-rs when your application is already Rust, you want embeddings and reranking in-process, and you are happy to pick from the model enum rather than bring arbitrary checkpoints. Do not adopt it if you need GPU throughput, dynamic model loading, or a model that is not on the supported list. Before committing, verify three things: that your chosen variant exists (including any Q suffix for quantized builds), that the feature flags you need are enabled in Cargo.toml, and that the download path and cache location work in your deployment environment, because the first run fetches model files from Hugging Face.
Frequently asked questions
What does fastembed-rs do?
It is a Rust library that generates vector embeddings and reranks locally, using ort for ONNX inference and the Hugging Face tokenizers crate for encoding. It supports synchronous usage and does not depend on Tokio.
Which models does fastembed-rs support?
The README lists four families: text embedding models such as BAAI/bge-small-en-v1.5 (the default), sparse text embedding models such as prithivida/Splade_PP_en_v1, image embedding models such as Qdrant/clip-ViT-B-32-vision, and rerankers such as BAAI/bge-reranker-base. Some models, including nomic-embed-text-v2-moe and the Qwen3 embedding models, require enabling a Cargo feature.
What exactly are embeddings?
The README describes fastembed-rs as generating vector embeddings, and its example turns a list of strings into embeddings through TextEmbedding. It does not define the concept further, so the crate's documentation is not the place to learn the theory.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/anush008-fastembed-rs)