Model or dataset
Blaizzy/mlx-embeddings avatar
Blaizzy/mlx-embeddings

MLX-Embeddings: Local Vision and Language Embeddings on Apple Silicon

MLX-Embeddings is the best package for running Vision and Language Embedding models locally on your Mac using MLX.

446 stars63 forksPythonNOASSERTION

At a glance

What is it?
A Python package that runs text and image embedding models on a Mac through MLX, with a model-agnostic loader and a multimodal process API. It is a good fit for offline retrieval work on Apple hardware and a poor fit for anyone without it.
Who is it for?
Adopt MLX-Embeddings if you are building retrieval or classification on an Apple Silicon machine and want embeddings computed on that machine rather than over an API. Do not adopt it if your serving target is Linux or CUDA, or if you need a stable public API: the package is at 0.1.0, the last push was on 2026-05-13, and the README does not document rollback.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 139 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem MLX-Embeddings solves, and who has it

Most embedding pipelines assume the model lives somewhere else. You send text to a hosted endpoint, get a vector back, and pay for the round trip in latency, money and the requirement that the text leave your machine. MLX-Embeddings moves that step onto the Mac itself. The README describes it as a package for running Vision and Language Embedding models locally on your Mac using MLX, and the supported architecture list backs that up: XLM-RoBERTa, BERT, ModernBERT, Qwen3, Qwen3-VL, Llama Bidirectional, Llama Nemotron VL and an OpenAI Privacy Filter variant.

The audience is narrow and specific. You need an Apple Silicon Mac, Python 3.10 or newer according to pyproject.toml, and a reason to keep embeddings local. That last part matters for anyone handling documents that cannot be shipped to a third party, and for anyone building retrieval-augmented-generation prototypes where the embedding step is the only piece that would otherwise require a network call. It is not a general-purpose embedding server. There is no HTTP surface described in the README, no batching daemon, no GPU cluster story. It is a library you import.

How the loader and the model.process API fit together

The package exposes two loading paths. The older one, used in the single-item examples, comes from mlx_embeddings.utils and returns a model plus a tokenizer. You encode text yourself, run the model, and read either the CLS token from last_hidden_state or the ready-made text_embeds field. The README notes that text_embeds uses mean pooling for BERT and XLM-RoBERTa, while ModernBERT takes its pooling strategy from the config file and defaults to mean. That distinction is easy to miss and will silently change your vectors if you swap architectures.

The newer path, shown for Qwen3-VL, returns a model and a processor and routes everything through model.process(...). Inputs are dictionaries: a text key, an image key, or both. The processor handles the modality-specific preprocessing, and the model returns a single embedding tensor. For reranking, the input shape changes: one instruction, one query, and a documents list, and the output is a score per document rather than a vector per item. The README's reranking example prints a shape of (3,) for three documents, which confirms the scorer returns one number per candidate rather than an embedding matrix.

What is not documented is how the processor decides image resolution, how instruction strings are templated, or whether process() batches internally. The README shows the call and the expected output shape and stops there.

Installing MLX-Embeddings and running a first embedding

Installation is a single pip command, as the README states. The dependency list in requirements.txt pins mlx>=0.31.1, mlx-vlm>=0.4.0, transformers[sentencepiece]>=5.0.0 and huggingface-hub>=0.25.1, so the first install pulls a meaningful amount of code and the model weights are downloaded separately on first load.

bash
pip install mlx-embeddings

After that, the smallest useful program loads a small sentence-transformer model and produces two vectors you can compare. The README's single-item example uses mlx-community/all-MiniLM-L6-v2-4bit. The model name is passed straight to load(), which resolves it through the Hugging Face hub.

python
from mlx_embeddings.utils import load

model, tokenizer = load("mlx-community/all-MiniLM-L6-v2-4bit")

text = "I like reading"
input_ids = tokenizer.encode(text, return_tensors="mlx")
outputs = model(input_ids)
raw_embeds = outputs.last_hidden_state[:, 0, :]
text_embeds = outputs.text_embeds

You should get a tensor from text_embeds with one row for the sentence. Note that raw_embeds is the CLS token and text_embeds is the pooled, normalized version; the README is explicit that these are different things and that the pooling rule depends on the architecture.

For multimodal work the entry point changes. The Qwen3-VL example builds a list of dictionaries, mixing text-only, image-only and combined items, then calls process() once and computes a similarity matrix with a single matrix multiply.

python
import mlx.core as mx
from mlx_embeddings import load

model, processor = load("Qwen/Qwen3-VL-Embedding-2B")

inputs = [
    {"text": "A woman playing with her dog on a beach at sunset.",
     "instruction": "Retrieve images or text relevant to the user's query."},
    {"image": "https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen-VL/assets/demo.jpeg"},
]

embeddings = model.process(inputs, processor=processor)
similarity = embeddings @ embeddings.T
mx.eval(embeddings, similarity)

The README states the expected embedding shape for the four-item version of this example is (4, 2048), which is a useful sanity check: if your output has a different second dimension, the model you loaded is not the one you think it is.

Where MLX-Embeddings breaks down

The first constraint is hardware. MLX is Apple's array framework for Apple Silicon, and nothing in the README suggests a CPU or CUDA fallback. If your inference target is a Linux box with an NVIDIA card, this package is the wrong tool and no amount of configuration will fix it. The same applies to CI runners on x86.

The second is the surface area of the API. The single-item path gives you outputs.last_hidden_state, outputs.text_embeds, outputs.pooler_output and outputs.logits depending on the model, and the README demonstrates each against a different architecture. There is no documented guarantee that a given output field is populated for a given model. Loading a sequence-classification checkpoint and reading text_embeds, or the reverse, is the kind of mistake the documentation does not warn you about.

The third is maturity. The most recent release listed is v0.1.0 from 2026-03-24, and the last push to the repository was on 2026-05-13. That is a young package with a moving API. The README itself says the team is continuously working to expand architecture support, which is a reasonable thing to say and also a signal that the supported-model list is a snapshot rather than a contract. The README does not document rollback, deprecation policy, or what happens to your pipeline when a model architecture is dropped.

Finally, there is the licence. The README and pyproject.toml both state GNU General Public License v3, while the repository's own licence field is reported as NOASSERTION. That mismatch is worth resolving with your own reading of the LICENSE file before you depend on the package in a closed product.

How this differs from sentence-transformers

The obvious comparison is sentence-transformers, which also loads embedding models from the Hugging Face hub and also returns pooled vectors. The difference is the runtime underneath. sentence-transformers targets PyTorch and runs wherever PyTorch runs, including CUDA and CPU; MLX-Embeddings targets MLX and runs on Apple Silicon. If you already have a PyTorch pipeline that works, switching buys you Apple GPU acceleration and nothing else, and it costs you portability.

The second difference is the multimodal story. sentence-transformers is primarily a text-embedding library, and the image side requires separate tooling. MLX-Embeddings treats image and text inputs as first-class alternatives in the same dictionary format, and the Qwen3-VL and Llama Nemotron VL entries in the supported list are explicitly described as vision-language models. If your retrieval corpus mixes screenshots and prose, that single-input-list design is the actual reason to pick this package over the alternative, not the speed.

The third difference is reranking. The README shows a reranker model loaded through the same load() call and scored through the same process() method, returning one score per document. With sentence-transformers you would typically reach for a separate cross-encoder class. Here it is the same code path with a different input shape.

Classification and PII detection beyond plain embeddings

The package is not limited to producing vectors. The README demonstrates masked language modeling with ModernBERT, sequence classification with NousResearch/Minos-v1, and token classification with openai/privacy-filter. The privacy filter entry is the most specific: the README describes it as a bidirectional 1.5B-parameter, 50M-active sparse-MoE token classifier that tags personally identifiable information with BIOES spans over eight categories, listed as person, email, phone, URL, address, date, account number and secret.

The example loads the model, reads id2label from the config, tokenizes a sentence containing a name, an email and a phone number, and takes an argmax over the logits to get per-token labels. That is a different workload from embedding, and it is worth being clear about the trade-off: a 1.5B-parameter model with 50M active parameters is not free to run, and the README gives no latency or memory figures. If your goal is redacting a large document set, the cost per token is the number you need and the documentation does not provide it.

Sequence classification follows the same pattern. The Minos-v1 example reads id2label, runs the model, and prints a logit per label. Note that the printed values are logits, not probabilities, and the README does not apply a softmax. Reading them as confidence scores would be a mistake.

Maintenance, upgrade cost and the licence question

The release history is short and the gaps are uneven. v0.0.4 landed on 2025-09-08, v0.0.5 on 2025-10-29, and v0.1.0 on 2026-03-24. The jump from 0.0.5 to 0.1.0 spans roughly five months, and the last push to the default branch was on 2026-05-13. Nothing in the repository is archived, but a package at 0.1.0 with a multimodal API introduced in the same release is not a stable target. Pin the version in your requirements file rather than tracking main.

Upgrade cost concentrates in two places. Model weights are fetched by name from the Hugging Face hub, so a model rename or a quantization change on the hub side affects you without a package upgrade. The output contract is the other. Because different architectures populate different fields on the output object, a version bump that changes pooling defaults, or a config change in a model you already use, can shift your vectors without raising an error. If you store embeddings in a vector index, that silently invalidates the index.

On licensing: the README and pyproject.toml both state GNU General Public License v3, and the repository metadata reports NOASSERTION. GPLv3 is a copyleft licence, which has implications for how the package can be combined with proprietary code. That is a question for your own legal review, not something this article can settle, and the disagreement between the two stated sources is the thing to resolve first.

Editorial conclusion

Adopt MLX-Embeddings if you are building retrieval or classification on an Apple Silicon machine and want embeddings computed on that machine rather than over an API. Do not adopt it if your serving target is Linux or CUDA, or if you need a stable public API: the package is at 0.1.0, the last push was on 2026-05-13, and the README does not document rollback. Verify first that the architecture you need appears in the supported list, then check the pooling strategy in the model's config, since ModernBERT takes it from the config file while BERT and XLM-RoBERTa use mean pooling.

Frequently asked questions

What is MLX-Embeddings?

It is a Python package for running vision and language embedding models locally on a Mac using MLX. It supports single-item and batch processing and includes utilities for comparing text similarities.

How do I install MLX-Embeddings?

The README gives a single command, pip install mlx-embeddings. The package requires Python 3.10 or newer according to pyproject.toml, and model weights are downloaded on first load.

Which model architectures does MLX-Embeddings support?

The README lists XLM-RoBERTa, BERT, ModernBERT, Qwen3, Qwen3-VL, Llama Bidirectional, Llama Nemotron VL and an OpenAI Privacy Filter variant. The README states the list is being expanded, so treat it as a snapshot.

Can I use an LLM as an embedding model with MLX-Embeddings?

The supported list includes Llama Bidirectional, described as Llama-based bidirectional embedding models such as NVIDIA NV-Embed, and Qwen3's embedding model. These are embedding checkpoints rather than general chat models, so a standard causal LLM is not what the loader expects.

Does MLX-Embeddings run on Linux or Windows?

The README describes the package as running on your Mac using MLX, and no other platform is mentioned. MLX is Apple's array framework, so a Linux or Windows target is not covered by the documentation.

What licence does MLX-Embeddings use?

The README and pyproject.toml both state GNU General Public License v3, while the repository metadata reports NOASSERTION. Resolve that discrepancy against the LICENSE file before depending on it.

Official sources

  1. Blaizzy/mlx-embeddings on GitHub
  2. Issues
  3. README
  4. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/blaizzy-mlx-embeddings.svg)](https://hysenlabs.com/projects/blaizzy-mlx-embeddings)