What is Embeddings?
Embeddings are dense numeric vectors that represent text, images or other content so that similar meanings land close together in vector space. A text embedding model turns a piece of content into such a vector, which is why search, clustering and recommendation systems depend on them.
How embeddings work
An embedding model is a neural network trained so that its output vector for a piece of content reflects meaning rather than surface form. Text is tokenised, passed through an encoder, and pooled into a single fixed-length vector, often a few hundred to a few thousand dimensions. Training objectives vary. Contrastive learning pulls paraphrases together and pushes unrelated pairs apart. Masked language modelling predicts hidden tokens. Static embedding methods such as those described by MinishLab/model2vec take an existing sentence transformer and distil it into a lookup table, so each token maps to a fixed vector and the cost of encoding a sentence collapses to averaging. The trade-off, as the model2vec documentation states, is a small drop in retrieval quality.
Once vectors exist, similarity is usually measured with cosine similarity or dot product. A search system embeds the query with the same model, then compares that vector against an index of document vectors. Approximate nearest neighbour indexes make this practical at scale, but they trade exactness for speed. The important constraint is that query and document must be embedded by the same model with the same pooling and normalisation. Mixing models produces vectors that are not comparable, and the failure is silent: results look plausible but are wrong.
Some models produce multiple vectors per input. huggingface/sentence-transformers wraps embedding, cross-encoder, sparse and multi-vector models behind four classes, according to its README. That matters because a cross-encoder scores a query and document pair jointly and tends to rank better than a bi-encoder, but it cannot be precomputed into an index. The common pattern is bi-encoder retrieval followed by cross-encoder reranking. Multimodal models extend the same idea to images and video. Tencent/WeMM-Embedding is a family of universal multimodal embedding models with 2B, 4B and 9B checkpoints and Matryoshka dimensions, which allow a single vector to be truncated to a shorter length with some loss of detail.
When you need embeddings and when you do not
You need embeddings when the matching criterion is meaning rather than exact tokens. Semantic search over documentation, deduplication of near-identical records, clustering of support tickets, retrieval-augmented generation and recommendation all rest on vector similarity. Keyword search still wins when the query is a known identifier, a code symbol or an exact phrase, and it is cheaper and easier to reason about. A hybrid of lexical and vector retrieval is common, but it adds a fusion step and a second set of parameters to tune.
You probably do not need embeddings for small, fixed vocabularies. If a category field has twenty values, a lookup table beats a vector index. You also do not need them if your corpus is tiny and changes constantly, because the overhead of maintaining an index and re-embedding updates can exceed the benefit. Embedding is not free: every document must be encoded once, and every model change forces a full re-embed. For a large corpus that is a real cost in GPU time and storage.
The choice of where inference runs is a separate decision from whether to use embeddings. huggingface/text-embeddings-inference serves embedding, re-ranking and sequence classification models behind an HTTP or gRPC API, and its README describes fast boot and clean containerisation. Its supported model list is narrower than the name suggests, so checking compatibility comes before planning a deployment. michaelfeil/infinity is an MIT-licensed Python serving engine that exposes text embeddings, rerankers, CLIP, CLAP and ColPALI models behind a FastAPI REST API, aimed at teams that want OpenAI-shaped endpoints without sending documents to a hosted provider. The trade-off is that you own the GPU, the model download and the capacity planning. If you would rather not run a server at all, brianpetro/obsidian-smart-connections indexes an Obsidian vault with a local embedding model and surfaces related notes in list and graph views, with zero setup and no API key by default.
Common pitfalls and limits
The first pitfall is treating similarity as truth. Cosine similarity between two vectors is a learned proxy, not a measure of correctness. Models trained on one domain often transfer poorly to another, and a model that ranks well on general benchmarks can fail on legal, medical or code text. Evaluation is therefore necessary. embeddings-benchmark/mteb evaluates text and multimodal embedding models against named tasks and benchmarks, and ships both a Python API and a CLI. Running a benchmark on your own data is more informative than reading a leaderboard, because leaderboard tasks may not resemble your queries.
The second pitfall is dimension and storage. A million documents at 768 dimensions in float32 occupy roughly three gigabytes before index overhead. Matryoshka models, including the checkpoints in Tencent/WeMM-Embedding, allow truncating vectors to shorter lengths, which reduces storage and speeds comparison at some cost in quality. Quantisation and product quantisation are alternatives, but each introduces approximation that must be measured.
The third pitfall is operational. An embedding index is derived data, and it drifts from the source unless re-embedding is scheduled. Changing the model invalidates every vector. The brianpetro/obsidian-smart-connections README does not document rollback or index recovery, which is a reminder to ask what happens when the index is corrupted or a model update changes results. The core of that plugin is source-available rather than conventionally licensed, so teams with strict licensing requirements should check before adopting it.
The fourth pitfall is chunking. Long documents must be split before embedding, and the split determines what can be retrieved. A chunk that cuts a sentence in half produces a vector that represents neither half well. Overlapping chunks mitigate this at the cost of index size and duplicate results. The right chunk size depends on the model's context window and on the shape of the questions being asked, and no default is universally correct.
How embeddings show up in open-source projects
huggingface/sentence-transformers is the default entry point for Python teams that need semantic search without training a model first, because it wraps several model families behind four classes. It is a library, not a service, so deployment is the user's problem.
huggingface/text-embeddings-inference and michaelfeil/infinity both turn models into HTTP services, but they differ in scope. TEI is a Rust toolkit for embedding, re-ranking and sequence classification, with a narrower model list. Infinity is Python, MIT-licensed, and covers text embeddings, rerankers, CLIP, CLAP and ColPALI behind a FastAPI REST API. The first is leaner; the second covers more modalities and asks you to operate a Python service.
rom1504/clip-retrieval computes CLIP embeddings and builds a searchable index with a Flask backend and a browser front end. It is a pipeline of separate commands rather than a hosted service, so it suits people who want to run each stage and inspect the artifacts. That shape is a constraint: there is no single process to start.
MinishLab/model2vec addresses the cost of inference rather than quality. It converts a sentence transformer into static embeddings, cutting size by up to 50x and CPU inference cost by up to 500x according to its documentation, with a small retrieval quality drop that the documentation acknowledges. It is the weaker option when maximum ranking accuracy is the priority.
apple/embedding-atlas renders millions of points in the browser with WebGPU, clusters and labels them automatically, and links charts to an embedding view. It is for inspection, not retrieval. embeddings-benchmark/mteb is for choosing a model, not serving one. pykeen/pykeen applies the same vector idea to knowledge graphs, wrapping 40 knowledge graph embedding models and 37 datasets behind a single pipeline call, which is a different problem from text retrieval even though the vocabulary overlaps. Tencent/WeMM-Embedding targets one vector space for images, video and visual documents, with checkpoints at 2B, 4B and 9B parameters and Matryoshka dimensions. brianpetro/obsidian-smart-connections shows the smallest useful deployment: a local model inside a note-taking plugin, no API key, with the licensing and recovery caveats noted above.
Choosing an approach without overbuilding
Start from the retrieval task, not the model. Write down what a correct result looks like, collect a few dozen queries with known answers, and measure. If lexical search already finds them, embeddings add cost without benefit. If it does not, an off-the-shelf sentence transformer is usually enough for a first pass, and huggingface/sentence-transformers is the shortest path to that. Only move to a served model when latency, throughput or data residency requires it, and then choose between TEI and Infinity based on model coverage and operational preference rather than raw claims.
Re-embedding is the cost that surprises teams. Model upgrades, chunking changes and schema changes all invalidate the index. Budget for a rebuild path before the first deployment, and keep the source text so the index can be regenerated. If a project does not document recovery, as is the case for brianpetro/obsidian-smart-connections, treat that as an open risk rather than an oversight to ignore.
In practice
Embeddings turn content into vectors so that similarity can be computed numerically, and the hard parts are not the vectors themselves but keeping query and document models aligned, re-embedding when anything changes, and measuring quality on your own queries. For a first pass, read the huggingface/sentence-transformers documentation, run a small evaluation with embeddings-benchmark/mteb, and only then decide whether a serving layer such as huggingface/text-embeddings-inference or michaelfeil/infinity is worth operating.