# EmbedAnything: Rust Ingestion, Inference and Indexing for RAG Pipelines

> EmbedAnything is a Rust embedding pipeline with Python bindings that turns files, images and audio into vectors and streams them to a vector database. It is a good fit when you want local inference without PyTorch, and a poor fit when you need a managed API.

**StarlightSearch/EmbedAnything** — Highly Performant, Modular, Memory Safe and Production-ready Inference, Ingestion and Indexing built in Rust 🦀

- Repository: https://github.com/StarlightSearch/EmbedAnything
- Website: https://embed-anything.com/
- Stars: 1,309 · Forks: 142
- Language: Rust
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/starlightsearch-embedanything

## The problem EmbedAnything solves, and for whom

Most retrieval-augmented generation stacks start with the same chore: walk a directory of PDFs, markdown, images and audio, split each file into chunks, run every chunk through an embedding model, and push the resulting vectors into a database. In Python that chore usually drags in PyTorch, which means a large container image, a CUDA dependency you have to match against your driver, and a memory footprint that grows with batch size. The README states the project's position plainly: "No Dependency on Pytorch: Easy to deploy on cloud, comes with low memory footprint."

EmbedAnything targets engineers building the ingestion half of a RAG system who want that half to run locally and cheaply. The README lists text, images, audio, PDFs and websites as inputs, and dense, sparse, ONNX, model2vec and late-interaction embeddings as outputs. The primary language is Rust, the Python package is built with maturin, and the workspace splits into processors, rust, python and server crates. If your team writes Python but is tired of shipping torch into a container, this is the audience the project is written for.

## Vector streaming: how ingestion, inference and indexing are separated

The mechanism the README spends the most words on is vector streaming. Rather than reading a file, embedding it, and writing the vector in one sequential pass, the pipeline separates document preprocessing from model inference, and the README claims this "transforms a sequential bottleneck into an efficient, concurrent workflow." The stated implementation detail is Rust MPSC channels: file processing, indexing and inferencing run on different threads, and embeddings are written straight to the vector database rather than accumulated in the process. The README frames the memory argument as "no memory leak as embeddings are directly saved to vector database."

That design has a visible consequence. Because embeddings leave the process as they are produced, the peak resident set is closer to the size of the in-flight batch than to the size of the corpus. It also means the vector database is on the critical path of ingestion: if the database is slow or unavailable, the pipeline backpressures rather than buffering everything in RAM. The repository layout reflects the split, with a processors/ crate for file handling, a rust/ crate for inference, and a server/ crate for the HTTP-facing deployment.

## Installing embed_anything and running a first embedding job

The Python package is named embed_anything and requires Python 3.10 or newer; pyproject.toml declares requires-python = ">=3.10" and classifiers up to 3.14. The build backend is maturin with the extension-module feature, so wheels are compiled Rust rather than pure Python. The single declared runtime dependency is onnxruntime==1.22.0.

```bash
pip install embed-anything
```

After installation the import name is embed_anything. The README points to a Colab notebook and to the examples/ directory for runnable scripts; examples/text.py is the smallest entry point for a text corpus.

The repository also ships a prebuilt server image, which the README describes with "Just pull it: starlightsearch/embedanything-server". There is a CUDA variant built from server-cuda.Dockerfile and a plain one from server.Dockerfile, so GPU and CPU deployments are separate images rather than one image with a runtime flag.

```bash
docker pull starlightsearch/embedanything-server
```

For the server build the Dockerfile sets ORT_DYLIB_PATH to the onnxruntime shared library inside site-packages, which is worth copying if you build your own image and hit a missing-library error at import time.

## Where the Python API is thinner than the Rust core

The README carries a warning that is easy to skim past: "WhichModel has been deprecated in pretrained_hf". A deprecated enum in the model-selection path is a real migration cost, because model selection is the first thing any new user writes. Anyone following an older blog post or notebook that passes WhichModel will need to check the current Python docs at embed-anything.com/references before assuming the snippet still runs.

The second limitation is structural. The core is Rust, exposed through a maturin-built extension, and the workspace version is 0.7.2 while the most recent release listed is v0.7.1 from 2026-07-10. Pre-1.0 versioning plus a compiled extension means that upgrading can require rebuilding against a matching onnxruntime, and a pinned onnxruntime==1.22.0 in pyproject.toml means you cannot freely float that dependency without testing the compiled module against a different runtime. If your environment already pins a different onnxruntime version for another service, that conflict is yours to resolve.

Finally, this is the wrong tool if you want a hosted embedding endpoint with an SLA. Everything here runs locally or in your own container, which is the point, but it also means you own the GPU, the model files and the failure modes.

## How EmbedAnything differs from sentence-transformers

The closest comparison is sentence-transformers, the Python library most teams reach for first. Both produce dense vectors from text, and both can run models locally. The difference is the dependency graph. sentence-transformers is built on PyTorch and Transformers, so installing it pulls the torch stack into your image; EmbedAnything is built on Candle and ONNX Runtime, and the README lists Candle as one of its backends alongside ONNX and cloud models. The README's own framing of the choice is "Why Candle?" followed by the no-PyTorch argument.

The second difference is scope. sentence-transformers embeds text and, with extra work, images. EmbedAnything treats multimodal ingestion as a first-class concern: the README lists PDFs, txt, md, JPG images and .WAV audio, and the examples directory contains audio.py, clip.py, colpali.py and text_ocr.py alongside plain text.py. If you only ever embed strings and you are already happy with torch in your image, sentence-transformers is the lower-friction choice. If your corpus is a folder of mixed file types and container size matters, the trade flips.

## Maintenance, licensing and what an upgrade actually costs

The repository is not archived and the last push was on 2026-08-12, roughly a month before this writing, so the project is being worked on. The release cadence visible in the listed releases is uneven: 0.6.6 in 2025-10-28, 0.7.0 in 2025-12-27, and v0.7.1 in 2026-07-10, with the workspace Cargo.toml already at 0.7.2. That gap between the last tagged release and the workspace version suggests unreleased changes on main, which is normal for pre-1.0 but means pinning to a tag is safer than tracking the branch.

On licensing, the workspace Cargo.toml declares license = "Apache-2.0" and the repository contains a LICENSE file, while pyproject.toml lists "License :: OSI Approved :: MIT License" in its classifiers. Those two statements do not agree, and the classifier is metadata rather than the licence text. Before redistributing a modified build, read the LICENSE file itself and, if the discrepancy matters to your legal review, raise it upstream rather than assuming either value. This is not legal advice.

The upgrade cost is dominated by the compiled extension. A version bump can change the Rust ABI of the extension module, so a wheel built for one release will not necessarily load against another. The Dockerfile shows the intended build path: cargo-chef caches the dependency layer, maturin builds the wheel with the mkl and extension-module features, and the runtime stage installs the resulting .whl. Reproducing that in CI is the cheapest way to keep upgrades routine.

## Conclusion

Adopt EmbedAnything when you want local, PyTorch-free embedding of mixed file types and you are willing to pin onnxruntime==1.22.0 and keep the Rust toolchain available for builds. Do not adopt it if you need a hosted embedding API, a stable long-term Python API, or a supported path for the deprecated WhichModel enum in pretrained_hf. Verify first that your target vector database has an adapter under examples/adapters, that your CUDA or MKL environment matches the Dockerfile's expectations, and that the model you intend to use is not one of the ones the README marks as deprecated.

## FAQ

### What does EmbedAnything do with my documents?

It reads text, PDF, markdown, image and audio sources, chunks them with methods the README describes as semantic and late-chunking, runs them through an embedding model, and streams the resulting vectors to a vector database. The README describes the streaming design as separating file processing, indexing and inferencing onto different threads.

### Can I use an LLM as the embedding model in EmbedAnything?

The README lists dense, sparse, ONNX, model2vec and late-interaction embeddings as supported, along with Candle, ONNX and cloud model backends, but it does not describe using a generative LLM as the embedding model. The examples directory includes reranker and ModernBert entries rather than a generative-LLM embedding path.

### What kinds of embeddings does EmbedAnything support?

According to the README, the project supports dense, sparse, ONNX, model2vec and late-interaction embeddings. The examples directory backs this up with files such as splade.py, colbert.py, colbert_onnx.py, colpali.py and model2vec.py.

## Sources

- [License: Apache-2.0](https://github.com/StarlightSearch/EmbedAnything/blob/main/LICENSE)
- [Project website](https://embed-anything.com/)
- [README](https://github.com/StarlightSearch/EmbedAnything/blob/main/README.md)
- [Releases](https://github.com/StarlightSearch/EmbedAnything/releases)
- [StarlightSearch/EmbedAnything on GitHub](https://github.com/StarlightSearch/EmbedAnything)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/starlightsearch-embedanything
