Infinity: a self-hosted embedding and reranking server for HuggingFace models
Infinity is a high-throughput, low-latency serving engine for text-embeddings, reranking models, clip, clap and colpali
At a glance
- What is it?
- Infinity is an MIT-licensed Python serving engine that exposes text embeddings, rerankers, CLIP, CLAP and ColPALI models behind a FastAPI REST API. It is aimed at teams that want OpenAI-shaped endpoints without sending documents to a hosted provider, and the trade-off is that you own the GPU, the model download and the capacity planning.
- Who is it for?
- Adopt Infinity if you already run HuggingFace embedding or reranking models in Python and want one process to serve several of them behind an OpenAI-compatible API on your own CUDA, ROCm, CPU, AWS INF2 or Apple MPS hardware. Do not adopt it if you need a stable, versioned API surface: the project is still on 0.0.x releases, and the README does not document rollback or a compatibility policy between versions.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Activity is slowing. The repository last received commits 6 months ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap Infinity fills between HuggingFace weights and an HTTP endpoint
Most teams that pick an embedding model start with sentence-transformers in a notebook. The model loads, the vectors look right, and then the notebook has to become a service. Someone writes a FastAPI wrapper, then discovers that batching, tokenization and GPU memory are all separate problems, and that a second model (a reranker, say) means a second process and a second port. Infinity is the answer to that second stage. It is a serving engine, not a library: you point it at a HuggingFace model id and it exposes an HTTP API, with dynamic batching and tokenization handled in dedicated worker threads rather than in the request handler. The scope is wider than text. The README lists text embeddings, reranking models, clip, clap and colpali, so image and audio encoders sit behind the same server. The audience is engineers running retrieval or RAG pipelines who want the model weights on their own hardware and an API their application code can call without a Python import.
How the server is put together: FastAPI, worker threads and selectable backends
The README describes the inference server as built on PyTorch, optimum (ONNX and TensorRT) and CTranslate2, with FlashAttention used to target NVIDIA CUDA, AMD ROCm, CPU, AWS INF2 or Apple MPS. That backend list matters more than it first appears. ONNX and TensorRT paths are not the same artifact as the PyTorch path, and the Docker images are published per accelerator family, so the choice of backend is partly a choice of container. The HTTP layer is FastAPI, and the README states the OpenAPI schema is aligned to OpenAI's embeddings specification, which is the reason a client written against a hosted embeddings endpoint can be repointed at a local server with a base URL change. Multi-model support is where the design gets interesting: the v2 CLI can launch several models in one process, including a reranker alongside an embedder, and the README says Infinity orchestrates them. That is a real operational simplification, because a reranking stage that would otherwise need its own deployment shares a tokenizer pool and a process. The cost is that a single process now holds every loaded model in accelerator memory, and the README does not document a per-model memory budget.
Installing Infinity and serving your first embedding model
There are two supported launch paths. The pip path installs the CLI with all optional dependencies:
pip install infinity-emb[all]After that, with the virtual environment active, the CLI is on your PATH. The README's first example launches a single small embedding model:
infinity_emb v2 --model-id BAAI/bge-small-en-v1.5You should see the server start and the model load, after which the FastAPI endpoints are reachable on the local port. To see every parameter the v2 CLI accepts, including the flags for multiple models and for an API key, the README points at the built-in help:
infinity_emb v2 --helpThe README calls the Docker route recommended. Instead of installing the CLI, you run the published image `michaelf34/infinity`, and you must mount your accelerator first, for example by installing `nvidia-docker`. The README does not spell out the full `docker run` invocation in the section available here, so check the docs site at https://michaelfeil.github.io/infinity for the exact flags before you write a compose file. One detail worth noting for anyone scripting a deployment: the README states that the v2 CLI lets you supply all arguments through environment variables as well as command-line arguments, which is what makes the container path practical.
Where Infinity is the wrong tool
Infinity serves models; it does not train, fine-tune or evaluate them, and it does not store vectors. If your problem is choosing between two embedding models on your own corpus, this server gives you no help with that, and a benchmark script using sentence-transformers directly will be faster to iterate on. The version numbering is the second constraint. The releases listed are 0.0.77, 0.0.76 and 0.0.75, published in August 2025, March 2025 and January 2025 respectively, and the last push to the repository was on 2026-03-24. A 0.0.x series carries no compatibility promise, and the README does not document a rollback procedure or a deprecation policy, so pinning a version and reading the release notes before upgrading is the only safe path. Third, the accelerator matrix is a hard boundary rather than a soft one. The README names CUDA, ROCm, CPU, AWS INF2 and Apple MPS, and mentions experimental int8 support for CPU and CUDA and fp8 for H100 and MI300, plus Blackwell support added in July 2025. If you are on hardware outside that list, the published images will not help you. Finally, if your workload is a handful of documents per minute, a batch job that imports sentence-transformers and writes vectors to a database has fewer moving parts than any server, and Infinity buys you nothing there.
Infinity versus a single-model server such as HuggingFace Text Embeddings Inference
The closest comparison is HuggingFace's own Text Embeddings Inference, which occupies the same niche: a compiled Rust server that exposes embedding and reranking models over HTTP. The difference in approach is the model surface and the runtime. TEI is a narrower tool built around a specific serving path, and the README's own model link points at HuggingFace models tagged `text-embeddings-inference`, which tells you the two projects overlap on the common case. Infinity's distinguishing claim is breadth: it mixes embedding, reranking, clip, clap and colpali models in one process, runs them on PyTorch, ONNX, TensorRT or CTranslate2 depending on your hardware, and keeps the Python stack in the loop. That Python dependency is also the cost. A Rust server has a smaller dependency surface and a simpler container; a Python server built on FastAPI, PyTorch and optimum gives you more backends and more model types, at the price of a heavier image and a larger set of things that can break during an upgrade. If you only ever serve one text embedding model on CUDA and never need reranking or multimodal encoders, the narrower tool is the lower-risk choice.
Licence, upgrades and the cost of staying current
Infinity is developed under the MIT License, and the LICENSE file sits at the top level of the repository. That is permissive: you can run it commercially, modify it and ship it inside a product, provided you keep the copyright notice and the licence text. The licence covers the server, not the model weights you load through it, and those carry their own terms on HuggingFace, so check each model separately. This is a description of the licence text, not legal advice. On maintenance, the last push to the repository was on 2026-03-24, and the most recent release listed is 0.0.77 from 2025-08-22. The upgrade cost is concentrated in two places: the CLI surface, which the README says accepts arguments through both flags and environment variables, and the backend stack. A bump that moves the optimum or CTranslate2 version can change which ONNX or TensorRT artifacts load, and because the project publishes separate Docker images per accelerator family, an upgrade means re-validating the image you actually deploy rather than the one you tested on a laptop.
Editorial conclusion
Adopt Infinity if you already run HuggingFace embedding or reranking models in Python and want one process to serve several of them behind an OpenAI-compatible API on your own CUDA, ROCm, CPU, AWS INF2 or Apple MPS hardware. Do not adopt it if you need a stable, versioned API surface: the project is still on 0.0.x releases, and the README does not document rollback or a compatibility policy between versions. Before committing, verify that your accelerator is covered by the published Docker images, that the model you want is loadable through one of the PyTorch, ONNX, TensorRT or CTranslate2 backends, and that your client can speak the OpenAI embeddings schema rather than a vendor-specific one.
Frequently asked questions
How do you install the Infinity embedding server?
The README gives two paths. You can run `pip install infinity-emb[all]` and then launch the CLI with `infinity_emb v2 --model-id BAAI/bge-small-en-v1.5`, or you can use the published Docker image `michaelf34/infinity`, which the README calls the recommended option, after mounting your accelerator.
How do you use Infinity from Python?
The README does not show a Python client example in the section available here. It states that the server's OpenAPI schema is aligned to OpenAI's embeddings specification, so a Python client written against an OpenAI-compatible embeddings endpoint can be pointed at the local server. The README also notes a separate `pip install infinity_client` package added in October 2024.
What models can Infinity serve?
The README states that you can deploy any embedding, reranking, clip and sentence-transformer model from HuggingFace, and the project description adds clap and colpali. The v2 CLI can launch multiple models in one process and Infinity orchestrates them.
Which accelerators does Infinity support?
The README names NVIDIA CUDA, AMD ROCm, CPU, AWS INF2 and Apple MPS, with FlashAttention used to get the most out of them. It also mentions experimental int8 support for CPU and CUDA, fp8 for H100 and MI300, and Blackwell support added in July 2025.
What licence is Infinity released under?
Infinity is developed under the MIT License, and the LICENSE file is at the top level of the repository. That covers the server itself; the model weights you load carry their own terms on HuggingFace.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/michaelfeil-infinity)