Model or dataset
infiniflow/infinity avatar
infiniflow/infinity

Infinity (infiniflow/infinity): a C++ hybrid search database for LLM retrieval

The AI-native database built for LLM applications, providing incredibly fast hybrid search of dense vector, sparse vector, tensor (multi-vector), and full-text.

4,727 stars452 forksC++Apache-2.0

At a glance

What is it?
Infinity is an Apache-2.0 database that puts dense vectors, sparse vectors, tensors and full text behind one query engine. It ships a single binary and a Python client, and the trade-off is that you run a server process rather than importing a library.
Who is it for?
Adopt Infinity if your retrieval layer needs dense, sparse, tensor and full-text search in one query and you can run a server with a volume at /var/infinity. Do not adopt it if you want an in-process library with no daemon, or if you need a managed cloud endpoint, since the README describes Docker and binary deployment only.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 7 days ago.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The retrieval problem Infinity is aimed at

Most RAG stacks end up gluing two systems together: a vector index for embeddings and a separate search engine for keyword matching. That split forces you to merge two ranked lists in application code, and it doubles the operational surface. Infinity's premise is that dense embedding, sparse embedding, tensor (multi-vector) and full-text search belong in one engine, so a single query can combine them with a reranker. The README lists RRF, weighted sum and ColBERT among the supported rerankers, and the example directory contains a ColBERT_reranker_example folder, so reranking is a first-class step rather than something you bolt on.

The audience is narrower than the tagline suggests. This is for teams building retrieval for LLM applications who are willing to operate a database process. If you are embedding a search index into a desktop app or a short-lived script, the client-server split is friction you do not need.

How the hybrid query path actually works

Infinity runs as a server that speaks a Thrift-based protocol; the Python package depends on thrift, and the repository has a thrift/ directory alongside client/, src/ and python/. Your application connects over a network address and port, gets a database handle, and works with tables defined by a schema dictionary. The README example creates a table with an integer column, a varchar column and a vector column declared as "vector, 4, float".

Queries are built as a chain: you call output() to select columns, then a match function, then a conversion method. In the README the chain ends in match_dense("vec", [...], "float", "ip", 2) followed by to_pl(), which returns a Polars frame. The "ip" argument is the distance metric, and the trailing 2 is the result count. That the SDK ships polars-lts-cpu and pyarrow as dependencies is consistent with the columnar output path.

The full-text side uses BM25, which appears in the repository topics, and the example directory includes fulltext_search.py, fulltext_search_zh.py and hybrid_search.py. The presence of a Chinese-language example plus the hanziconv and nltk dependencies in pyproject.toml suggests tokenization and text normalization are handled server-side or in the SDK rather than left entirely to you. The README does not document how to configure analyzers, so treat tokenizer behaviour as something to verify against the docs site rather than assume.

Installing the Infinity server with Docker and running a first vector search

The README's get-started path is Docker with client and server as separate processes. It asks for an x86_64 CPU with AVX2 support, Linux with glibc 2.17+, Windows 10+ under WSL/WSL2, or macOS, and Python 3.11 or newer. Create the data directory and start the container. Note the --network=host flag and the raised file descriptor limit:

bash
sudo mkdir -p /var/infinity && sudo chown -R $USER /var/infinity
docker pull infiniflow/infinity:nightly
docker run -d --name infinity -v /var/infinity/:/var/infinity --ulimit nofile=500000:500000 --network=host infiniflow/infinity:nightly

The container stores its data under /var/infinity on the host. The README's Windows instructions add a step before this: enable systemd inside WSL2, install docker-ce, and if you run Docker Desktop 4.29 or later, turn on host networking under Settings, Features in development.

Install the client and pin it to the release version:

bash
pip install infinity-sdk==0.7.3

Then connect to the server on port 23817, create a table and insert two rows. The README shows this sequence:

python
import infinity

infinity_obj = infinity.connect(infinity.NetworkAddress("<SERVER_IP_ADDRESS>", 23817))
db_object = infinity_obj.get_database("default_db")
table_object = db_object.create_table("my_table", {"num": {"type": "integer"}, "body": {"type": "varchar"}, "vec": {"type": "vector, 4, float"}})
table_object.insert([{"num": 1, "body": "unnecessary and harmful", "vec": [1.0, 1.2, 0.8, 0.9]}])
table_object.insert([{"num": 2, "body": "Office for Harmful Blooms", "vec": [4.0, 4.2, 4.3, 4.5]}])

The query chains output, match_dense and to_pl. The metric is inner product, and 2 is the number of neighbours to return:

python
res = table_object.output(["*"]).match_dense("vec", [3.0, 2.8, 2.7, 3.1], "float", "ip", 2).to_pl()
print(res)

What you should see is a Polars frame with both rows, ordered by inner product against the query vector. The README does not show the expected output, so read the ordering rather than expecting an exact printout. If you prefer a binary deployment over Docker, the README points to a separate "Deploy infinity using binary" guide, and building from source has its own guide.

Where Infinity is the wrong tool

The single-binary claim in the README is about the server having no external dependencies, not about your application having no server. If your workload is a batch job that indexes a few thousand documents and exits, a client-server database adds a container lifecycle, a volume to manage and a port to expose for very little gain. An embedded index would be simpler.

The AVX2 prerequisite is a hard gate. The README states x86_64 with AVX2 support as a requirement, so older CPUs and some virtualized or emulated environments will not run it, and the README does not describe a fallback path.

Version pinning deserves attention. The README's install command pins infinity-sdk==0.7.3, matching the latest tagged release, while the Docker command pulls the nightly tag. Those two are not the same artifact stream. If you follow the README literally you pair a tagged client with a nightly server, and the README does not state a compatibility guarantee between them. pyproject.toml also constrains Python to >=3.11,<3.14, so a 3.14 interpreter will not install the SDK even though the README only says Python 3.11+.

Finally, the README does not document rollback, backup or restore procedures. The repository does contain run_snapshot_stress_test.py and example/export_data.py, which suggests snapshotting and export exist, but the README itself is silent on recovery, so that is a question for the docs site and the project's FAQ page before you put production data in it.

Infinity compared with pgvector and Qdrant

The closest comparison in the Python ecosystem is pgvector, which keeps vectors inside PostgreSQL. That approach inherits transactions, backups, replication and SQL tooling you already run, and it avoids a second datastore. The difference in approach is that pgvector extends a general relational engine, while Infinity is a purpose-built engine whose query surface is a method chain rather than SQL. Infinity's README claims hybrid search across dense, sparse, tensor and full text with ColBERT reranking in one query; pgvector's scope is vector similarity inside Postgres, so combining it with keyword ranking means using Postgres full-text search and fusing the results yourself.

Qdrant is the nearer architectural match: a standalone vector database with its own server and client. The repository's own benchmark extra lists qdrant_client and elasticsearch as optional dependencies, so the project benchmarks against both. The README's performance figures (0.1 millisecond query latency and 15K+ QPS on million-scale vector datasets, 1 millisecond and 12K+ QPS for full-text search on 33M documents) come from the project's own benchmark report, which is linked from the README. Treat those as vendor-published numbers and reproduce them on your data and hardware before relying on them.

Where Infinity differs most is the tensor and multi-vector types, which the README lists alongside dense and sparse vectors. If your retrieval uses late-interaction representations rather than a single pooled embedding, that is the feature that separates it from a plain vector store.

Licence, release cadence and upgrade cost

Infinity is Apache-2.0, and pyproject.toml declares the same licence for the infinity-sdk package. That is a permissive licence with a patent grant and no copyleft obligation on your application code. It does not answer questions about the third-party components vendored under third_party/ or the dependencies listed in vcpkg.json, so if your legal review cares about the transitive tree, that is where to look. This is not legal advice.

The release history shows v0.7.2 on 2026-07-15, v0.7.3 on 2026-08-06, and a nightly build dated 2026-01-20, with the last push to the repository on 2026-09-09. The tagged cadence is roughly monthly, which is manageable, but the nightly tag moving independently of the release tags is the real upgrade cost: any deployment that follows the README's docker pull infiniflow/infinity:nightly is tracking an unpinned artifact. Pinning the server image to a version tag that matches your SDK version is the obvious mitigation, and the README does not currently show that pattern.

The SDK's dependency list is another upgrade consideration. It pins numpy to <2.0.0, pandas to <3.0.0, pyarrow to >=21.0.0,<23.0.0 and polars-lts-cpu to <2.0.0. Those upper bounds protect you from breakage but also mean the SDK will lag behind major releases of the data stack you already use, and a conflict there will surface at pip install time.

Editorial conclusion

Adopt Infinity if your retrieval layer needs dense, sparse, tensor and full-text search in one query and you can run a server with a volume at /var/infinity. Do not adopt it if you want an in-process library with no daemon, or if you need a managed cloud endpoint, since the README describes Docker and binary deployment only. Before committing, verify the AVX2 requirement on your CPU, check that the SDK pins match your Python version (3.11 to 3.13), and confirm the nightly image tag is the one you intend to run in production.

Frequently asked questions

How do you use Infinity in Python?

Install the client with pip install infinity-sdk==0.7.3, then call infinity.connect with a NetworkAddress pointing at your server on port 23817. From there you get a database handle, create a table from a schema dictionary, insert rows, and build a query by chaining output(), match_dense() and to_pl().

How do you install Infinity?

The README's path is Docker: create /var/infinity, pull infiniflow/infinity:nightly, and run the container with --network=host and --ulimit nofile=500000:500000. The client is installed separately with pip. The README also links to a binary deployment guide and a build-from-source guide.

What are the system requirements for running the Infinity server?

The README asks for an x86_64 CPU with AVX2 support, Linux with glibc 2.17+, Windows 10+ with WSL/WSL2, or macOS, plus Python 3.11 or newer. pyproject.toml narrows the SDK's supported Python range to 3.11 through 3.13.

Can Infinity do full-text search as well as vector search?

Yes. The README describes hybrid search over dense embedding, sparse embedding, tensor and full text, and the repository topics list BM25 and full-text search. The example directory includes fulltext_search.py, fulltext_search_zh.py and hybrid_search.py.

Does Infinity support reranking of search results?

The README lists RRF, weighted sum and ColBERT among the supported rerankers, and the example directory contains a ColBERT_reranker_example folder. The README does not document the configuration keys for choosing a reranker, so check the docs site for that.

Official sources

  1. infiniflow/infinity on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/infiniflow-infinity.svg)](https://hysenlabs.com/projects/infiniflow-infinity)