Library / SDK
RyanCodrai/turbovec avatar
RyanCodrai/turbovec

turbovec: TurboQuant-Powered Vector Index with Python and Rust APIs

A vector index built on TurboQuant, written in Rust with Python bindings.

17,262 stars1,476 forksPythonMIT

At a glance

What is it?
turbovec is a Rust vector index with Python bindings that applies Google Research's TurboQuant quantization algorithm to fit a 10-million-document corpus in 4 GB instead of 31 GB, while searching faster than FAISS IndexPQFastScan. It operates entirely locally with no training phase and supports incremental persistence.
Who is it for?
turbovec is the right fit for RAG pipelines, semantic search systems, and recommendation engines where memory is a hard constraint and query latency matters. The pure-local design suits air-gapped environments and privacy-sensitive deployments.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 16 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

What turbovec Is and the Problem It Addresses

turbovec is a vector index built in Rust with Python bindings, designed around Google Research's TurboQuant algorithm, a data-oblivious quantizer that the README describes as achieving near-optimal distortion with no separate training phase. The practical outcome the README states is this: a 10-million-document corpus that takes 31 GB of RAM as float32 fits in 4 GB with turbovec at 4-bit quantization.

The target audience is anyone building retrieval-augmented generation pipelines, semantic search systems, or recommendation engines where memory footprint and query speed are constraints. The README names the motivating scenario directly: RAG where privacy, memory, or latency matters. Because turbovec runs entirely locally with no external service, it suits air-gapped environments and systems where data must not leave the machine or VPC.

The library is published under the MIT license. The PyPI package is turbovec. The Rust crate is also named turbovec on crates.io. The repository has no GitHub releases; the most recent push was on 2026-09-13.

TurboQuant Algorithm: Online Ingest and SIMD Search

TurboQuant is the quantization algorithm from a Google Research paper (arxiv.org/abs/2504.19874). The README describes it as data-oblivious: the quantizer does not require a separate training pass over the corpus before indexing begins. Vectors can be added and they are indexed immediately, with no rebuild step as the corpus grows.

The search kernels are hand-written SIMD code. The README lists specific targets: NEON SDOT/SMMLA on ARM, AVX-512 VNNI and vpermb on x86, with AVX2 and scalar fallbacks. The README states that these kernels beat FAISS IndexPQFastScan in every measured configuration, averaging 3.4 times faster at 4-bit and 23% faster at 2-bit across eight measurement cells on both ARM and x86 architectures.

For recall comparison against FAISS IndexPQ, the README's recall section reports that calibrated TurboQuant (TQ+) beats FAISS at R@1 on three of four measurement cells for OpenAI d=1536 and d=3072 embeddings, with a 0.7-point trail in one configuration. On GloVe d=200, which the README describes as the harder regime due to lower dimensionality, TQ+ lands ahead of FAISS at R@1 at both bit widths.

Installing and Indexing Vectors in Python

Install turbovec from PyPI:

bash
pip install turbovec

The Python API provides TurboQuantIndex for position-based access and IdMapIndex for stable external IDs. The README's primary Python example shows adding vectors and searching:

python
from turbovec import TurboQuantIndex

index = TurboQuantIndex(dim=1536, bit_width=4)
index.add(vectors)
index.add(more_vectors)

scores, indices = index.search(query, k=10)

index.write("my_index.tv")
loaded = TurboQuantIndex.load("my_index.tv")

index.sync("my_index.tv")   # after more changes: durable incremental save

The README states that vectors and query must be 2-D float32 arrays. Other dtypes are rejected rather than silently converted, so the README advises casting with np.asarray(x, dtype=np.float32) before calling add or search.

The sync method persists only what changed since the last sync call: one fsync per call, crash-safe at any byte. The README notes that a removal or a small append costs milliseconds regardless of index size. The write and load methods remain available for whole-file snapshots.

IdMapIndex for Stable IDs and Filtered Search

When vectors are deleted and re-added over time, TurboQuantIndex positions shift. IdMapIndex solves this by mapping user-controlled uint64 IDs to internal slots:

python
import numpy as np
from turbovec import IdMapIndex

index = IdMapIndex(dim=1536, bit_width=4)
index.add_with_ids(vectors, np.array([1001, 1002, 1003], dtype=np.uint64))

scores, ids = index.search(query, k=10)
index.remove(1002)

index.write("my_index.tvim")
loaded = IdMapIndex.load("my_index.tvim")

Removal is O(1) by ID. Persistence uses the .tvim extension for IdMapIndex files and .tv for TurboQuantIndex files.

Filtered search restricts results to a caller-provided allowlist without over-fetching. The README shows passing an allowlist of numpy uint64 IDs to search():

python
scores, ids = idx.search(query, k=10, allowlist=allowed)

Filtering happens inside the SIMD kernel at 32-vector block granularity. Blocks with no allowed slots are short-circuited before any lookup or scoring work. The README states that selective allowlists avoid most of the SIMD cost rather than paying it and discarding the result. The output length is min(k, n_allowed), where n_allowed counts distinct allowed vectors.

Rust API and Framework Integrations

Add the crate with:

bash
cargo add turbovec

The Rust API mirrors the Python surface: TurboQuantIndex::new(dim, bit_width) creates the index, add feeds vectors, search returns scores and positions, and write/load handle persistence. IdMapIndex provides the same stable-ID and remove semantics as the Python version.

For framework integration, turbovec ships optional extras that replace the default in-memory vector stores in common RAG frameworks without changing public surfaces or pipeline wiring. The README lists four integrations with their install commands and the specific store they replace:

- LangChain: pip install turbovec[langchain], replaces langchain_core.vectorstores.InMemoryVectorStore - LlamaIndex: pip install turbovec[llama-index], replaces llama_index.core.vector_stores.SimpleVectorStore - Haystack: pip install turbovec[haystack], replaces haystack.document_stores.in_memory.InMemoryDocumentStore - Agno: pip install turbovec[agno], replaces agno.vectordb.lancedb.LanceDb

The Cargo.toml workspace structure separates the Rust library (turbovec/) from the Python bindings (turbovec-python/). The downstream-smoke example under examples/ is intentionally excluded from the workspace to test the library exactly as an external downstream user would experience it.

Limitations: No Training Flexibility and No Clustering

TurboQuant is data-oblivious: the quantizer uses no information about the specific distribution of the corpus vectors. The README notes that GloVe d=200 is the harder regime because at low dimensionality the asymptotic assumption in the Beta distribution is looser. Applications with very low-dimensional embeddings may see more recall variation than applications using standard 1536-dimensional OpenAI embeddings.

The library targets single-machine deployments. The README describes no clustering, sharding, or distributed search capability. For a corpus that exceeds a single machine's memory or requires horizontal scaling, turbovec does not provide built-in solutions.

The project has no GitHub releases at the time of this writing. The most recent push was on 2026-09-13. Developers relying on semver-pinned releases must use PyPI or crates.io version tags rather than GitHub release tags.

Comparison with FAISS and Annoy

FAISS is a Meta AI library for dense vector search that supports many index types, including flat exhaustive search, IVF approximate search, and product quantization. FAISS requires a training step for most non-flat index types, during which it learns a codebook from a sample of the corpus. TurboQuant's data-oblivious design eliminates that step. The README benchmarks specifically against FAISS IndexPQFastScan, the production-grade PQ variant with fast SIMD scanning.

Annoy (Approximate Nearest Neighbors Oh Yeah) is a library from Spotify that builds random projection trees for approximate search. Annoy is designed around read-heavy workloads where the index is built once and queried many times. It does not support online ingest: adding a new vector requires rebuilding the index trees. turbovec supports online ingest by design, making it better suited to scenarios where new vectors arrive continuously without rebuild windows.

Editorial conclusion

turbovec is the right fit for RAG pipelines, semantic search systems, and recommendation engines where memory is a hard constraint and query latency matters. The pure-local design suits air-gapped environments and privacy-sensitive deployments. The project has no GitHub releases; the most recent repository activity was on 2026-09-13. Developers who need a production-grade managed index with audit logging and cluster scaling should evaluate dedicated vector database offerings instead.

Frequently asked questions

What is turbovec?

turbovec is a Rust vector index with Python bindings that uses Google Research's TurboQuant algorithm to quantize and search dense embedding vectors. It fits a 10-million-document corpus in 4 GB instead of 31 GB as float32 and requires no separate training phase.

How do I use turbovec?

Install with pip install turbovec, create a TurboQuantIndex with the embedding dimension and bit width, call add() with float32 numpy arrays, then call search() with a query vector and k to get top-k results. For stable external IDs, use IdMapIndex with add_with_ids() instead.

What does turbovec do?

turbovec quantizes high-dimensional embedding vectors using the TurboQuant algorithm and searches them with hand-written SIMD kernels. It supports online ingest with no rebuild step, incremental persistence via sync(), and filtered search against an ID allowlist.

Is turbovec from Google?

turbovec is an independent Rust implementation of the TurboQuant algorithm from a Google Research paper (arxiv.org/abs/2504.19874). The project is not an official Google product; the repository is published by RyanCodrai under the MIT license.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/ryancodrai-turbovec.svg)](https://hysenlabs.com/projects/ryancodrai-turbovec)
Community notes

Community notes