Model or dataset
tonbistudio/turboquant-pytorch avatar
tonbistudio/turboquant-pytorch

turboquant-pytorch: KV Cache Compression That Admits What Fails

From-scratch PyTorch implementation of Google's TurboQuant (ICLR 2026) for LLM KV cache compression. 5x compression at 3-bit with 99.5% attention fidelity.

1,049 stars140 forksPythonMIT

At a glance

What is it?
A from-scratch PyTorch implementation of Google's TurboQuant, with a V3 branch that drops the paper's QJL stage and reports honest generation results. It is a research harness, not a drop-in cache.
Who is it for?
Adopt turboquant-pytorch if you are researching KV cache quantization and want a readable PyTorch reference with a generation test that reports failures. Do not adopt it as a drop-in cache for a production serving stack: the README states 3-bit keys break generation, V2 stored tensors 38% larger than uncompressed, and no release exists.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 159 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

What turboquant-pytorch solves, and for whom

The KV cache grows linearly with context length and is stored in fp16. Compressing it is the only way to fit longer contexts on a fixed GPU. TurboQuant, the paper this repository implements, attacks that with a random rotation followed by scalar quantization. The repository's own framing is narrower than a serving library: it is a from-scratch PyTorch implementation of the algorithm, tested on Windows with NVIDIA GPUs, and its most valuable output so far is a negative result. The README states that the paper's key innovation, QJL, "actually hurts in practice" for KV cache, and that the author built an improved version (V3) informed by findings from 8+ independent community implementations. That makes the audience researchers and engineers who want to read the math in Python, reproduce the attention-score tables, and decide for themselves whether the bit allocation is worth it. It is not aimed at someone who wants to add two lines to a vLLM config.

Random rotation, Lloyd-Max centroids, and why QJL was dropped

The core mechanism is stated plainly in the README. Each vector is multiplied by a random orthogonal matrix, which makes every coordinate follow a predictable bell-curve distribution. An optimal scalar quantizer (Lloyd-Max) then rounds each coordinate independently to the nearest precomputed centroid. Quantizing means normalize, rotate, round each coordinate, store indices plus the norm. Dequantizing means look up centroids, reverse the rotation, restore the norm. The solver lives in lloyd_max.py, and the V2 classes (TurboQuantMSE, TurboQuantProd) live in turboquant.py.

The interesting design decision is what V3 removes. The paper's stage two, QJL, stores one bit of sign information per residual to make inner-product estimates mathematically unbiased. The README's argument against it is specific: QJL is unbiased for raw inner products, but attention runs scores through softmax, and softmax exponentially amplifies variance. MSE-only has biased inner products but lower variance, and lower variance wins after softmax. The README cites scos-lab measuring +300% error with QJL versus +7.6% without on GPT-2, and reports the project's own V2 with QJL at 0/27 generation tests passed against V3 at 18/18. It also notes QJL does work for vector search, where there is no softmax, which is the paper's other use case. That is a fair reading of the trade-off rather than a dismissal of the paper. V3's other changes are asymmetric K/V bit allocation, bit-packed storage, and protected first and last layers.

Installing turboquant-pytorch and running a first real test

The package is not on PyPI in the README given; the install path is from the repository. Requirements are Python 3.10+, a CUDA-capable NVIDIA GPU (tested on an RTX 3060 with 12GB), and Windows 11 or Linux. Install the dependencies first:

bash
pip install -r requirements.txt

For editable local development, the README gives:

bash
pip install -e .

If your environment needs a CUDA build of PyTorch, the README points at the cu128 index:

bash
pip install torch --index-url https://download.pytorch.org/whl/cu128

The first real use is the generation test, which the README labels the real test because it checks whether the model produces correct text. It hides a fact ("The secret project code name is AURORA-7749") in a long document and asks the model to find it, logging compressed token counts so you can confirm compression actually happened:

bash
python -m turboquant.generation_test

The first run downloads Qwen2.5-3B-Instruct, roughly 2GB, and tests multiple configs across context lengths. If you want the attention-score comparison instead, or a run that needs no model download, the README offers these:

bash
python -m turboquant.validate_v3
python -m turboquant.test_turboquant

validate_v3 compares V2 and V3 attention accuracy side by side. test_turboquant validates the core algorithm against the paper's theoretical bounds without loading a model, which is the cheapest way to confirm the install works before pulling 2GB of weights.

The compression numbers do not survive the generation test

This is the section that matters most, and the repository is unusually direct about it. The README carries a correction dated 2026-03-30 stating that an earlier version claimed "18/18 perfect generation at 5x compression," and that this came from a bugged test where residual_window=0 caused no compression to happen. The corrected generation table tells a different story. K6/V4 with a 128-token fp16 residual window gives about 2x compression with EXACT output at both 2K and 4K context. K8/V4 gives about 1.6x, also EXACT. At 4-bit keys the model finds the needle at short context but garbles it slightly, dropping the hyphen. At 3-bit keys, the README states generation is broken. Without a residual window, 3-4 bit compression produces garbage, same as V2.

The attention-score table is where the 99.5% figure comes from, and the README explicitly warns against reading it as a guarantee: high attention score similarity does not guarantee working generation. V3 K4/V2 reaches 5.1x compression with 0.9996 cosine similarity and 94% top-1 match on captured KV tensors, but that measurement bypasses V3Cache. So the same configuration that scores beautifully on attention accuracy is the one that returns PARTIAL or MISS in generation. If you only read the cosine column you will ship something broken. There is a second, quieter limitation: V2 stored tensors that were 38% larger than uncompressed, which is why bit-packed storage is listed as a V3 improvement. Theoretical compression ratios and on-disk or on-GPU memory are different numbers.

RaBitQ and the difference between unbiased and low-variance

The closest comparison in the README is RaBitQ, which the search data shows people asking about directly. Both are vector quantization schemes for approximate inner products, and both lean on random rotation to make coordinates well behaved before quantizing. The difference this repository exposes is what happens after quantization. RaBitQ-style approaches keep a correction term to make the inner-product estimate unbiased, which is the same instinct as TurboQuant's QJL stage. The README's finding is that for attention specifically, unbiasedness is the wrong target: softmax amplifies variance, so an estimator with lower variance and some bias beats an unbiased estimator with more noise. That is a real difference in approach, not a feature checklist. It also means the comparison is workload-dependent. The README concedes QJL works for vector search, where there is no softmax, and suggests it may work with non-softmax attention such as sigmoid, linear or gated. If your retrieval stack uses inner-product search over compressed embeddings, the V2 path in this repository is closer to the paper and may be the more relevant code to read.

Maintenance, licence, and what upgrading costs

The repository is not archived, and the last push was on 2026-04-23. That is roughly five months before today, so treat it as a project that was moving earlier this year and has been quiet since. There are no retrieved releases, and pyproject.toml pins version 0.1.0, so there is no tagged artifact to depend on. Installation is from source, which means an upgrade is a git pull plus re-running pip install -e . and re-running the test suite, not a version bump in a lockfile. The dependency set is small but not light: torch>=2.0.0, scipy>=1.10.0, transformers>=4.40.0, accelerate>=0.25.0 and bitsandbytes>=0.43.0. Those lower bounds are open-ended, so an upgrade of transformers or torch can change behaviour underneath the quantizer without any change to this repository, and the generation test is the only thing that would catch it. Licence is MIT, declared both in the LICENSE file and as license = {text = "MIT"} in pyproject.toml. MIT is permissive and places no copyleft obligation on your own code, but the paper implementation itself is a separate matter, and nothing here is legal advice.

Where turboquant-pytorch is the wrong tool

If you need a KV cache compressor in a serving path today, this is not it. The README's own results say 3-bit keys break generation, the working configuration is about 2x compression with a 128-token fp16 window, and the attention validation runs on captured tensors rather than through the cache class. A 2x figure that requires keeping 128 recent tokens in fp16 also scales badly with batch size, since every sequence carries that window. The Windows-first testing note is a second constraint: the README says it also works on Linux, but the tested configuration is Windows 11 with an RTX 3060 at 12GB, so anything you run on a different stack is your own validation. The third case is anyone who wants a stable dependency. With version 0.1.0, no releases, and a correction notice rewriting the headline result, the API surface should be expected to move. Use it to understand the algorithm, to reproduce the V2 versus V3 comparison, or to test whether QJL helps on a non-softmax attention variant. Do not use it as the compression layer under a latency-sensitive endpoint.

Editorial conclusion

Adopt turboquant-pytorch if you are researching KV cache quantization and want a readable PyTorch reference with a generation test that reports failures. Do not adopt it as a drop-in cache for a production serving stack: the README states 3-bit keys break generation, V2 stored tensors 38% larger than uncompressed, and no release exists. Before anything else, run python -m turboquant.generation_test and confirm the compressed token counts are non-zero on your own model and GPU, because the project's own correction note traces an earlier false pass to exactly that check.

Frequently asked questions

What is TurboQuant?

TurboQuant is a vector quantization algorithm for compressing LLM key-value caches, published as an ICLR 2026 paper. It rotates each vector with a random orthogonal matrix so coordinates follow a predictable distribution, then applies a Lloyd-Max optimal scalar quantizer to each coordinate independently.

What is Google TurboQuant?

It is the same algorithm: the paper this repository implements is described in the README as Google's vector quantization algorithm for compressing LLM key-value caches. The repository's V3 version departs from the paper by removing the QJL residual-correction stage.

What are the key differences between TurboQuant and RaBitQ?

The README does not compare the two directly, so the traceable difference is about estimator design. TurboQuant's paper adds QJL to make inner-product estimates unbiased, while this repository's V3 drops it because softmax amplifies variance, and low variance beats unbiasedness after softmax. The README notes QJL still works for vector search, where there is no softmax.

Official sources

  1. Issues
  2. License: MIT
  3. README
  4. tonbistudio/turboquant-pytorch on GitHub
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/tonbistudio-turboquant-pytorch.svg)](https://hysenlabs.com/projects/tonbistudio-turboquant-pytorch)
Community notes

Community notes