# TurboQuant: KV Cache Quantization for LLM Inference

> TurboQuant is a Python library implementing the ICLR 2026 KV cache compression paper, reducing keys to 3-bit and values to 2-bit precision with Triton kernels and vLLM integration. It nearly doubles max token capacity on dense transformers, but 2-bit value quantization introduces measurable quality degradation that its own adversarial audit documents.

**0xSero/turboquant** — TurboQuant: Near-optimal KV cache quantization for LLM inference (3-bit keys, 2-bit values) with Triton kernels + vLLM integration

- Repository: https://github.com/0xSero/turboquant
- Stars: 1,788 · Forks: 197
- Language: Python
- License: GPL-3.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/0xsero-turboquant

## The Memory Bottleneck TurboQuant Was Built to Solve

Every token a large language model processes grows the key-value cache stored in GPU memory. At 30,000 tokens on a single RTX 5090, the baseline KV cache for Qwen3.5-27B-AWQ occupies enough VRAM to halve the maximum batch size or context window. The problem is not arithmetic precision but the sheer number of tensor entries that accumulate with each decoding step.

TurboQuant targets this bottleneck by applying near-optimal scalar quantization to the KV cache entries as they are written, then reconstructing approximate inner products during attention computation. The project is a research implementation of a paper presented at ICLR 2026 (arXiv:2504.19874) by Zandieh and co-authors. It is not a generic quantization library: it touches only the KV cache, leaving model weights at their original precision.

The intended users are ML engineers running vLLM inference at scale, specifically those who find that context length or concurrency is GPU-memory-bound rather than compute-bound. According to the README benchmark on a single RTX 5090, using TurboQuant on Qwen3.5-27B-AWQ freed 30 GB of KV cache across four GPUs and doubled the max token capacity from 457,072 to 914,144 tokens.

## Five Stages: How the Compression Pipeline Works

TurboQuant compresses each KV cache entry through five sequential operations. First, a random orthogonal rotation spreads information evenly across all dimensions of the key vector, so that no single dimension carries disproportionate signal before quantization. Second, Lloyd-Max optimal scalar quantization at b-1 bits is applied to the rotated values, where b is the target bit-width; this is the theoretically optimal quantizer for continuous distributions and is what the paper's distortion bounds in Theorem 1 and Theorem 3 describe.

Third, a QJL (Quantized Johnson-Lindenstrauss) projection provides one residual sign bit per dimension. Fourth, values are quantized with group quantization at 2-bit or 4-bit precision, using per-group scales and zero-points to handle distribution shifts across the head dimension. Fifth, bit-packing stores four 2-bit values per byte or two 4-bit values per byte.

The combined key estimator is unbiased: the expected inner product between a query and a compressed key equals the true inner product. The paper's Theorem 2, validated in the repository's proof.py test suite, confirms that relative bias stays below 0.1 percent under controlled conditions.

The quantization quality measurements in the README are precise. Key compression at 3-bit achieves a cosine similarity of 1.000000 on the test hardware, classified as near-lossless. Value quantization at 2-bit achieves a cosine similarity of 0.940, which the project itself identifies as the quality bottleneck. Switching to 4-bit values raises this to 0.997 and is recommended for quality-sensitive applications.

## Installing TurboQuant and Running the Validation Suite

The repository provides a setup.py that defines three optional dependency groups. The core library requires torch 2.1 or newer and numpy. The vllm extra requires vllm 0.16 or newer. The triton extra requires triton 3.0 or newer. The test extra adds pytest. Python 3.10 or newer is required across all configurations.

The package name registered in setup.py is turboquant, and the repository lives at github.com/0xSero/turboquant. The README does not document a pip install command, so the intended path is cloning the repository and installing from source. The setup.py extras_require entries show which extras to request when building from source with vLLM or Triton support.

The repository includes proof.py, which runs nine tests validating the theoretical claims from the paper. A passing run confirms that MSE distortion bounds, codebook accuracy, unbiasedness, distortion scaling, recall at rank 8, rank correlation, needle retrieval, and compression ratio all match published values. Running pytest exercises the same claims through the test harness.

The benchmark.py script generates the throughput and VRAM measurements shown in the README tables. These require actual GPU hardware. The README specifies the exact setup: vLLM 0.18.0, gpu_memory_utilization=0.90 for the RTX 5090 run and 0.92 for the RTX 3090 cluster. Reproducing the results on different hardware or vLLM versions may yield different numbers, and the README does not document rollback if a newer vLLM version changes the integration interface.

## Benchmark Readings From the RTX 5090 and RTX 3090 Cluster

The README documents two distinct hardware configurations. On a single RTX 5090 running Qwen3.5-27B-AWQ (a dense model with 4-bit weights), TurboQuant with 3-bit keys and 2-bit values increased prefill throughput from 1,804 to 1,907 tokens per second at 30,000 context tokens, a 5.7 percent gain. Decode throughput rose from 1.264 to 1.303 tokens per second, a 3.1 percent gain. Max token capacity doubled from 457,072 to 914,144.

On an eight-GPU RTX 3090 cluster running the MoE model Qwen3.5-35B-A3B (pruned, 205 experts, tensor parallelism 8), the results are more constrained. That model has 10 full-attention layers and 30 linear-attention layers. TurboQuant compresses only the full-attention layers, which hold 40 percent of the total KV cache. The measured savings are consistently 30.9 percent of total KV cache across all context lengths from 8,000 to 131,000 tokens. Context capacity increased from 1,411,680 to 2,043,808 tokens, a 1.45x multiplier rather than the near-2x seen on the dense model.

At 131,000 context tokens on the MoE cluster, VRAM utilization stayed essentially flat at 22,306 MB per GPU, because the compressed KV cache for the full-attention layers occupied only about 190 MB per GPU. No CPU or KV offloading was used. All data remained in VRAM.

## What the Adversarial Audit Found

The repository includes an audit_claims.py script and an adversarial section in the README that rates eight specific claims from the paper and marketing materials. Several findings are worth reading before deploying the library.

The advertised 5.1x compression ratio is described as misleading. The accurate figure, according to the audit, is approximately 4.6x at 4,000 tokens and about 5x at 32,000 tokens and above, because the Pi and S matrices from the rotation step and the ring buffer overhead are not counted in the 5.1x figure. The measured value in proof.py is 4.41x at head_dim=256.

The needle-in-haystack retrieval pass is rated as true but trivial. The test uses the query as a copy of the key, which the README notes is far easier than real LLM inference where queries are not copies of their corresponding keys.

The claim that TurboQuant is faster at 30,000 tokens context is rated as within noise, because the measurement was a single run, and the total wall time with TurboQuant was actually slightly slower than baseline.

The 200,000-context claim is rated as unverified. The model did not crash, but output quality was not checked. Finally, decode compression is noted as saving storage but not compute, because the implementation dequantizes the full history to float32 on each decode step.

## Where TurboQuant Cannot Help

Three situations make TurboQuant the wrong tool. The first is MoE models with a high fraction of linear-attention layers. On the Qwen3.5-35B-A3B configuration tested in the README, 75 percent of KV cache entries come from linear-attention layers that TurboQuant cannot compress. The savings are bounded by the fraction of full-attention layers in the model. On a pure dense transformer, the README projects 77 percent savings; on the tested MoE, only 30.9 percent materialized.

The second situation is quality-sensitive applications at 2-bit value precision. The measured cosine similarity of 0.940 for combined 3-bit keys and 2-bit values means inner product estimates carry a small but consistent error. The README recommends 4-bit values for quality-sensitive use, which reduces savings but raises cosine similarity to 0.997.

The third is any codebase not using vLLM. TurboQuant's integration is built against the vLLM serving framework. The base library installs without vLLM, and proof.py and benchmark.py can run without it, but the serving path assumes vLLM 0.16 or newer. The project is labeled version 0.1.0 and classified as Alpha in its setup.py classifiers, meaning the API and vLLM integration interface may change without backward-compatibility guarantees.

RaBitQ is an alternative approach to quantized attention that also targets KV cache compression using binary random projections. The two methods differ in their quantization strategy: TurboQuant applies Lloyd-Max optimal scalar quantization with a QJL residual, while RaBitQ relies on random binary codes without the per-group value quantization stage that TurboQuant uses for its value path.

## License, Maintenance, and Upstream Dependency Risk

TurboQuant is released under GPL-3.0. This license requires that any software that incorporates or links TurboQuant as part of a serving pipeline also be distributed under GPL-3.0. For commercial inference infrastructure that is not intended to be open-source, this is a hard constraint. The setup.py confirms the GPL-3.0 classifier.

The repository had its last push on 2026-09-03, placing it within a few weeks of the review date. There are no GitHub releases, and the version in setup.py is 0.1.0, signaling early-stage status. The test suite is small (nine paper validation tests in proof.py), so there is limited coverage of edge cases outside the documented hardware configurations.

The vLLM dependency introduces a significant external risk. vLLM's internal APIs change frequently between minor versions. The README documents testing against vLLM 0.18.0 specifically, and makes no guarantee of compatibility with earlier or later releases. Any upgrade to vLLM in a production environment would need to be tested against TurboQuant's integration code before deployment.

## Conclusion

Engineers deploying dense transformer models on vLLM who need more context in the same VRAM are the primary audience. Those running MoE architectures should expect substantially lower savings, because TurboQuant cannot compress linear-attention layers, which make up the majority of KV cache in models like Qwen3.5-35B-A3B. Before depending on the published compression ratios, run the adversarial audit in audit_claims.py: several claims in the README are either misleading or unverified at the 200k context range. The GPL-3.0 license requires that any derivative serving code also be released under GPL-3.0, which rules out most commercial deployments without a separate agreement.

## FAQ

### What is TurboQuant?

TurboQuant is a Python library implementing the ICLR 2026 paper on near-optimal KV cache quantization for large language model inference. It compresses the key-value cache to 3-bit keys and 2-bit values using Lloyd-Max scalar quantization and QJL projections, reducing GPU memory usage without changing model weights.

### What models use TurboQuant?

The README documents benchmarks on Qwen3.5-27B-AWQ (a dense model) on a single RTX 5090 and Qwen3.5-35B-A3B MoE (a mixture-of-experts model) on an eight-GPU RTX 3090 cluster. TurboQuant compresses only the full-attention layers of a model, so any model with linear-attention layers will see proportionally lower memory savings.

### How do I install TurboQuant?

TurboQuant installs via pip from the repository. Run pip install turboquant[vllm,triton] to include the vLLM integration and Triton kernel extras. Python 3.10 or newer is required, along with torch 2.1 or newer as a core dependency.

### How do I use TurboQuant?

After installation, the benchmark.py script in the repository demonstrates integration with vLLM, and proof.py validates the paper's theoretical claims on your hardware. The README does not document a standalone Python API for custom inference frameworks outside vLLM.

## Sources

- [0xSero/turboquant on GitHub](https://github.com/0xSero/turboquant)
- [Issues](https://github.com/0xSero/turboquant/issues)
- [License: GPL-3.0](https://github.com/0xSero/turboquant/blob/main/LICENSE)
- [README](https://github.com/0xSero/turboquant/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/0xsero-turboquant
