# KVarN: a vLLM KV-cache quantization backend that keeps FP16 throughput

> KVarN is a vLLM fork from Huawei CSL that quantizes the KV cache to int4 with no calibration, selected by one dtype flag. It targets agentic and long-context serving, where the cache, not the weights, is what runs out first.

**huawei-csl/KVarN** — KVarN is a native vLLM KV-cache quantization backend for your agents: 3-5x more context, throughput above FP16, and FP16-level accuracy. Calibration-free, one flag.

- Repository: https://github.com/huawei-csl/KVarN
- Website: https://arxiv.org/abs/2606.03458
- Stars: 508 · Forks: 37
- Language: Python
- License: Apache-2.0
- Published: 2026-09-20 · Updated: 2026-09-20 · Language: en
- Canonical page: https://hysenlabs.com/projects/huawei-csl-kvarn

## What KVarN changes about the KV cache, and who it is for

The KV cache grows with every token in the context and every concurrent request, so on long-context or agentic workloads it is usually the thing that exhausts GPU memory first. Quantizing it is the obvious answer, and the README is direct about why people avoid it: citing the vLLM TurboQuant blog, it states that existing methods report 40 to 52 percent lower throughput for 2.3 to 3.7 times the capacity, and that aggressive low-bit quantization tends to cost accuracy. Losing speed and quality at the same time is why the feature stays off in production.

KVarN is an attempt to keep both. It is a native vLLM attention backend, so the integration point is the KV-cache dtype rather than a wrapper around the engine. The README claims 3 to 5 times more KV-cache capacity and up to roughly 1.3 times the throughput of FP16, with FP16-level accuracy, and says it is calibration-free: no dataset pass, no per-model scale search. The intended audience is narrow and clear. If you serve long contexts or many concurrent agent sessions and you already run vLLM, this is aimed at you. If your contexts are short, the cache is not your bottleneck and the whole exercise is pointless.

## How the backend works: tile size, presets and the MLA path

The mechanism visible in the README is a mapping between vLLM's paged KV cache and KVarN's quantization tiles. One vLLM block is one KVarN tile, so the page size equals --block-size, and the preset name encodes the granularity: kvarn_k4v2_g128 means 128-token groups, kvarn_k4v2_g64 means 64. Both 128 and 64 are supported, and the README calls 128 the design point. The 64 preset gives finer quantization granularity at the cost of a little KV capacity, because more per-tile scale overhead is paid per token, at essentially the same throughput. That is a real trade-off stated as one, not a free knob.

KVarN runs in float16 compute, and the kernels are Triton, JIT-compiled at runtime. That last detail explains the install shape: the upstream precompiled wheel supplies the compiled parts, and the KVarN kernels compile on first use rather than at build time.

The MLA path is the most interesting design decision. On Multi-head Latent Attention models, the same --kv-cache-dtype flag routes to a latent path that quantizes the compressed KV latent to int4, with no code or environment change, and the fast path is on by default. The README claims this makes KVarN the first vLLM-compatible sub-8-bit KV-cache quantization method to support MLA-based models. Note what the README itself concedes for this case: MLA's latent is already tiny, so KVarN is not a latency play there. The reported GLM-4.7-Flash numbers show 377 tok/s against 401 for bf16, or 0.94 times, with 865K tokens of capacity against 313K, or 2.77 times, and AIME25 accuracy at 53.3 percent for both. The gain is capacity, not speed.

Hybrid models that interleave linear-attention or Mamba layers with full-attention layers are handled by compressing only the layers that hold a KV cache. The recurrent state of the other layers is left alone, and the fp16 decode pool is sized from the full-attention layer count, so the default flags work without manual pool tuning.

## Installing KVarN and serving a first model

KVarN ships as a vLLM fork, and the README says to install it like vLLM. Clone the repository, then install with the upstream precompiled wheel. The VLLM_USE_PRECOMPILED=1 variable is what selects that wheel; the KVarN kernels themselves are Triton and compile at runtime, so the build is not doing that work.

```bash
git clone https://github.com/huawei-csl/KVarN.git
cd KVarN
VLLM_USE_PRECOMPILED=1 pip install -e .
```

The README's Python example constructs a vLLM LLM object with three settings that matter: float16 compute, the KVarN KV-cache dtype, and a block size of 128 to match the g128 preset.

```python
from vllm import LLM, SamplingParams

llm = LLM(
    model="Qwen/Qwen3-32B",
    dtype="float16",
    kv_cache_dtype="kvarn_k4v2_g128",
    block_size=128,
)
print(llm.generate("Explain KV-cache quantization in one sentence.",
                    SamplingParams(max_tokens=64))[0].outputs[0].text)
```

Serving uses the same three settings as command-line flags. The README gives this exact invocation.

```bash
vllm serve Qwen/Qwen3-32B --dtype float16 --kv-cache-dtype kvarn_k4v2_g128 --block-size 128
```

One operational detail is worth reading twice. KVarN reaches its full KV-cache capacity only when there is room to amortize a small fixed decode workspace. On multi-GPU or generous --gpu-memory-utilization setups that happens automatically. On a tight single-GPU budget, vLLM's CUDA-graph memory profiler can over-reserve and shrink the KV pool, and the README's remedy is to set VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS=0 or raise --gpu-memory-utilization. If you install this on one card and see less capacity than the headline number, that variable is the first thing to check.

## Speculative decoding and weight quantization together

The README states that KVarN composes with speculative decoding, including Multi-Token Prediction, and the flags are unchanged: pass --speculative-config as you normally would, alongside the KVarN dtype and block size.

```bash
vllm serve Qwen/Qwen3.6-27B \
    --dtype bfloat16 \
    --kv-cache-dtype kvarn_k4v2_g128 \
    --block-size 128 \
    --speculative-config '{"method":"mtp","num_speculative_tokens":3}'
```

The correctness argument in the README is specific: the verify step attends over the full cached context, which KVarN reconstructs from the quantized cache, and a block is written to the quantized cache only once all of its tokens are accepted. Rejected draft tokens therefore never enter history. That is the right place to put the commit point, and it is the kind of detail that decides whether speculative decoding and cache quantization can coexist at all.

Weight quantization is a separate axis. KVarN quantizes the KV cache independently of the model weights, so it stacks with weight-quantized checkpoints such as compressed-tensors or AWQ INT4, and with MTP at the same time. The README says this combination was validated on Qwen3.6-27B. If you already run an INT4 checkpoint, you do not have to choose between the two.

## Where KVarN is the wrong tool, and what to check first

The clearest limitation is one the README states for MLA models and that generalizes: capacity is the product, not latency. On GLM-4.7-Flash the reported throughput is 0.94 times bf16. If your problem is tokens per second on a model that already fits, KVarN does not solve it, and on MLA it slightly costs you. The dense claim of up to roughly 1.3 times FP16 throughput is a different regime and a different model.

The second constraint is that this is a fork, not a plugin. Installing it replaces your vLLM with a build pinned in the README badge to vLLM v0.23.0. The pyproject.toml describes the package as name = "vllm" with the vLLM Team as author and the vLLM homepage, and setup.py carries the upstream build machinery including the Rust frontend and CUDA_HOME detection. If your deployment tracks upstream vLLM releases, or if anything else in your stack depends on a specific vLLM version, that is the cost of entry, and the repository does not document a rollback path.

The third is hardware. The README's examples assume CUDA-class GPUs and tensor parallelism, and the capacity tip is written around vLLM's CUDA-graph memory profiler. Nothing in the repository describes a CPU or non-CUDA serving path.

Before adopting, verify three things on your own workload: that your model architecture is one of the supported shapes (dense, MLA, or hybrid with full-attention layers), that the capacity you actually get matches the claim once the decode workspace is accounted for, and that accuracy on your task holds, since the published parity numbers are for AIME25 on specific models.

## The honest alternative: TurboQuant and plain FP16

The README positions KVarN against vLLM TurboQuant, and it does so with the competitor's own reported numbers: 40 to 52 percent lower throughput for 2.3 to 3.7 times capacity, plus an accuracy cost at low bit widths. The claimed edge is up to roughly 2.4 times TurboQuant's throughput at the same capacity and higher accuracy. Those are the project's own comparisons, not independent results, and the README does not describe how TurboQuant's quantization is implemented, so the difference in approach is not spelled out beyond the variance-normalization idea in the project name and the tile-based presets.

The alternative that needs no argument is plain FP16. If your KV cache already fits, FP16 is faster than KVarN on MLA and simpler everywhere. KVarN only pays off when the cache is the binding constraint and the extra tokens are worth more to you than the last few percent of throughput.

## Licence and maintenance

KVarN is Apache-2.0, and the LICENSE file is at the repository root. Because the project is a vLLM fork, the licence question is not only KVarN's: the pyproject.toml declares license = "Apache-2.0" with license-files = ["LICENSE"], and the setup.py header carries SPDX-License-Identifier: Apache-2.0 with a copyright line for the vLLM project contributors. Apache-2.0 is permissive and includes an explicit patent grant, which matters for a quantization kernel, but this is a description of the files, not legal advice; read LICENSE and the upstream notices before shipping.

The repository is not archived, and the last push was on 2026-06-22. No releases were retrieved. That means upgrades arrive as commits on main rather than as tagged versions, so pinning to a commit hash is the only way to make a deployment reproducible. Budget for the rebase cost against upstream vLLM: the fork carries csrc, rust, cmake and docker directories, and the README badge names vLLM v0.23.0 as the base, so moving forward means reconciling with whatever upstream has changed in the attention backend since.

## Conclusion

Adopt KVarN if you serve long-context or agentic workloads on vLLM, your KV cache is the binding memory constraint, and you can accept running a vLLM fork pinned to a commit. Do not adopt it if you are on MLA and want lower latency, if your contexts are short enough that the cache already fits, or if you cannot absorb a fork's upgrade cost. Verify three things first: that your model architecture matches a supported path, that VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS=0 or a higher --gpu-memory-utilization actually recovers the capacity you expect on your GPU, and that accuracy holds on your own evaluation set rather than on AIME25.

## FAQ

### What is KVarN and which models does it support?

KVarN is a native vLLM KV-cache quantization backend that runs in float16 compute and is selected with the --kv-cache-dtype flag. The README describes support for dense models, Multi-head Latent Attention models through a latent path, and hybrid models that interleave linear-attention or Mamba layers with full-attention layers.

### How do I install KVarN for vLLM?

Clone the repository and install it like vLLM, using VLLM_USE_PRECOMPILED=1 pip install -e . so the upstream precompiled wheel is used. The KVarN kernels are Triton and compile at runtime.

### Does KVarN need calibration?

No. The README describes KVarN as calibration-free, and the only change needed at serving time is the KV-cache dtype flag plus a matching block size.

### Why is my KV-cache capacity lower than expected with KVarN?

The README says KVarN needs room to amortize a small fixed decode workspace, and that on a tight single-GPU budget vLLM's CUDA-graph memory profiler can over-reserve and shrink the KV pool. It suggests setting VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS=0 or raising --gpu-memory-utilization.

### Does KVarN work with speculative decoding and weight-quantized checkpoints?

The README states that KVarN is compatible with speculative decoding including MTP, passed through --speculative-config, and that it composes with weight-quantized checkpoints such as compressed-tensors or AWQ INT4 because the KV cache is quantized independently of the weights.

## Sources

- [huawei-csl/KVarN on GitHub](https://github.com/huawei-csl/KVarN)
- [Issues](https://github.com/huawei-csl/KVarN/issues)
- [License: Apache-2.0](https://github.com/huawei-csl/KVarN/blob/main/LICENSE)
- [Project website](https://arxiv.org/abs/2606.03458)
- [README](https://github.com/huawei-csl/KVarN/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/huawei-csl-kvarn
