Model or dataset
huawei-csl/KVarN avatar
huawei-csl/KVarN

KVarN is a vLLM fork that installs itself under the name vllm

KVarN is a native vLLM KV-cache quantization backend for your agents: 3-5x more context, throughput above FP16, and FP16-level accuracy. Calibration-free, one flag.

507 stars37 forksPythonApache-2.0

At a glance

What is it?
A KV-cache quantization backend for long-context serving, claiming more cache capacity and accuracy at parity with half-precision compute, delivered as a fork of the inference engine rather than a plugin. The technical claims are specific and mostly checkable, and one of them inverts on the one model where a full table is published.
Who is it for?
KVarN is worth evaluating if you are serving long contexts on vLLM and have measured that your KV pool is what limits you, because it is the rare quantization backend that does not ask you to trade throughput for memory.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 106 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.

Editorial analysis

It is a fork, and the package installs as vllm

The distribution strategy is stated in one line: KVarN ships as a vLLM fork, and you install it the way you install vLLM.

bash
# 1. Clone
git clone https://github.com/huawei-csl/KVarN.git
cd KVarN

# 2. Install (uses the upstream precompiled wheel; KVarN kernels are Triton, JIT-compiled at runtime)
VLLM_USE_PRECOMPILED=1 pip install -e .

That choice is visible in the metadata. The project name in the packaging file is `vllm`, the authors are listed as the vLLM team, the description is the upstream one about a high-throughput inference and serving engine, and the homepage and repository URLs point at the upstream project. Only the licence file and the readme differ.

The practical consequence is that `pip install -e .` produces a distribution named vllm. You cannot have the fork and upstream in one environment, and a wheel built from this tree is indistinguishable from upstream by name.

The build requirements are equally opinionated. The build backend pins the deep learning framework to one exact version, requires a specific minimum of the CMake build tool and the Ninja build system, and caps the packaging tool below a specific release rather than leaving it open. A comment above that list notes it should be mirrored in a requirements file in the build directory, which is a manual consistency obligation carried by every contributor.

The build script imports the framework at module level and loads the engine's own environment module by file path, with a comment explaining that a direct import is impossible because that module depends on the engine, which is not installed yet. On macOS the script warns and forces the target device to CPU.

The headline throughput win inverts on the one model with a full table

The positioning is a direct attack on a specific weakness. The argument runs that KV-cache quantization usually costs you throughput, citing a blog post on an existing compression method that reports throughput dropping between forty and fifty-two percent in exchange for capacity gains of two to three and a half times, and that aggressive low-bit quantization also tends to cost accuracy. The conclusion drawn is that losing both speed and quality is why this feature is rarely enabled in production.

KVarN's answer is that it keeps both, and the headline numbers are up to about 1.3 times the throughput of half-precision compute alongside three to five times the cache capacity. On a dense thirty-two-billion-parameter model it is said to match half-precision accuracy, beat its throughput, and deliver roughly four times the capacity.

Then the one model with a complete published table tells a different story. On a model using a compressed latent attention layout, tensor-parallel over two devices, the comparison is:

| Metric | bf16 | KVarN | ratio | | --- | --- | --- | --- | | Burst throughput at 32K (tok/s) | 401 | 377 | 0.94x | | KV-cache capacity (tokens) | 313K | 865K | 2.77x | | AIME25 accuracy | 53.3% | 53.3% | parity |

Six percent slower, in exchange for nearly three times the cache, at identical accuracy.

The README does not hide this. Its own summary says the latent is already tiny so this is not a latency play there, and that the value is fitting more concurrent context in the same memory. But it does mean the headline figure describes the dense case only, and the table most people will read first is the one where throughput regresses.

Your advertised capacity needs a memory profiler flag on a tight GPU

There is a tip in the documentation that determines whether the headline capacity number is real on your hardware, and it is easy to skip.

The mechanism is an amortization requirement. The full cache capacity is only realized when there is room to spread a small fixed decode workspace over enough of the pool. On multi-GPU setups, or on any setup where you are generous with the GPU memory utilization setting, that room exists and everything happens automatically.

On a tight single-GPU budget it often does not. The inference engine's own CUDA-graph memory profiler can over-reserve, and an over-reserved pool shrinks the KV cache that KVarN is trying to expand. The fix is an environment variable that turns off the profiler's graph estimate, optionally combined with raising the memory utilization setting.

So the failure mode is quiet. The kernel loads, the flag is accepted, the server starts, and you get less than half the capacity improvement you read about, with no error and nothing in the log. A single-GPU operator benchmarking against the published table would see the same configuration and a very different number.

The same reasoning applies to the tile size choice, which is the other knob that trades capacity for granularity.

One tile is one block, so the tile size is your page size

The quantization granularity is not set by a separate option. It is tied to the engine's own block size, because one block is one tile.

Two sizes are supported. The default, and the stated design point, is 128. A second preset at 64 is available and is described as giving finer quantization granularity.

The trade is specific rather than vague. The smaller tile costs a little KV capacity, because there is more per-tile scale overhead for each token, and it does so at essentially the same throughput. So the choice is capacity against precision of the quantization bins, with speed held roughly constant.

The presets encode the parameters in their names. The dense configuration carries a four-bit designation, a version number, and the group size, with a matching preset for each supported tile size. Switching between them is a change to one string in the dtype field for the library interface or one flag for the server.

One constraint travels with the method. KVarN runs its compute in half precision, so the model dtype is not something you can leave at whatever the model card suggests; the examples set it explicitly for both the library and the server paths.

Rejected draft tokens never reach the cache

Compatibility with speculative decoding is the section that shows the most careful thinking, because speculative decoding has an obvious way to corrupt a quantized cache.

The method composes with multi-token prediction and with draft models, configured through the engine's normal speculative configuration alongside the cache dtype flag. The design detail that matters is the commit rule.

The verify step attends over the entire cached context, which KVarN reconstructs from the quantized representation. A block is only committed to the quantized cache once every token in it has been accepted. Rejected draft tokens therefore never enter the history at all, rather than being written and then corrected.

That is the right invariant for a compressed cache. Once a lossy representation holds a rejected token, no later correction makes the stored value match what the model actually believed, so refusing to store it is cleaner than storing and hoping.

Cache quantization is also independent of weight quantization, which the documentation states explicitly. It composes with compressed checkpoints at four bits and with multi-token prediction at the same time, validated on one model in both a normal and a weight-quantized precision.

A third case extends the same idea. A parallel-drafting method is supported whose drafter attends to the cached context with bidirectional rather than causal attention, which a causal cache does not normally allow. The backend advertises non-causal support so that drafter reads the same quantized cache the target model reads, with no additional flags.

Hybrid models are deliberately left alone on the non-attention layers

A growing class of models interleaves standard attention layers with linear-attention or recurrent layers, and those do not have a KV cache at all.

KVarN's handling is to compress only the full-attention layers, which are the layers that actually hold a cache, and leave the recurrent layers untouched. The half-precision decode pool is then sized from the full-attention layer count alone, which is why a hybrid model loads with default flags and no manual pool tuning.

That is the correct behaviour and it is worth reading closely for a second reason. A cache-only method that did not know which layers held a cache would either quantize state that is not a cache or mis-size the pool. Saying explicitly that the untouched layers keep their own recurrent state is the documentation of that awareness.

There is a small inconsistency in precision across the examples. The note states that KVarN runs its compute in half precision, and the dense example passes a half-precision dtype. The hybrid example passes a brain-float dtype, and the latent-attention comparison uses brain-float as its baseline. So half precision is a property of the kernel path rather than of every invocation, and the two document sections do not agree on which flag you need.

The hybrid capacity gain is also scoped the same way as the implementation: it applies to the full-attention layers, so a model that is mostly recurrent state gains proportionally less.

The latent-attention work sits at the fork root as four loose scripts

A repository that is a fork of a large project usually keeps its own work inside the package. KVarN does not.

Four Python files sit at the top level of the repository alongside the readme and the licence: a decode reference implementation, a dequantization kernel, a tile validator, and a tile packer. Two further directories hold scripts for the dense path and the latent-attention path separately, and two documents at the root cover a backend specification and a status report for the latent-attention work.

The names suggest a development sequence. A reference implementation to define correct behaviour, a kernel to make it fast, a validator to check the kernel against the reference, and a packer to build the artefacts the kernel consumes. Having the reference beside the kernel is the right arrangement, since it is what you test against.

The rest of the top level is the upstream project's own infrastructure carried by the fork: a C++ source directory, the engine package itself, a build kit configuration, a CMake file, a large examples tree with twenty or more upstream example directories, a Rust toolchain file and a Rust build script, and the usual documentation and lint configuration.

The novel claim here is hedged. The assertion that this is the first engine-compatible sub-eight-bit KV-cache quantization method to support these latent-attention models is qualified as being to the best of the authors' knowledge, which is the correct way to make a priority claim you cannot fully verify.

Editorial conclusion

KVarN is worth evaluating if you are serving long contexts on vLLM and have measured that your KV pool is what limits you, because it is the rare quantization backend that does not ask you to trade throughput for memory. Plan for the fork rather than against it, since the package installs as vllm and cannot coexist with upstream, and measure your own capacity on the hardware you will actually run on, because the advertised figure depends on a memory profiler setting you may have to change.

Frequently asked questions

What does KVarN do?

It is a KV-cache quantization backend for vLLM that compresses the cache so more of it fits in the same memory, letting you serve longer contexts or more concurrent requests. It is claimed to keep accuracy at the level of half-precision compute and, on dense models, to exceed its throughput.

Is KVarN a plugin or a fork?

A fork. It ships as a vLLM fork and its packaging metadata declares the package name as vllm, with the upstream authors, description and URLs, so installing it produces a distribution named vllm that cannot coexist with upstream.

Does KVarN slow down inference?

On the one model with a full published table it does. On a latent-attention model the burst throughput is 377 tokens per second against 401 for brain-float compute, a ratio of 0.94x, in exchange for 2.77 times the cache capacity at identical accuracy. The readme says the latent is already tiny so that path is not a latency play.

How do I get the full KVarN cache capacity?

Make sure there is room to amortize a fixed decode workspace. On multi-GPU or generous memory-utilization setups that happens automatically, but on a tight single GPU the engine's CUDA-graph memory profiler can over-reserve and shrink the pool, so set the environment variable that disables its graph estimate and optionally raise the memory utilization setting.

What do the KVarN preset names mean?

They encode the quantization parameters: a four-bit designation, a version number, and the group size. Because one engine block is one tile, the tile size is set with the block size option, and both 128 and 64 are supported through matching presets, with 128 as the design point and 64 giving finer granularity for slightly less capacity.

Does KVarN work with speculative decoding and weight quantization?

Yes. It is compatible with multi-token prediction and draft models, and a block is committed to the quantized cache only once all of its tokens are accepted, so rejected draft tokens never enter the history. Cache quantization is independent of weight quantization, so it composes with four-bit compressed checkpoints at the same time.

Official sources

  1. huawei-csl/KVarN on GitHub
  2. Issues
  3. License: Apache-2.0
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/huawei-csl-kvarn.svg)](https://hysenlabs.com/projects/huawei-csl-kvarn)