Model or dataset
Zefan-Cai/KVCache-Factory avatar
Zefan-Cai/KVCache-Factory

KVCache-Factory: A Unified Benchmark for KV Cache Compression in LLM Inference

Unified KV Cache Compression Methods for Auto-Regressive Models

1,384 stars180 forksPythonMIT

At a glance

What is it?
KVCache-Factory is a Python research toolkit for evaluating and comparing KV cache compression, retrieval, merging, and quantization methods for long-context LLM inference. It started from the PyramidKV paper and now provides a single evaluation interface for more than a dozen KV cache strategies on Llama and Mistral architectures.
Who is it for?
KVCache-Factory is suited to researchers and ML engineers who want to evaluate and compare KV cache compression methods under controlled conditions without reimplementing each baseline. It is not ready for production deployment: it targets Llama and Mistral architectures, quantization support requires reading the specific method's constraints before running, and the repository notes that some newer methods have narrower model coverage.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 47 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

The Problem KVCache-Factory Addresses

Long-context LLM inference is bottlenecked by the size of the key-value cache that grows with sequence length. Different research papers have proposed different strategies for managing this: evicting low-importance tokens, retrieving only the tokens most relevant to a query, merging adjacent-layer representations, or quantizing cache values to reduce memory per token. Each proposal typically ships with its own implementation and evaluation setup, making it difficult to compare methods on equal footing.

KVCache-Factory provides a common evaluation interface for these methods. It started from the PyramidKV paper, which introduced a layer-wise pyramidal cache budget allocation, and has expanded to include multiple baselines. The intended audience is LLM inference researchers and engineers who need to reproduce paper results, compare methods on standard benchmarks, or profile a compression strategy before integrating it into their own inference stack.

Supported Methods: From Eviction to Quantization

The repository supports 18 methods organized into five categories. The eviction and retrieval group covers StreamingLLM (attention-sink plus sliding window), H2O (heavy-hitter token retention), SnapKV (observation-window attention pooling), NACL (proxy-token score reduction), Scissorhands (historical importance accumulation), and L2Norm (norm-based token selection). The query-aware retrieval group includes Quest, which builds page-level key min/max metadata and selects tokens at query time.

Cross-layer and adaptive strategies include MiniCache (SLERP direction sharing across adjacent layers), PyramidKV (layer-wise pyramidal budget), CAM (cache merging with attention-informed value aggregation), AdaKV (head-adaptive budgets), and HeadKV (head-aware retrieval and reasoning cache allocation). ThinK provides query-driven key-channel pruning for Llama LongBench runs.

Three quantization backends are available: KIVI, KVQuant, and GEAR. GEAR additionally accepts `--rank` and `--outlier_ratio` arguments. All three are activated with `--quant_method kivi`, `--quant_method kvquant`, or `--quant_method gear`. The quantization path is separate from the compression methods; both can be combined.

HeadInfer is the one lossless method. Rather than approximating the cache, it offloads KV cache heads to CPU memory with async prefetch. It requires `flash_attention_2` and ignores `--max_capacity_prompts` because it retains the full cache.

Installing and Running the LongBench Evaluation

KVCache-Factory requires `transformers==4.44.2`, `torch`, and optionally `flash-attn>=2.4.0.post1`. The README notes that `flash-attn` must be installed manually after `torch` because it needs PyTorch at build time.

bash
git clone https://github.com/Zefan-Cai/KVCache-Factory.git
cd KVCache-Factory
pip install -r requirements.txt
export PYTHONPATH="$PWD:${PYTHONPATH}"

For `flash-attn` specifically:

bash
pip install flash-attn --no-build-isolation

Beyond LongBench, the repository provides two additional evaluation runners. The needle-in-a-haystack runner (`run_needle_in_haystack.py`) tests retrieval at varying sequence lengths between `--s_len` and `--e_len`. The RULER runner (`run_ruler.py`) accepts the same method flags. Both accept the same `--attn_implementation` and `--method` arguments as the LongBench runner.

Attention visualization is available through `examples/visualization.ipynb` and the utilities in `pyramidkv/viztools/`. Generated maps are stored under `./attention` by default.

The LongBench evaluation runner takes a method name, a model path, and configuration flags.

bash
python3 run_longbench.py \
  --method pyramidkv \
  --model_path /path/to/Llama-3-8B-Instruct \
  --max_capacity_prompts 128 \
  --attn_implementation flash_attention_2 \
  --save_dir ./results_long_bench \
  --use_cache True

The `--method` flag accepts one of `FullKV`, `pyramidkv`, `snapkv`, `streamingllm`, `h2o`, `cam`, `l2norm`, `adakv`, `headkv`, `think`, `headinfer`, or `minference`. The `--datasets` flag accepts a comma-separated list of LongBench datasets and defaults to the full 16-dataset list. The `--dtype` flag defaults to `float16`.

For GPUs without FlashAttention v2 support, `--attn_implementation sdpa` is available. The `--method think` specifically requires `eager`.

HeadInfer and the Distinction Between Compression and Offloading

Most methods in KVCache-Factory reduce accuracy in exchange for memory savings. HeadInfer takes a different approach: it offloads KV cache entries head-by-head to CPU memory with asynchronous prefetch, keeping the full cache intact. The README classifies this as lossless.

The trade-off is bandwidth. Asynchronous prefetch hides some of the CPU-GPU transfer latency, but the total bandwidth required depends on the model's number of attention heads and the sequence length. HeadInfer is most useful when GPU memory is the primary constraint and CPU memory is abundant. It is the only method in the repository where `--max_capacity_prompts` has no effect because the full KV cache is preserved.

This design distinction matters for evaluation: comparing HeadInfer with an eviction-based method like H2O on LongBench measures different things. HeadInfer's accuracy matches FullKV by design; the variable is memory usage and throughput, not quality degradation.

Model Coverage and Constraints to Verify Before Running

The README explicitly states that Llama and Mistral attention paths are supported for the main compression methods, and that some newer methods have narrower runner and model coverage. It advises checking the runner argument choices before launching large experiments. This is practical guidance: not every method works on every model, and running a method on an unsupported architecture may silently produce incorrect results rather than raising an error.

For grouped query attention (GQA) models, the `--kv_cache_granularity` flag selects the cache layout: `query_head` (the default, legacy layout) or `kv_head` (a GQA-efficient layout supported for `snapkv`, `pyramidkv`, `h2o`, `streamingllm`, `cam`, `l2norm`, `adakv`, and `headkv`). The README notes GPU validation for `adakv` and `headkv` with `kv_head` is pending.

The MInference integration is optional and kept out of the base `requirements.txt`. Installing it requires `pip install -r requirements-minference.txt` separately.

For quantization, KIVI uses key axis `1` and value axis `0` by default. The `--q_group_size`, `--axis_key`, and `--axis_value` flags control the quantization layout for advanced use. The `--quant_residual_length` flag sets the full-precision residual cache window and defaults to `max_new_tokens`. These options give fine-grained control over how quantization interacts with different model architectures, but each combination should be validated on a small experiment before committing to a full LongBench run.

Alternatives for Production KV Cache Management

KVCache-Factory is a research evaluation harness, not a production inference library. For production use, NVIDIA's TensorRT-LLM and vLLM both implement KV cache management strategies as part of their serving infrastructure. vLLM, for example, uses paged attention for memory-efficient KV cache allocation and is widely deployed in production LLM serving. These tools prioritize throughput and deployment stability over the ability to swap in arbitrary compression baselines.

KVCache-Factory's distinctive value is the unified evaluation interface that lets a researcher run PyramidKV, SnapKV, and H2O on the same model with the same benchmark and the same evaluation script. That comparison is difficult to replicate with production serving infrastructure, which does not expose the level of control over individual compression strategies that KVCache-Factory provides.

For teams that need to justify a specific KV cache strategy to stakeholders, the ability to run FullKV as a baseline alongside the compressed variant on LongBench or the needle-in-a-haystack test provides a reproducible reference point. The repository's `results_long_bench` directory accumulates output across runs for this purpose.

Maintenance State and MIT License

The last push to the repository was on 2026-08-13. The repository has no GitHub releases; version history is tracked through commits. Notable additions include multi-GPU inference support for large models including Llama-3-70B-Instruct (added in June 2024) and FlashAttention v2 and SDPA paths for several methods (also June 2024).

The project is licensed under MIT. No contributor agreement is required. A Chinese-language README is available at `README_zh.md` for readers who prefer that.

Editorial conclusion

KVCache-Factory is suited to researchers and ML engineers who want to evaluate and compare KV cache compression methods under controlled conditions without reimplementing each baseline. It is not ready for production deployment: it targets Llama and Mistral architectures, quantization support requires reading the specific method's constraints before running, and the repository notes that some newer methods have narrower model coverage. Before running large experiments, verify that the chosen method supports the target model by checking the runner argument choices as the README advises.

Frequently asked questions

What is KV cache compression in LLMs?

In LLM inference, the key-value cache stores intermediate attention computations for each generated token. As sequence length grows, this cache grows proportionally and becomes a memory bottleneck. KV cache compression methods reduce this memory by selectively evicting, merging, or quantizing cached values, accepting some accuracy trade-off in exchange for lower memory use and faster inference.

Which models does KVCache-Factory support?

The README states that Llama and Mistral attention paths are supported for the main compression methods. Some newer methods have narrower model coverage. The repository advises checking the runner argument choices before running large jobs to confirm compatibility between a chosen method and a target model.

How does KVCache-Factory differ from a production LLM serving framework?

KVCache-Factory is a research evaluation harness for comparing KV cache compression baselines on benchmarks like LongBench. It is not a serving system. Production inference frameworks such as vLLM manage KV cache allocation for throughput and deployment stability, but do not expose the per-method configurability that KVCache-Factory provides for controlled research comparisons.

Official sources

  1. Issues
  2. License: MIT
  3. README
  4. Zefan-Cai/KVCache-Factory on GitHub
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/zefan-cai-kvcache-factory.svg)](https://hysenlabs.com/projects/zefan-cai-kvcache-factory)
Community notes

Community notes