R-KV: Redundancy-Aware KV Cache Compression for Reasoning Models
[Neurips 2025] R-KV: Redundancy-aware KV Cache Compression for Reasoning Models
At a glance
- What is it?
- R-KV is a decoding-time KV cache compression method for reasoning models, shipped as patches over vLLM, SGLang, Nano-vLLM and Mini-SGLang plus a HuggingFace path. It evicts redundant tokens during generation and reports lossless accuracy at budget=512 on GSM8K.
- Who is it for?
- R-KV is worth trying if you already serve a reasoning model on vLLM v0.25.1 or SGLang v0.5.14 and your KV pool, not your compute, is the limit on concurrent requests.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Yes. The repository last received commits 71 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
The problem R-KV targets: reasoning traces that outgrow the KV pool
Long chain-of-thought decoding keeps every past token's key and value in memory for the whole generation. On a single H100 serving a 7B or 8B reasoning model, that pool, not the arithmetic units, decides how many requests can be in flight. The repository frames the method as discarding repetitive tokens on the fly, "delivering full-accuracy reasoning with only a fraction of the memory". The intended reader is someone running R1-distill style models on math benchmarks and hitting a memory ceiling, not someone tuning prefill latency. The project carries the NeurIPS 2025 label and topics kvcache, llm and reasoning-models, which matches that audience.
How R-KV decides what to drop: budget, buffer and cross-layer scoring
Compression is a decode-time operation, not a prefill pass. In the vLLM port it activates automatically once budget and buffer are both greater than zero, read from the environment variables VLLM_V1_R_KV_BUDGET and VLLM_V1_R_KV_BUFFER. The HuggingFace path exposes the same idea through flags: --method rkv, --kv_budget, --window_size, --mix_lambda, --divide_method step_length and --divide_length. The window_size and divide_length parameters point at a local window plus a segmentation scheme, and mix_lambda blends two scoring signals. The release notes describe batched cross-layer redundancy scoring and bounded-footprint block freeing (FREE_BLOCKS), which returns evicted KV to the allocator rather than leaving it stranded. In the tensor-parallel case every rank is said to evict the identical set, which is what makes the scheme usable above a single GPU. The SGLang port adds a fused Triton redundancy kernel, two-phase compaction and compression-aware admission, and the patch applies an eviction-score all-reduce across ranks for TP.
Installing R-KV and running a first math evaluation
The HuggingFace route is the shortest path to a first number. Install the minimal dependencies, then build the rkv package itself, because the run scripts import it.
pip install -r requirements.txt
cd HuggingFace
pip install -r requirements.txt
pip install -e .The repository notes that the evaluation toolkit has its own dependencies, including a local build of latex2sympy under HuggingFace/evaluation/latex2sympy. Without that step the math answer checking will not import. The README also gives a one-line entry point, bash scripts/run.sh, and the explicit form below, which runs DeepSeek-R1-Distill-Llama-8B over the AIME24 file with a KV budget of 1024.
export CUDA_VISIBLE_DEVICES=0
cd HuggingFace
python3 ./run_math.py \
--dataset_path ./data/aime24.jsonl \
--save_path ./outputs/output.jsonl \
--model_path deepseek-ai/DeepSeek-R1-Distill-Llama-8B \
--max_length 32768 \
--eval_batch_size 1 \
--method rkv \
--kv_budget 1024 \
--window_size 8 \
--mix_lambda 0.1 \
--divide_method step_length \
--divide_length 128Expect a JSONL file at the save path with one record per problem. The README states that HuggingFace defaults to flash attention via attn_implementation="flash_attention_2", so a machine without FlashAttention will need that changed. For serving, the project ships patches rather than forks. The vLLM port pins v0.25.1 and the script clones the upstream tree, drops in rkv/ and applies the wiring patch; the source build needs CUDA and a GPU.
cd vLLM
scripts/apply_rkv.sh
pip install -e vllm-srcThe SGLang port follows the same shape against v0.5.14, with apply_rkv.sh, a pip install of sglang-src/python from the SGLang wheel index, and requirements-rkv.txt. The README warns not to pass --enforce-eager to the vLLM port, because R-KV auto-selects PIECEWISE cudagraph. The Nano-vLLM reference engine is the opposite case: it is roughly 1.2k lines and requires enforce_eager=True.
The stale-generation bug, and why you should rerun old GSM8K numbers
One entry in the news list is a correction rather than a feature. Commits before e9f54c45 shipped a data/gsm8k.jsonl whose generation field came from an old experiment, and eval_math.py preferred that field over the fresh run's output. The effect was that GSM8K scores froze around 40 percent regardless of method or budget, which is exactly the kind of number a reader would have used to judge whether compression hurt accuracy. MATH and AIME24 were unaffected. Anyone who computed GSM8K numbers with the repository before that commit should rerun them. The incident says something about the project's maturity: the evaluation harness had a silent precedence rule that masked the variable under test, and it survived long enough to reach users.
Where R-KV is the wrong choice
The reported wins are memory-bound throughput wins, and the repository says so. The vLLM port is described as holding far more requests in flight than Full-KV as the pool shrinks, with reported gains of +17 percent to +32 percent tokens per second at gpu_mem 0.40 to 0.25. If your workload is compute-bound, or your sequences are short enough that the KV pool never fills, there is nothing to win and you have added a patch to maintain. The SGLang port's throughput is described as sitting within a few percent of the fair Full-KV baseline, so that port is about staying lossless while compressing, not about going faster. The accuracy claims are scoped to GSM8K at budget=512 and Qwen2.5-Math-7B on H100; the README does not claim lossless behaviour at other budgets, other datasets or other model families. The TODO list still names GPT-OSS, VeRL, QwQ, GPQA and liveCodeBench as unfinished, and it asks for expanded Qwen-3 evaluation coverage. Treat any budget below the validated setting as untested. There is also no rollback procedure documented for the serving patches; the apply scripts clone and patch a pinned checkout, and the README does not describe how to revert a running deployment.
How R-KV differs from static KV cache compression
PyramidKV, which appears in the related searches and is a known point of comparison for this class of work, allocates a fixed, pyramid-shaped budget across layers ahead of time. R-KV is a decoding-time method: it scores redundancy as tokens arrive and evicts on the fly, with the budget and buffer set at launch. The practical difference is that a static allocation cannot react to a particular reasoning trace, while an online method can, at the cost of scoring work on the critical path. That cost is visible in the ports: the SGLang integration needed a fused Triton kernel and batched cross-layer scoring to keep throughput near baseline, and the vLLM port needed in-graph decode and async scheduling. R-KV also ships two lightweight reference engines, Nano-vLLM and Mini-SGLang, which a static scheme would not need in order to be understood. If you want a method you can read in an afternoon, the reference ports are the honest comparison; if you want one that adapts per trace, the scoring overhead is the price.
Maintenance, licence and upgrade cost
The last push was on 2026-07-20, and the repository is not archived. There are no releases, so adoption means tracking main. The layout is deliberately patch-not-fork, which keeps the diff small but ties you to pinned upstream versions: vLLM v0.25.1 and SGLang v0.5.14. Upgrading either server means revalidating the patch against a new checkout, and the repository gives no compatibility matrix beyond those pins. The licence is not stated in the README and there is no LICENSE file in the top-level listing, so you cannot tell from the repository what terms apply. That is a question for whoever owns your distribution, and it should be answered before R-KV goes anywhere near a shipped product. The four serving integrations plus the HuggingFace path are described as GPU-validated on A100 as of the 2026-07-02 run in results/validation-2026-07-02-a100/, which is the closest thing to a regression suite the repository offers.
Editorial conclusion
R-KV is worth trying if you already serve a reasoning model on vLLM v0.25.1 or SGLang v0.5.14 and your KV pool, not your compute, is the limit on concurrent requests. It is the wrong tool if you need a stable tagged release, a documented rollback path, or a licence you can read before shipping: the repository has no releases and no licence file, so verify both, plus the storage and memory cost of the full pinned upstream checkouts the apply scripts clone, before you put it in front of traffic.
Frequently asked questions
What is R-KV?
R-KV is a decoding-time KV cache compression method for reasoning models, presented as a NeurIPS 2025 paper and released as code under Zefan-Cai/R-KV. It discards repetitive tokens on the fly rather than allocating a fixed budget ahead of time, and it ships as patches over vLLM v0.25.1 and SGLang v0.5.14 plus a HuggingFace path and two lightweight reference ports.
How does KV cache compression work in R-KV?
Compression runs during decode, not prefill. In the vLLM port it activates once budget and buffer are both greater than zero, read from VLLM_V1_R_KV_BUDGET and VLLM_V1_R_KV_BUFFER, and the release notes describe batched cross-layer redundancy scoring with bounded-footprint block freeing that returns evicted KV to the allocator.
Which serving frameworks does R-KV support?
The repository ships ports for vLLM v0.25.1, SGLang v0.5.14, Nano-vLLM and Mini-SGLang, plus a HuggingFace implementation. The vLLM and SGLang ports are patch-not-fork layouts applied by scripts/apply_rkv.sh against pinned upstream checkouts.
Does R-KV lose accuracy compared with a full KV cache?
The release notes report lossless results at budget=512 on GSM8K for both the vLLM and SGLang ports, and the SGLang port is described as on par with the baseline at 95% GSM8K on Qwen2.5-Math-7B. The README does not claim lossless behaviour at other budgets or on other datasets.
Why did R-KV GSM8K scores come out around 40 percent?
A stale generation field in data/gsm8k.jsonl was preferred by eval_math.py over the fresh run's output, freezing GSM8K scores near 40 percent regardless of method or budget. The repository asks anyone who computed GSM8K numbers before commit e9f54c45 to rerun them; MATH and AIME24 were unaffected.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/zefan-cai-r-kv)
Community notes