NVIDIA/kvpress: a Python library for KV cache compression in transformers
LLM KV cache compression made easy
At a glance
- What is it?
- kvpress collects training-free KV cache compression methods behind a transformers pipeline, so you can swap presses and compression ratios without rewriting inference code. It is a research and evaluation tool first, and the decoding path is still labelled experimental.
- Who is it for?
- Adopt kvpress if you are evaluating KV cache compression methods on transformers models and want a single pipeline to swap presses and ratios. Do not adopt it as a serving-layer optimisation: the README does not document vLLM integration, and DecodingPress is described as experimental and limited to ScorerPress base presses.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 2 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The memory bill that kvpress is trying to cut
The README opens with a number: handling 1M tokens with Llama 3.1-70B in float16 requires up to 330GB of memory, because the key-value cache grows linearly with sequence length. That growth is the target. kvpress is aimed at researchers and developers working on long-context inference who want to test compression methods without rebuilding the surrounding inference stack, and the repository states that goal directly: it implements multiple KV cache compression methods and benchmarks them using transformers.
The scope is deliberately narrow. This is not a serving framework and not a quantisation library. It is a collection of presses plus the plumbing to attach them to a transformers pipeline. If your problem is that a long-context model will not fit on your GPU, kvpress is one way to attack it. If your problem is throughput under concurrent load, kvpress does not address it.
How a press hooks into the prefilling phase
Every press inherits from BasePress, and the presses that prune inherit from ScorerPress. The mechanism is uniform: during prefilling, each key-value pair gets a score, and the pairs with the lowest importance are dropped. What differs between presses is only how that score is computed. RandomPress uses a random score, KnormPress uses the inverse norm of the key, SnapKVPress uses the average attention weight of the last queries, TOVAPress uses the attention weight of the last query averaged across heads, and ObservedAttentionPress uses the average attention weight observed during prefilling. ExpectedAttentionPress estimates the expected attention weight during the generation phase instead of measuring it.
Two presses change the allocation rather than the scoring. PyramidKVPress keeps pyramid-like cache sizes, giving more budget to lower layers and less to higher layers. QFilterPress projects key representations onto the main SVD component of the query vectors to approximate attention scores. StreamingLLMPress takes a different route entirely and keeps only the initial and recent tokens.
Each press carries a compression_ratio attribute that measures the compression of the cache. The pipeline applies compression to the context tokens only, which the README says lets you evaluate compression against different questions without recompressing. That design choice is what makes the library usable for benchmarking: the compressed context is a fixed artefact you can query repeatedly.
Installing kvpress and running a first compression
The published package installs from PyPI. The README gives a single command and no configuration step:
pip install kvpressFor a local checkout the project uses uv. Cloning and syncing gives you the repository plus its development environment:
git clone https://github.com/NVIDIA/kvpress.git
cd kvpress
uv syncOptional extras exist for evaluation tooling and for FlashAttention. The README shows both in one command:
uv sync --extra eval --extra flash-attnOnce installed, importing kvpress registers a transformers pipeline under the name "kv-press-text-generation". The README's example builds that pipeline for Qwen/Qwen3-8B, wraps a long context in an ExpectedAttentionPress at a compression ratio of 0.5, and asks an optional question about the compressed context:
from transformers import pipeline
from kvpress import ExpectedAttentionPress
model = "Qwen/Qwen3-8B"
pipe = pipeline("kv-press-text-generation", model=model, device_map="auto", dtype="auto")
context = "A very long text you want to compress once and for all"
question = "\nA question about the compressed context" # optional
press = ExpectedAttentionPress(compression_ratio=0.5)
answer = pipe(context, question=question, press=press)["answer"]The pipeline handles chat templates and tokenisation, so the return value is a dictionary whose "answer" key holds the generated text. The README points to a Wikipedia notebook demo for a longer walkthrough, and a Colab link is in the badge row. Note the dependency floor from pyproject.toml: Python 3.10 or newer, torch 2.3.1 or newer, and transformers 4.56.0 or newer.
DecodingPress and the compatibility wall
Compression during prefilling is the default. A newer, explicitly experimental path compresses during generation instead, through a wrapper called DecodingPress. It takes a base_press, a compression_interval defaulting to 512 steps, a target_size defaulting to 2048 tokens, and a hidden_states_buffer_size defaulting to 256. Some presses do not need buffered hidden states and can set that last value to 0.
The important difference is the control knob. Prefilling presses use a compression_ratio; DecodingPress uses target_size, and the ratio is computed automatically so the cache lands at that size after each interval. The README's example compresses every 10 steps to a target of 512 tokens with KnormPress as the base.
Here is the constraint that matters most. The README states plainly that not all existing presses are fully compatible with DecodingPress, because compression during decoding differs fundamentally from compression during prefilling, and that only ScorerPresses are supported as base presses. So the presses that make kvpress interesting for layer-aware allocation, such as PyramidKVPress, are not available on the decoding path. If your workload needs compression during generation, check the base press against the ScorerPress hierarchy before planning around it.
There is a second, quieter limitation. The README does not document rollback. If a compression ratio turns out to be too aggressive for a task, the documented behaviour offers no way to recover the dropped key-value pairs; the context is compressed once and the discarded pairs are gone.
What kvpress is not, and where a serving stack fits
kvpress attaches to transformers. The repository lists a kvzap/ directory and the related searches include KVzap, but the README does not describe what kvzap does, so treat it as an unverified part of the tree rather than a documented feature.
For teams already running inference through a dedicated serving engine, the difference in approach is structural. A serving engine owns the cache, the scheduler and the attention kernel, and applies whatever memory policy it supports inside that stack. kvpress instead sits on top of a Hugging Face pipeline, wraps the model's attention, and scores key-value pairs in Python. That is why the library can offer ten interchangeable scoring methods with no training step and no kernel work: it is optimising for comparability between methods, not for tokens per second. The trade-off is that anything kvpress does, it does at the level of a single pipeline call, and the README does not describe an integration path into a serving engine.
If you want an off-the-shelf serving configuration, kvpress is the wrong tool. If you want to know whether ExpectedAttentionPress at ratio 0.5 beats KnormPress at ratio 0.5 on your own long-context task, kvpress is built for exactly that question.
Maintenance, licence and the cost of tracking transformers
The last push to the default branch was on 2026-09-07, and the most recent release is v0.5.4 from 2026-07-02, following v0.5.3 in April 2026 and v0.5.2 earlier that month. The repository is not archived. Release cadence is irregular rather than scheduled, so plan upgrades around releases you choose, not a fixed train.
The upgrade cost is dominated by one dependency. pyproject.toml pins transformers to >=4.56.0,<5.3, torch to >=2.3.1,<3, and accelerate to >=1.0.0,<2. Because presses hook into attention internals, a transformers release that changes those internals is the realistic source of breakage. The project's own Makefile runs flake8, mypy with --check-untyped-defs, and a check that every Python file carries an SPDX-FileCopyrightText header, which tells you the codebase is held to a typed, linted standard; passing your own test suite against a new transformers version is still the only way to confirm a press behaves.
kvpress is Apache-2.0. The presses implement methods from published papers, several of which are linked in the README, and the repository ships a CITATION.cff. Apache-2.0 permits commercial use, but the licence covers the code, not any patent or licensing position attached to the underlying methods. Check the papers and your own counsel if that distinction matters to you.
Editorial conclusion
Adopt kvpress if you are evaluating KV cache compression methods on transformers models and want a single pipeline to swap presses and ratios. Do not adopt it as a serving-layer optimisation: the README does not document vLLM integration, and DecodingPress is described as experimental and limited to ScorerPress base presses. Before committing, verify that your chosen press appears in the available presses list, that your transformers version falls inside the declared range, and that your attention implementation matches what the press expects.
Frequently asked questions
What is the difference between the KV cache and the context window?
The context window is the number of tokens a model can attend to; the KV cache is the stored key and value tensors for those tokens. kvpress targets the cache, which the README says grows linearly with sequence length, and compressing it reduces the memory the model needs during inference.
How to compress KV cache with kvpress?
Import kvpress so the "kv-press-text-generation" pipeline is registered, build the pipeline for your model, then pass a press with a compression_ratio to the call. The README's example uses ExpectedAttentionPress at a ratio of 0.5 and reads the result from the "answer" key.
Does kvpress require training or fine-tuning?
No. The README states that all current presses are training free, and each inherits from BasePress. Compression happens at inference time by scoring key-value pairs and pruning the lowest-importance ones.
Can kvpress compress the cache during generation instead of prefilling?
Yes, through the DecodingPress wrapper, which the README describes as a new experimental feature. It compresses every compression_interval steps down to a target_size, and only ScorerPress instances are supported as base presses.
Which Python and transformers versions does kvpress need?
pyproject.toml requires Python 3.10 or newer, transformers >=4.56.0 and <5.3, and torch >=2.3.1 and <3. The published version at the time of writing is 0.5.4.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/nvidia-kvpress)