What is KV cache?
A KV cache (key-value cache) stores the key and value tensors a transformer has already computed for earlier tokens, so each new token is generated without recomputing the whole prefix. It trades GPU memory for speed, and it is the reason long conversations and long documents get expensive to serve.
How the KV cache works
A decoder-only transformer generates text one token at a time. Inside every attention layer, each token is projected into a query, a key and a value. To produce token number N, the layer needs the keys and values of all previous tokens, not just the current one. Without a cache, the model would recompute those tensors from scratch at every step, so generating a sequence of length L would cost work proportional to L squared.
The KV cache removes that repetition. After a token is processed, its key and value tensors are kept in memory. When the next token arrives, the model computes only the new query, key and value, then attends over the stored keys and values plus the new ones. The result is that per-step cost becomes roughly linear in the sequence length, and the prefix is encoded once instead of L times.
The cost is memory. For each layer, each attention head and each token, two tensors must be held. Their size depends on the head dimension and the numeric precision used. A typical estimate is 2 (key and value) times the number of layers times the number of KV heads times the head dimension times the number of tokens times the bytes per element. On a long context, this can exceed the memory used by the model weights themselves, which is why serving systems spend so much effort on cache layout, paging and compression.
Two further details matter in practice. First, the cache is append-only during a single generation: entries are added, not rewritten. Second, it is tied to a specific prompt prefix. If the beginning of a request changes, the stored entries no longer apply, which is the reason prefix caching only helps when requests genuinely share a prefix.
When you need a KV cache and when you do not
You need a KV cache whenever a transformer model generates tokens autoregressively and you care about latency or throughput. Interactive chat, code completion, streaming summarisation and any multi-turn session all fall into this group. The same applies to batched serving: without a cache, every request would re-encode its whole prompt on every decoding step, and batching would multiply that waste.
You do not need one when the model is used in a single forward pass with no generation, such as embedding extraction or classification, because there is no future token that needs past keys and values. You also do not need one for architectures that do not use attention over a growing context. BlinkDL/RWKV-LM describes RWKV as an RNN that can be trained like a GPT transformer while offering linear time and constant space, explicitly stating there is no kv-cache. In that design, state replaces the growing cache, so memory does not scale with context length.
There is also a middle ground. If your prompts are short and generation is brief, the cache saves little and the memory bookkeeping may not be worth it. The decision usually comes down to sequence length and concurrency: the longer the context and the more simultaneous requests, the more the cache matters, and the more its memory footprint becomes the binding constraint.
Limits and common pitfalls
The first limit is memory fragmentation. Naive implementations allocate a contiguous block sized for the maximum possible sequence length, which wastes space for short requests and blocks admission for long ones. Paged or virtualised schemes address this by separating virtual addressing from physical GPU memory. ovg-project/kvcached describes itself as an Apache-2.0 Python library that decouples KV cache virtual addressing from physical GPU memory so SGLang and vLLM can share a GPU elastically. The trade-off is that you now depend on an extra component in the serving path.
The second limit is persistence. A cache held in GPU memory disappears when the engine restarts, so a warm prefix must be recomputed. LMCache/LMCache moves KV cache out of GPU memory into a tiered store that survives engine restarts and can be shared between engines. That helps long-context, multi-turn or RAG traffic, but the operational cost is real: you run a cache tier alongside your model server.
The third limit is precision and compression. Keys and values can be quantised to save memory, but the accuracy loss is not uniform across layers or heads. 0xSero/turboquant implements a KV cache quantisation scheme as Triton kernels plus a vLLM patch, and its README retracts several headline numbers and narrows the real gain to the full-attention layers. That is a useful reminder that compression claims should be read against the layers and budgets they were measured on.
A fourth pitfall is prompt mutation. Any change to the prefix invalidates the cached entries for that region. Systems that insert retrieved documents at the front of a prompt, or rewrite system messages between turns, will see low cache hit rates. microsoft/LLMLingua takes a different route: it compresses the prompt and KV-Cache before the request reaches a black-box LLM, trading a small local model pass for fewer billed tokens. The README states up to 20x compression with minimal performance loss, but this is a poor fit when every character of the input must survive intact.
Finally, cache reuse is not free accuracy. shepherd-agents/shepherd reports about 95% KV-cache reuse on replay while turning an agent run into a reversible, Git-like trace. Reuse depends on the replayed prefix matching exactly, so any nondeterminism in the environment or prompt assembly reduces the benefit.
How open-source projects expose the KV cache
Some projects make the cache a first-class object. LMCache/LMCache is built around a KV cache layer that can be shared across engines and survive restarts. ovg-project/kvcached virtualises the cache so multiple engines can share one GPU. uccl-project/uccl is a C++ GPU communication library covering collectives, P2P transfers such as KV cache transfer and RL weight transfer, and expert parallelism; it keeps the NCCL API while replacing the transport, and ships as a Python wheel that loads as a net plugin. These are infrastructure pieces: they assume you already run a transformer server and want the cache to move or persist.
A second group treats the cache as a compression target. NVIDIA/kvpress collects training-free KV cache compression methods behind a transformers pipeline, so presses and compression ratios can be swapped without rewriting inference code; the decoding path is still labelled experimental. Zefan-Cai/KVCache-Factory gathers PyramidKV, SnapKV, H2O, StreamingLLM, HeadInfer, KIVI and other methods behind a single evaluation runner, and its README describes it as a research harness rather than a library. Zefan-Cai/R-KV is a decoding-time compression method for reasoning models, shipped as patches over vLLM, SGLang, Nano-vLLM and Mini-SGLang plus a HuggingFace path; it evicts redundant tokens during generation and reports lossless accuracy at budget=512 on GSM8K. 0xSero/turboquant sits here too, with Triton kernels and a vLLM integration for 3-bit keys and 2-bit values.
A third group avoids the cache entirely or reuses it at the agent layer. BlinkDL/RWKV-LM states that RWKV-7 has constant space and no kv-cache, which changes the memory profile of long contexts. shepherd-agents/shepherd records an agent run as a durable trace with retained outputs, so a supervising meta-agent can inspect, fork, replay or revert it; the alpha is Python 3.11+, MIT licensed, and macOS or Linux only. Its reported KV-cache reuse on replay is a consequence of replaying the same prefix, not a general property of the runtime.
None of these projects is a drop-in replacement for another. The compression tools change what you store, the infrastructure tools change where it lives, and RWKV changes whether it exists at all.
Sizing and operational questions to ask
Before adopting any cache-related tool, estimate the cache size for your own model and context length using the formula above, then compare it with the free memory left after weights and activations. If the cache dominates, compression or offloading is worth evaluating; if it does not, the added operational surface may not pay for itself.
Next, measure the cache hit rate on real traffic. Prefix sharing is the assumption behind most reuse schemes, and it is easy to break with dynamic system prompts or retrieved context inserted at the front. A cache that never hits is pure overhead.
Finally, check the maintenance signals. BlinkDL/RWKV-LM, LMCache/LMCache, microsoft/LLMLingua, shepherd-agents/shepherd, 0xSero/turboquant, uccl-project/uccl, ovg-project/kvcached, Zefan-Cai/KVCache-Factory, NVIDIA/kvpress and Zefan-Cai/R-KV all show recent activity, but the archived flag and last push date are the only reliable indicators, and you should read them at the time you evaluate a dependency. Research harnesses in particular may not track upstream inference engines closely.
In practice
The KV cache is a memory-for-compute trade: it makes autoregressive generation linear per step but makes long contexts memory-bound. Start by measuring cache size and hit rate on your own workload, then decide whether you need a persistence layer such as LMCache, a virtualisation layer such as kvcached, a compression method such as kvpress or R-KV, or an architecture such as RWKV that removes the cache entirely.