kvcached: elastic KV cache allocation for shared GPUs
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
At a glance
- What is it?
- kvcached is an Apache-2.0 Python library that decouples KV cache virtual addressing from physical GPU memory, so SGLang and vLLM can share a GPU elastically. Here is how its mechanism works, how to install it, and where it stops being the right tool.
- Who is it for?
- Adopt kvcached when you already run SGLang or vLLM on a GPU you must divide between several models, and when you can match the engine version to the tested range in the README table. Do not adopt it for a single model that already fills the device, or for an engine outside that table.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
The memory partitioning problem kvcached attacks
Serving engines reserve KV cache memory up front. Once a vLLM or SGLang process claims its share of the device, that share is gone for the lifetime of the process, even when the model behind it is idle. The practical result on a single GPU is rigid partitioning: two models each get a fixed slice, and whichever one is quiet wastes its slice while the busy one queues.
kvcached targets exactly that. Its README frames the goal as making GPU sharing flexible and easy, and the repository description calls it a virtualized elastic KV cache for dynamic GPU sharing. The intended users are people running multi-model serving, serverless LLM deployments, and compound AI systems on hardware they cannot expand. The README lists those three as the example use cases, and notes that Red Hat's Sardeenz project builds on kvcached for dynamic multi-model serving with Kubernetes and OpenShift.
It is not a general GPU memory manager and it is not a scheduler. It is a KV cache layer that sits under a serving engine you already run.
How the virtual memory abstraction is wired into the engine
The core mechanism is a split between GPU virtual addressing and physical memory allocation. According to the README, serving engines initially reserve virtual memory only, and physical GPU memory is attached later, when the cache is actively used. Allocation and reclamation then follow live load rather than a startup estimate.
The repository layout shows where the pieces live. csrc/ holds the C++ and CUDA sources that setup.py compiles into an extension, kvcached/ holds the Python package, controller/ and engine_integration/ hold the glue that connects the library to a serving engine, and tools/ plus the kvctl and kvtop console scripts expose control and observation. The pyproject.toml declares three runtime dependencies: numpy, posix_ipc and wrapt. The posix_ipc dependency is consistent with a design where separate processes coordinate over shared memory rather than through a single in-process allocator.
Two features build on top of that base. A frontend router with a sleep mode routes requests to the target model and puts models to sleep when idle, which is what makes the elastic behavior visible to a caller. Prefix caching arrived in the 2026-04 update: automatic prefix caching for vLLM and RadixCache for SGLang, both with a configurable memory bound. The examples directory mirrors the same scope, with directories for a model router with sleep, memory control, serverless serving, and prefix caching.
The honest trade-off here is that elasticity is only as good as the reclamation path. Nothing in the README describes what happens to in-flight requests when memory is reclaimed, and there is no documented rollback procedure. That is a gap worth probing in a staging cluster before it matters in production.
Installing kvcached and running a first shared-GPU setup
The build is not a pure-Python install. setup.py imports torch and raises ImportError with the message "Torch not found, please install torch>=2.6.0 first." if the import fails, so install PyTorch before anything else. The same file then inspects torch.version.hip and torch.version.cuda to pick a backend, and raises a RuntimeError if neither is set. A CPU-only PyTorch build will not produce a working extension.
With a CUDA or ROCm PyTorch in place, install the package from the repository root:
pip install torch>=2.6.0
pip install -e .The editable install compiles the sources under csrc/ and registers the two console scripts declared in pyproject.toml. The README describes the CLI as enforcing memory limits, and the script entry points are kvctl and kvtop. After installation, kvtop should be on your PATH; it is the observation tool for the cache state, while kvctl is the control path for limits.
The repository ships numbered examples rather than a single getting-started script. The relevant entry points for a first run are examples/01_simple_two_models/ for two models on one GPU, examples/02_memory_control/ for the kvctl limit path, and examples/03_model_router_sleep/ for the router with sleep mode. Each directory is self-contained; read the files there rather than expecting a top-level quickstart, because the README does not provide one.
Engine support is version-bounded. The README table lists SGLang at v0.4.9 or newer, tested up to v0.5.15, and vLLM at v0.8.4 or newer, tested up to v0.24.0. Attention types covered are MHA, GQA, MLA, sliding window and hybrid. If your engine is older than the floor in that table, the integration is not the supported configuration.
Where kvcached is the wrong tool
The clearest failure case is a single model that already saturates the device. If one process can use all the memory it needs, there is nothing to share and the virtualization layer adds indirection without returning anything. The README's own framing is multi-model serving, serverless, and compound systems, and none of those describe a saturated single-tenant GPU.
The second boundary is engine coupling. kvcached integrates with SGLang and vLLM, and the README gives a tested version range for each. Running it under a different engine, or under a fork whose KV cache layout has diverged, is outside what the documentation claims. The README points to issue #425 for per-model results on each engine and KV layout, which implies the model-by-engine matrix is not uniform: some combinations are validated and others are not, and the README does not present a single blanket guarantee.
The third boundary is maturity. pyproject.toml carries the classifier "Development Status :: 3 - Alpha". The most recent release listed is v0.1.5 from 2026-04-07, with v0.1.4 and v0.1.3 before it. For a component that sits between a serving engine and physical GPU memory, an alpha classifier is a statement about API and behavior stability, and it should be read as one.
Finally, the documentation has holes that matter operationally. The README describes elastic allocation and reclamation but does not document rollback, and it does not describe behavior under memory pressure when reclamation cannot keep up. Those are the questions to answer yourself before trusting it with a latency-sensitive endpoint.
How this differs from LMCache and from static partitioning
LMCache is the natural comparison point, and the two solve adjacent problems differently. LMCache is built around reusing and moving KV cache across storage tiers, so the emphasis is on where cached state lives and how to get it back cheaply. kvcached instead applies an OS-style virtual memory abstraction to the GPU itself: the README describes decoupling GPU virtual addressing from physical memory allocation, reserving virtual memory first and backing it with physical memory when the cache is actively used. One is about persistence and reuse across a hierarchy; the other is about the timing of physical allocation inside a shared device.
The second alternative is the status quo: static partitioning, where each model is given a fixed memory slice at launch. That approach is predictable and has no extra dependency, and for a stable two-model deployment with known traffic ratios it is often good enough. kvcached's argument is that dynamic and mixed workloads make fixed slices wasteful, because the quiet model holds memory the busy one needs. Whether that argument holds for your traffic is an empirical question about variance, not a general truth.
A third reference point is the project's own research context. The README links two arXiv papers, one on a GPU OS vision and one on multi-LLM serving, plus a blog post on GPU cost. Those describe the design intent; they are not evidence about your workload.
Licence, maintenance and upgrade cost
kvcached is Apache-2.0. The repository carries a LICENSE file, a .license-header.txt used by the pre-commit configuration, and SPDX headers in the sources, for example the line "SPDX-License-Identifier: Apache-2.0" at the top of setup.py. Apache-2.0 includes an explicit patent grant and requires preservation of notices, which is friendlier for commercial embedding than a copyleft licence. This is a description of the licence text and not legal advice; check the terms against your own distribution model.
The upgrade cost is dominated by the engine version matrix rather than by the library's own API. Because kvcached hooks into vLLM and SGLang internals, an engine upgrade can move you outside the tested range in the README table, and the README already warns that the tested ceiling for vLLM is v0.24.0 and for SGLang is v0.5.15. Budget for re-validating your model and KV layout after either side moves, and use issue #425 as the starting point for what has been checked.
On activity: the last push to the default branch was on 2026-08-23, and the repository is not archived. The release cadence visible in the release list is roughly one release per one to two months across v0.1.3 through v0.1.5, with the most recent of those on 2026-04-07.
Editorial conclusion
Adopt kvcached when you already run SGLang or vLLM on a GPU you must divide between several models, and when you can match the engine version to the tested range in the README table. Do not adopt it for a single model that already fills the device, or for an engine outside that table. Before committing, verify the per-model KV layout results linked from the README for your exact model and engine pair, and confirm the build finds a CUDA or ROCm PyTorch, since setup.py raises an error when neither torch.version.cuda nor torch.version.hip is set.
Frequently asked questions
Does kvcached work with vLLM and SGLang?
Yes, those are the two engines the README lists. vLLM is supported from v0.8.4 and tested up to v0.24.0, and SGLang from v0.4.9 and tested up to v0.5.15, across MHA, GQA, MLA, sliding window and hybrid attention types.
What is the difference between virtual memory and cache in kvcached?
kvcached borrows the OS virtual memory idea for the GPU: the serving engine reserves virtual addressing first, and physical GPU memory is attached later when the KV cache is actually used. The cache is the data being stored; the virtual memory abstraction is the mechanism that decides when physical memory gets committed to it.
How much GPU memory does kvcached need?
The README does not state a fixed memory requirement. Allocation is demand-driven, and the kvctl CLI is documented as the way to enforce memory limits, so the ceiling is something you set rather than a number the project publishes.
How do I install kvcached?
Install PyTorch 2.6.0 or newer first, then run pip install -e . from the repository root. setup.py raises an error if torch is missing or if PyTorch reports neither a CUDA nor a ROCm backend.
Does kvcached support prefix caching?
Yes. The 2026-04 update added automatic prefix caching for vLLM and RadixCache for SGLang, both with a configurable memory bound. The README points to examples/09_prefix_caching for details.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/ovg-project-kvcached)
Community notes