# LMCache: a KV cache layer that outlives the inference engine

> LMCache moves KV cache out of GPU memory into a tiered store that survives engine restarts and can be shared between engines. It is aimed at teams serving long-context, multi-turn or RAG traffic, and the trade-off is operational: you now run a cache tier alongside your model server.

**LMCache/LMCache** — LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

- Repository: https://github.com/LMCache/LMCache
- Website: https://lmcache.ai/
- Stars: 11,939 · Forks: 1,968
- Language: Python
- License: Apache-2.0
- Published: 2026-08-27 · Updated: 2026-08-27 · Language: en
- Canonical page: https://hysenlabs.com/projects/lmcache-lmcache

## The problem LMCache targets: prefill work that gets thrown away

Every request to an LLM serving engine runs a prefill pass over the prompt. That pass produces key and value tensors, one pair per layer per token, and those tensors are what the decode step attends over. In a plain vLLM or SGLang deployment they live in GPU memory and are evicted when the block table fills up. The next request that shares the same system prompt, the same retrieved document set, or the same conversation history pays for that prefill again.

LMCache takes the position that this state should not be treated as scratch space. The README describes the goal as turning KV cache from temporary state into reusable "AI-native knowledge" that can be stored persistently, reused across serving engines, monitored, and transformed. The workloads it names are long-context agentic runs, multi-turn conversation, and knowledge-augmented generation such as RAG. Those are exactly the cases where prompts are long and largely identical between calls, so the ratio of reused tokens to fresh tokens is high.

The intended audience is infrastructure engineers running their own inference, not application developers calling a hosted API. You need control over the serving engine process, the GPU host, and whatever storage tier you point the cache at.

## How the cache layer sits beside the engine instead of inside it

The architectural choice that separates LMCache from the prefix caching built into serving engines is that it runs as a standalone daemon process. The README states this directly under engine-independent deployment: LMCache manages KV cache independently from the inference engine process, so the cache is not lost if the engine crashes. The README calls this "no fate-sharing with engines".

In practice that means the engine hands KV blocks to LMCache over a transport, and LMCache decides where they live. The README describes a tiered hierarchy spanning CPU memory, local storage, and remote backends, with a pluggable backend interface. The backends it lists are CPU RAM, local disk (SSD), Redis/Valkey, Mooncake, InfiniStore, S3-compatible object storage, NIXL, and GDS. Because the cache is not owned by the engine, two engines pointed at the same LMCache deployment can reuse each other's blocks.

Two features go beyond ordinary prefix caching. Non-prefix KV reuse lets cached blocks be reused at any position in the prompt rather than only as a leading prefix, and the README says this uses CacheBlend to selectively recompute tokens for quality recovery. That recomputation is the honest part of the design: reusing a block that was computed in a different context changes the attention result, so some tokens are recomputed to limit the damage. Separately, prefill-decode disaggregation is supported, with KV cache transferred from prefill workers to decode workers over NVLink, RDMA, or TCP through transport layers such as NIXL.

The repository layout matches this story. There is a csrc/ directory for C++ extensions, a rust/ directory, a separate operator/ directory that suggests Kubernetes-facing work, and an examples/ directory with subdirectories for disagg_prefill, disagg_prefill_mp, p2p, redis_lookup, cache_controller, observability, and a kv_cache_calculator. The presence of a calculator example is a small signal that sizing the cache tier is a real question rather than an afterthought.

## Installing LMCache and running a first request against vLLM

The README gives a single install command and points to the documentation site for anything beyond it.

```bash
pip install lmcache
```

The package on PyPI is named lmcache. The README does not inline the engine-side configuration, so the concrete wiring for a vLLM deployment has to come from the documentation, which the README links at docs.lmcache.ai under an Installation page. What can be said from the repository itself is that the package ships a CLI entry point (there is a pyproject_cli.toml at the top level alongside pyproject.toml), and that the examples directory contains runnable configurations rather than prose, including examples/cache_with_configs/ for configuration-driven setups and examples/kv_cache_reuse/ for the reuse path.

The build metadata is worth reading before you install on a fresh machine. pyproject.toml pins torch==2.13.0 in the build requirements and notes in a comment that the recommended source install avoids build isolation so the torch version stays flexible, which means the invocation for a source build is:

```bash
pip install -e . --no-build-isolation
```

The same file declares requires-python as >=3.10,<3.14, and the classifiers list Linux and GPU environments only. There is no Windows or macOS classifier. If you are on a Mac, the package metadata does not claim to support you.

For a first real use, the honest starting point is examples/kv_cache_reuse/ or examples/cache_with_configs/ rather than a from-scratch config, because the README does not document the engine-side flags inline and inventing them would be guesswork. Point the engine at the LMCache daemon, send the same long prompt twice, and watch the second request's prefill work drop. That is the behaviour the project is built around.

## Where LMCache is the wrong tool

The package classifiers say Development Status :: 3 - Alpha. That is the project's own label, not an outside assessment, and it should set expectations about API stability across releases. The release list shows nightly builds tagged by CUDA version (nightly, nightly-cu129) and by ROCm version (nightly-rocm), which tells you the project tracks upstream accelerator stacks closely. Tracking closely cuts both ways: a nightly that moves to a new CUDA or ROCm baseline can leave your pinned engine version behind.

There is also a real operational cost. LMCache adds a daemon, a transport, and at least one storage backend to a deployment that previously needed only a model server. The README lists Redis/Valkey, Mooncake, InfiniStore, S3-compatible object storage, NIXL, and GDS as options, and each of those is something you have to run, size, and monitor. If your prompts are short and mostly unique, the cache hit rate will be low and you will have paid that cost for nothing. Prefix caching already built into the serving engine handles the simple repeated-system-prompt case without a second process.

The non-prefix reuse path deserves particular caution. Because it reuses blocks computed in a different context, it relies on selective recomputation for quality recovery. That is a deliberate accuracy-for-speed trade, and the README does not quantify the accuracy impact. If your application is sensitive to small output differences, evaluate that path on your own evaluation set before enabling it, rather than assuming reuse is free.

## LMCache compared with vLLM's built-in prefix cache and with Mooncake

The most common comparison is LMCache against vLLM, and the two are not competitors in the usual sense. vLLM is the serving engine; LMCache is meant to be used with it. vLLM ships a prefix cache that lives in GPU memory and is scoped to the engine process and its lifetime. LMCache externalizes that state into a tiered store and keeps it alive across requests, sessions, and engine instances. If you restart your vLLM process, the built-in prefix cache is gone; the README's stated design goal is that LMCache's cache is not.

The second comparison the search data surfaces is LMCache against Mooncake. Mooncake appears in the README as one of the supported storage backends behind LMCache's pluggable interface, which makes the relationship closer to integration than rivalry. The difference in approach is where the intelligence lives: a dedicated KV store like Mooncake owns its own storage and transfer design, while LMCache positions itself as an engine-independent layer that can talk to several backends, including Mooncake, through one interface. The README frames this as vendor neutrality, letting users switch between serving engines and storage vendors while reusing stored KV caches. That framing is the project's actual differentiator, and it is also the thing to test: neutrality only helps if the backends you care about are actually implemented.

A third point of comparison is the observability surface. The README claims production-level KV cache metrics covering request-level and token-level prefix cache hit rates, cache lifecycle, request-level performance, and per-user usage. A built-in engine prefix cache typically exposes far less than that. If you need to explain to a team why TTFT moved, this is the feature that matters.

## Licence, maintenance and the cost of keeping up

LMCache is Apache-2.0, declared in pyproject.toml as license = "Apache-2.0" with license-files = ["LICENSE"]. The repository also carries a DCO file and a SECURITY.md, and the project states it joined the PyTorch Foundation in October 2025. Apache-2.0 is a permissive licence that allows commercial use and modification; it also includes an explicit patent grant. Whether that fits your organisation's policy is a question for your legal team, not something this article can settle.

The last push to the default branch was on 2026-08-29, and the nightly releases carry the same date. That is recent enough that the project is clearly being worked on, and the update log in the README runs through May 2026 with entries about an AMD MI300X agentic workload benchmark and a multiprocess architecture release. The default branch is dev rather than main, which is worth knowing if you clone the repository expecting a stable trunk.

Upgrade cost is the part that deserves more attention than it usually gets. The build requirements pin grpcio==1.78.0, grpcio-tools==1.78.0, and torch==2.13.0, and the pyproject comment explains that the torch pin exists so wheels can be released in sync with vLLM. That is a coupling: if you upgrade vLLM, you may need to move LMCache with it. The nightly tags are split by accelerator stack (CUDA 12.9, CUDA 13.0, ROCm 7.2 on gfx942 and gfx950), so the artefact you install is tied to your hardware generation. Budget for a coordinated upgrade of engine, LMCache, and accelerator stack rather than treating LMCache as a dependency you can bump independently.

## Conclusion

Adopt LMCache if your traffic repeats long prefixes across requests, sessions or engine instances and you can run a cache daemon next to your model server. Do not adopt it if your prompts are short and unique, or if you cannot operate an extra storage tier. Before committing, verify three things against your own deployment: which engine version your pinned LMCache build targets, whether your chosen storage backend is on the supported list, and what your TTFT looks like with the cache cold.

## FAQ

### What does LMCache do?

It is a KV cache management layer for LLM inference. It moves KV cache out of GPU memory into a tiered store spanning CPU memory, local storage and remote backends, so cached state can be reused across requests, sessions and engine instances instead of being recomputed.

### What are the key differences between LMCache and vLLM?

vLLM is the serving engine; LMCache is a layer used alongside it. vLLM's prefix cache lives in GPU memory and dies with the engine process, while LMCache runs as a standalone daemon, keeps KV cache in a tiered store, and per the README does not share fate with the engine.

### How do you use LMCache with vLLM?

The README gives only the install step, pip install lmcache, and links the documentation site for setup. The concrete engine-side configuration is not inline in the README, so the examples directory, particularly examples/cache_with_configs/ and examples/kv_cache_reuse/, is the place to start from.

### What is L2 cache used for in LMCache?

The README describes a tiered storage hierarchy spanning CPU memory, local storage and remote backends, but it does not define an L2 tier by name. The repository does contain an examples/lmc_external_l2_adapter/ directory, which indicates an adapter for an external second-level cache, though the README does not document its configuration.

### What is LMCache?

It is described in the README as a KV cache management layer for scalable LLM inference, and in pyproject.toml as a serving engine extension that reduces time-to-first-token and increases throughput, especially in long-context scenarios.

### How does LMCache compare with other KV cache options?

The README positions LMCache as vendor-neutral: it can serve as a KV cache layer for a range of serving engines and storage systems, and it lists Redis/Valkey, Mooncake, InfiniStore, S3-compatible object storage, NIXL and GDS as backends behind one interface. The README does not include a feature-by-feature comparison table.

## Sources

- [Official documentation](https://lmcache.ai/)
- [Official README](https://github.com/LMCache/LMCache#readme)
- [Project repository](https://github.com/LMCache/LMCache)
- [Release notes](https://github.com/LMCache/LMCache/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/lmcache-lmcache
