# Mooncake: the KVCache-centric transfer engine behind Kimi's serving stack

> Mooncake is Moonshot AI's open source serving platform for Kimi, built around a transfer engine and a distributed KV cache store. It is infrastructure for teams running multi-node LLM inference and training, not a drop-in library.

**kvcache-ai/Mooncake** — Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.

- Repository: https://github.com/kvcache-ai/Mooncake
- Website: https://kvcache-ai.github.io/Mooncake/
- Stars: 6,658 · Forks: 1,267
- Language: C++
- License: Apache-2.0
- Published: 2026-08-08 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/kvcache-ai-mooncake

## The problem Mooncake targets: KV cache that cannot move between nodes

In a disaggregated LLM deployment, prefill and decode run on different machines. The prefill worker produces a KV cache, the decode worker needs it, and the two are connected by a network. Doing that transfer through ordinary sockets wastes the interconnect you paid for, and it copies data through host memory that has no reason to touch it.

Mooncake's answer is a transfer engine that moves KV cache blocks directly between nodes, plus a distributed store that keeps those blocks addressable across instances. The README frames the whole project as "A KVCache-centric Disaggregated Architecture for LLM Serving", and states that under real workloads the architecture let Kimi handle 75% more requests while adhering to SLOs. That number comes from Moonshot's production deployment, not from a reproducible public benchmark, so treat it as a description of what the design targets rather than a number you can expect to reproduce.

The audience is narrow and specific: platform engineers running inference clusters, and researchers wiring rollout data between RL training and inference. If you serve one model on one GPU, nothing here applies to you.

## How the transfer engine and the store fit together

The repository is split into directories that map to distinct components rather than one monolith. mooncake-transfer-engine holds the RDMA-capable data movement layer. mooncake-store holds the distributed KV cache engine. Around those sit mooncake-p2p-store, mooncake-ep, mooncake-pg, mooncake-rl, mooncake-reshard and mooncake-conductor, plus mooncake-integration for the connectors that serving frameworks call into.

The data flow the README describes for integrations is consistent: a serving framework decides that a KV block should live somewhere other than the local GPU, hands the buffer to Mooncake, and Mooncake moves it zero-copy over RDMA to another node or into the store. The vLLM integration is documented as a KV Connector in PD-disaggregated setups, and SGLang uses Mooncake Store as a hierarchical KV caching storage backend that extends RadixAttention across device, host and remote storage tiers. SGLang's RDMA-based P2P weight transfer, described in the update notes, moves the 1T-parameter Kimi-K2 model's weights in 7.2s against 53s previously, again a figure reported by the integration's authors.

What makes this an architecture rather than a library is that the cache is treated as the primary resource. Placement decisions follow where KV blocks are, not where requests land.

## Installing mooncake-transfer-engine and running a first transfer

The Python package is published on PyPI as mooncake-transfer-engine, with separate wheels for different accelerators: mooncake-transfer-engine-cuda13, mooncake-transfer-engine-non-cuda, mooncake-transfer-engine-npu, mooncake-transfer-engine-musa, mooncake-transfer-engine-efa and mooncake-transfer-engine-rocm. Picking the wrong one is the most likely first failure.

```bash
pip install mooncake-transfer-engine
```

After installation the importable module is mooncake. The pyproject.toml declares dependencies on aiohttp, msgpack and requests, so those arrive automatically. Python 3.10 or newer is required; the classifiers list 3.10 through 3.13 and Linux only.

The repository also ships a Docker image under the kvcacheai/mooncake namespace, which is the path to take if you do not want to build the C++ side yourself. Building from source uses CMake, with dependencies.sh present at the repository root to fetch what the build needs.

The realistic first use is not a standalone script but an integration. For vLLM, the documented path is the Mooncake connector in a PD-disaggregated configuration; the vLLM docs page for mooncake_connector_usage is where the configuration lives. For SGLang, Mooncake Store is configured as the hierarchical KV caching backend. Both require you to have the serving framework already running across at least two nodes, which is why there is no meaningful single-machine hello world here.

## Where Mooncake is the wrong tool

The project is Linux-only. The classifiers in pyproject.toml list POSIX Linux and nothing else, and the transfer engine's value depends on RDMA hardware. On a single machine, or on nodes connected by plain Ethernet without RDMA, the engine has little to offer over a local cache, and the operational surface you take on is real: a C++ build, accelerator-specific wheels, and a store that has to be deployed alongside your serving stack.

The second limitation is documentation shape. The top-level README is dominated by a dated list of integrations and blog links. It does not document rollback, failure recovery for the store, or what happens when a node holding cached blocks disappears mid-transfer. Those questions have to be answered from the component directories and the linked integration docs, and for some of them the README is simply silent.

Third, the project is tightly coupled to the frameworks it integrates with. Mooncake Store is not a general-purpose Redis replacement you point arbitrary applications at. It exists to serve KV cache blocks to vLLM, SGLang, TensorRT LLM and similar systems, and its API surface is shaped by that.

## Mooncake against LMCache and plain prefix caching

The closest comparison in this space is LMCache, which also targets KV cache reuse across instances. The difference in approach is where the intelligence sits. LMCache is built as a caching layer that plugs into vLLM and other engines and manages reuse at the request level, with storage backends underneath. Mooncake's design starts from an RDMA transfer engine and a distributed store, and the serving framework calls into it as a connector; the cache is the architecture, not an add-on to the scheduler.

That distinction shows up in what each project integrates with. Mooncake's update list includes TensorRT LLM, vLLM Ascend, SGLang HiCache, FlexKV, LightX2V, TorchSpec and Speculators, and it has joined the PyTorch Ecosystem. Several of those are training and RL paths rather than pure inference, which reflects the transfer engine's generality: it moves hidden states and weights, not only KV cache.

Plain prefix caching inside a single engine is the other alternative, and for a single-node deployment it is the correct one. Mooncake earns its place when the cache has to cross a network boundary between nodes.

## Maintenance, licence and what an upgrade costs

The repository is not archived, and the most recent push recorded is 2026-08-26, which is the same date as the v0.3.13 release. Before that came v0.3.12.post1 on 2026-07-25 and v0.3.12 on 2026-07-23. The cadence over that window is roughly monthly, with a post-release immediately after v0.3.12, which suggests fixes are shipped promptly rather than batched.

Upgrade cost depends on which layer you consume. If you use the Python package, the version in pyproject.toml is 0.3.12.post1 while the latest tagged release is v0.3.13, so the wheel and the tag are not always in lockstep; check which one your environment resolves to. If you build from source, you inherit the CMake and dependency chain in dependencies.sh and the accelerator-specific wheel matrix, which means an upgrade is a rebuild rather than a pip bump.

The licence is Apache-2.0, with the file present as LICENSE-APACHE and declared in pyproject.toml. Apache-2.0 is permissive and includes a patent grant, but this is not legal advice; if you are redistributing Mooncake inside a product, read the licence text and your own obligations rather than relying on the classifier line that says Apache Software License.

## Conclusion

Adopt Mooncake if you run multi-node LLM inference or RL training on Linux with RDMA-capable hardware and you already know which integration you need, such as the vLLM KV connector or the SGLang HiCache backend. Do not adopt it if you serve a single model on one GPU, if you need Windows or macOS, or if you want a self-contained cache server with no engine-side changes. Before committing, verify which of the wheel variants matches your accelerator (the plain mooncake-transfer-engine, the cuda13, non-cuda, npu, musa, efa or rocm builds), confirm the Python version your environment runs against the requires-python >=3.10 floor, and read the integration docs for your serving stack rather than the top-level README, which is mostly a changelog of adoptions.

## FAQ

### What is Mooncake?

Mooncake is the serving platform for Kimi, the LLM service operated by Moonshot AI. The README describes it as a KVCache-centric disaggregated architecture for LLM serving, built around a transfer engine and a distributed KV cache store.

### How do I install Mooncake?

The Python package is published on PyPI as mooncake-transfer-engine, installed with pip. Separate wheels exist for cuda13, non-cuda, npu, musa, efa and rocm accelerators, and a Docker image is published as kvcacheai/mooncake.

### Which operating systems does Mooncake support?

The pyproject.toml classifiers list POSIX Linux only, and the package requires Python 3.10 or newer. There is no Windows or macOS classifier in the project metadata.

### What licence is Mooncake released under?

Apache-2.0. The repository contains a LICENSE-APACHE file and pyproject.toml declares the license as that file.

## Sources

- [Official documentation](https://kvcache-ai.github.io/Mooncake/)
- [Official README](https://github.com/kvcache-ai/Mooncake#readme)
- [Project repository](https://github.com/kvcache-ai/Mooncake)
- [Release notes](https://github.com/kvcache-ai/Mooncake/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/kvcache-ai-mooncake
