Model or dataset
llm-d/llm-d avatar
llm-d/llm-d

llm-d: a Kubernetes inference stack that sits above vLLM and SGLang

Achieve state of the art inference performance with modern accelerators on Kubernetes

4,660 stars795 forksShellApache-2.0

At a glance

What is it?
llm-d is a CNCF sandbox project that adds routing, KV-cache tiering, prefill/decode disaggregation and autoscaling on top of existing model servers. The judgment: it is a platform team's tool, not a single-node serving shortcut.
Who is it for?
Adopt llm-d if you already run Kubernetes, already serve models through vLLM or SGLang, and have a measured problem with prefix cache reuse, TTFT or GPU utilization at scale. Do not adopt it if you are serving one model on one GPU, if you have no Kubernetes operational capacity, or if you cannot yet say what your current TTFT and tokens/sec numbers are.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 3 days ago.
What is it written in?
Mainly Shell, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 26, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap llm-d fills between a model server and a production cluster

vLLM and SGLang are good at running one model on one set of accelerators. They are not trying to be a cluster scheduler. The llm-d README states this division directly: model servers "handle efficiently running large language models on accelerators," while llm-d "provides state-of-the-art orchestration and optimizations above model servers to serve high-scale real-world traffic efficiently and reliably."

That framing matters because it tells you who the project is for. The intended reader is a platform or infrastructure engineer who already has a working vLLM or SGLang deployment and has hit a wall that is not about the model server itself. The wall usually looks like one of four things: requests that repeat long prefixes are being routed to replicas that have never seen that prefix; a single GPU cannot hold the KV cache for the concurrency you want; a model is too large for one node; or traffic spikes are being handled by autoscaling rules that watch CPU instead of inference signals.

llm-d organizes its work into five themes, and they map onto those four problems almost one to one. Intelligent routing covers prefix-cache-aware and load-aware balancing plus predicted-latency scheduling. Advanced KV-cache management covers tiered offloading to CPU or disk with a global index of cache state. Serving large models covers prefill/decode disaggregation and wide expert parallelism. Operational excellence covers flow control for multi-tenant serving and SLO-aware autoscaling. Batch processing covers offline workloads through OpenAI-compatible Batch APIs.

None of that is a model server feature. All of it requires a control plane that knows about the pods, the nodes and the network between them. That is why the project is Kubernetes-native rather than a Python library you import.

Intelligent routing and the EPP: what actually moves requests

The routing layer is the part most people will touch first, and it is the part with the clearest published numbers. The README cites partner benchmarks showing 3x higher output throughput and 2x faster TTFT with prefix-cache-aware routing versus round-robin for Llama 3.1 70B on 4 AMD MI300X, and a 40% reduction in TTFT and ITL with predicted-latency scheduling versus heuristics on NVIDIA GPUs.

The mechanism behind the first number is worth stating plainly, because it explains why the gain is so large. Round-robin spreads requests evenly across replicas. If your traffic has repeated prefixes, that even spread is actively harmful: each replica ends up holding a partial, mostly cold cache, and every request pays prefill cost again. Prefix-cache-aware routing sends a request to the replica that already holds the matching KV blocks. The compute you save is the prefill pass you no longer run.

The second mechanism, predicted-latency scheduling, is the more experimental one. The README describes it as "experimental predicted latency-based scheduling" in the theme list, and the release notes for v0.7 record that predicted-latency scheduling reached GA. Those two statements are from different points in the project's history, so check which one applies to the version you deploy. This is the kind of detail that decides whether a feature is safe for production traffic.

The component that carries this out is referred to in the project's own vocabulary as the EPP, and the well-lit path guides describe how to configure the intelligent router. The guides are the authoritative source for the configuration keys; the README does not reproduce them.

Installing llm-d and deploying a first optimized baseline

There is no pip install and no single binary. llm-d is a set of Helm charts, kustomize manifests and container images, and the documented entry point is the Quickstart Guide at llm-d.ai, with the Optimized Baseline well-lit path given as the recommended starting configuration for most users.

The repository also carries a Makefile for building images, which is useful to understand because it tells you what the project actually ships. The device, architecture and OS are all build-time variables, and the CUDA variant is the default. The Makefile declares the image target as follows:

makefile
IMG := $(IMAGE_BASE):$(VERSION)

The Makefile defaults set DEVICE to cuda, ARCH to amd64, OS to rhel, CUDA_VERSION to 13.0 and BUILD_TYPE to dev. The DEVICE variable also accepts xpu and cpu, and the OS variable accepts rhel or ubuntu, which selects the ubi9 or ubuntu24.04 base image suffix respectively. The image base is ghcr.io/llm-d/$(PROJECT_NAME)-$(DEVICE), with a -dev suffix appended when BUILD_TYPE is dev. Note that the Makefile's own VERSION default is v0.2.1 while the latest tagged release is v0.9.0; do not assume the Makefile default matches the release you want.

For an actual deployment, follow the Quickstart rather than the Makefile. The sequence the documentation describes is: set up the llm-d stack on Kubernetes, configure the intelligent router, then validate with the project's benchmarks. The Optimized Baseline guide is the one to read first, because it is the configuration the project describes as a high-performance foundation for a wide range of serving cases.

What you should see after a successful baseline deployment is an inference endpoint that behaves like a standard OpenAI-compatible service, with the router in front of it. The README does not document a rollback procedure for the stack, so plan your own before you apply manifests to a cluster that matters.

Where llm-d is the wrong tool

The most common mismatch is scale. If you are serving one model on one GPU for an internal team, llm-d adds a router, additional pods and a set of configuration surfaces that have no work to do. Round-robin across a single replica is not a problem you have. The project's own performance claims are all framed around multi-replica, multi-node or multi-accelerator topologies, which is a fair signal about where the value lives.

The second mismatch is operational. Every theme in the README assumes a Kubernetes cluster you control, with accelerators attached and a network fabric between nodes. Prefill/decode disaggregation and wide expert parallelism depend on "fast accelerator interconnects," and the README does not claim those topologies work well without them. If your nodes are connected by ordinary Ethernet, the disaggregated paths are not a tuning exercise, they are a different cluster build.

The third is version drift. The README badge shows Version 0.8 while the latest release is v0.9.0, and the Makefile still defaults to v0.2.1. Predicted-latency scheduling is described as experimental in one place and GA in the release notes for another version. Treat the documentation as versioned, not universal, and pin your guides to your release.

Finally, llm-d does not replace a model server. It has no inference engine of its own. If vLLM or SGLang cannot run your model on your hardware, llm-d has nothing to orchestrate.

llm-d versus vLLM alone, and versus other serving stacks

The comparison people search for most is llm-d versus vLLM, and the honest answer is that it is not a versus. vLLM is a dependency, not a competitor. The README lists vLLM and SGLang as the model servers llm-d orchestrates. If you compare them as alternatives you will pick the wrong one, because the question is whether you need orchestration above your model server, not which model server to use.

A more useful comparison is against building the same orchestration yourself. A team that wants prefix-cache-aware routing can write a custom scheduler that queries each vLLM replica's cache state and scores replicas per request. That is a real option and it has the advantage of fitting whatever internal conventions you already have. What llm-d offers instead is a maintained set of these components, benchmarked together and shipped with Helm charts, plus the KV tiering, disaggregation and autoscaling pieces that a custom router alone would not give you. The trade is control for maintenance burden.

The other comparison worth making is between llm-d's own well-lit paths. The Optimized Baseline is the general-purpose starting point. The tiered prefix cache path targets multi-turn workloads where the working set outgrows GPU memory. The wide expert parallelism path targets massive models such as DeepSeek-R1 and GPT-OSS. These are not interchangeable configurations; picking the wrong guide is the most likely way to get disappointing results from a correct installation.

Maintenance, release cadence and licence

The repository is not archived, and the last push was on 2026-09-09, which is recent. The release history shows v0.9.0 on 2026-08-17, v0.8.1 on 2026-06-26 and v0.8.0 on 2026-06-24. That is roughly a two-month gap between the v0.8 line and v0.9.0, with a patch release in between, which is a fast enough cadence that you should expect to track releases rather than pin and forget.

The maintenance cost is not in the code you write, because you write almost none. It is in the cluster. Each release can change the router configuration, the Helm charts and the well-lit path guides together, and the project's own news entries describe structural changes at that level: the v0.7 release notes mention a renamed and stabilized optimized baseline, kustomize-first migrated guides, expanded nightly CI and a batch gateway marked experimental. A team upgrading across that boundary is re-reading guides, not bumping a version string.

The licence is Apache-2.0, which is a permissive licence that permits commercial use and modification and includes an explicit patent grant. This is a statement about the licence text, not legal advice; if your organisation has policies about CNCF sandbox projects or about the patent terms in your jurisdiction, route it through whoever handles that. The project is a CNCF sandbox project, founded by Red Hat, Google Cloud, IBM Research, CoreWeave and NVIDIA, with listed support from AMD, Cisco, Hugging Face, Intel, Lambda, Mistral AI, UC Berkeley and University of Chicago. Sandbox is the earliest CNCF maturity stage, which is worth weighing if your procurement process distinguishes between sandbox and graduated projects.

Editorial conclusion

Adopt llm-d if you already run Kubernetes, already serve models through vLLM or SGLang, and have a measured problem with prefix cache reuse, TTFT or GPU utilization at scale. Do not adopt it if you are serving one model on one GPU, if you have no Kubernetes operational capacity, or if you cannot yet say what your current TTFT and tokens/sec numbers are. Before committing, verify three things: that a well-lit path guide exists for your accelerator and model pair, that the router version you deploy matches the release you are on, and that your cluster can supply the interconnect the disaggregated topologies assume.

Frequently asked questions

What is llm-d?

llm-d is a high-performance distributed inference serving stack optimized for production deployments on Kubernetes. It provides orchestration and optimizations above model servers such as vLLM and SGLang, and is a CNCF sandbox project.

What is the difference between vLLM and llm-d?

vLLM runs large language models efficiently on accelerators, while llm-d provides orchestration above model servers for high-scale traffic. The README lists vLLM as one of the model servers llm-d works with, so they are layered rather than competing.

What does LLM-d stand for?

The repository does not expand the name anywhere in the README or the project files, so there is no documented meaning to report. The name appears only as llm-d, with the tagline about achieving state of the art inference performance on any accelerator.

Is llm-d open source?

Yes. The repository carries an Apache-2.0 licence and is a Cloud Native Computing Foundation sandbox project.

Official sources

  1. License: Apache-2.0
  2. llm-d/llm-d on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/llm-d-llm-d.svg)](https://hysenlabs.com/projects/llm-d-llm-d)