# vLLM Production Stack: a Helm-deployed Kubernetes inference stack

> The vLLM Production Stack is a reference implementation for running vLLM across a Kubernetes cluster, with a request router, Prometheus and Grafana observability, and optional KV cache offloading. It is a good fit if you already run Kubernetes with GPUs and want the OpenAI-compatible API without rewriting clients.

**vllm-project/production-stack** — vLLM’s reference system for K8S-native cluster-wide deployment with community-driven performance optimization

- Repository: https://github.com/vllm-project/production-stack
- Website: https://docs.vllm.ai/projects/production-stack
- Stars: 2,631 · Forks: 509
- Language: Python
- License: Apache-2.0
- Published: 2026-09-28 · Updated: 2026-09-28 · Language: en
- Canonical page: https://hysenlabs.com/projects/vllm-project-production-stack

## The gap the vLLM Production Stack fills between one vLLM process and a cluster

A single vLLM process serves an OpenAI-compatible API on one machine. The moment you want more than one replica, someone has to decide which replica receives each request, how clients discover the endpoints, and where the metrics go. The README frames the project as "a reference implementation on how to build an inference stack on top of vLLM", and its stated goals are scaling from a single instance to a distributed deployment without changing application code, monitoring through a web dashboard, and gains from request routing and KV cache offloading.

The audience is therefore narrow and specific: teams that already run Kubernetes with GPUs and want the deployment shape decided for them rather than assembled from scratch. If you serve one model on one box, this stack adds a control plane you do not need. The repository is Apache-2.0 licensed, the default branch is main, and the last push was on 2026-09-22.

## Serving engine, router and observability stack: how the pieces connect

The architecture section names three parts. The serving engine is one or more vLLM engines running the models. The request router directs requests to backends "based on routing keys or session IDs to maximize KV cache reuse". The observability stack scrapes backend metrics through Prometheus and renders them in Grafana. Helm ties the three together.

The router is the part with the most documented behaviour. It routes to endpoints running different models, exports per-instance metrics (QPS, time to first token, pending, running and finished request counts, uptime), discovers services and tolerates faults through the Kubernetes API, supports model aliases, and offers round-robin and session-ID based routing. Prefix-aware routing is marked WIP in the README. That label matters more than the feature list: if your workload depends on prefix cache hits across replicas, the routing algorithm you want is not finished.

Because the router runs inside the cluster and speaks the OpenAI API, client code does not change when you go from one vLLM instance to many. That is the whole design bet, and it is a reasonable one.

## Installing the vLLM production-stack helm chart and sending a first request

The prerequisites are a running Kubernetes environment with GPUs. The README points at `cd utils && bash install-minikube-cluster.sh` for a local cluster, or the tutorial in tutorials/00-install-kubernetes-env.md. The deployment itself is three commands from the README: clone the repository, add the vllm Helm repository, and install the chart with the minimal example values file.

```bash
git clone https://github.com/vllm-project/production-stack.git
cd production-stack/
helm repo add vllm https://vllm-project.github.io/production-stack
helm install vllm vllm/vllm-stack -f tutorials/assets/values-01-minimal-example.yaml
```

After the install completes, the stack exposes the same OpenAI API interface as vLLM, reachable through a Kubernetes service. The README does not print the exact service name or port in the section quoted here, so read tutorials/01-minimal-helm-installation.md for the validation query rather than guessing the endpoint. To remove everything, the uninstall is a single Helm command:

```bash
helm uninstall vllm
```

Configuration lives in the chart's values file. The README directs you to helm/values.yaml for the available keys and to the tutorials for changes such as loading weights from a persistent volume, launching multiple models, or offloading KV cache with LMCache.

## Where the reference stack stops short: autoscaling, prefix routing and rollback

The roadmap is the honest part of the README. Autoscaling based on vLLM-specific metrics, support for disaggregated prefill, and router improvements (a faster router in a non-Python language, KV-cache-aware routing, better fault tolerance) are all listed as work the project will release "soon". None of them are described as shipped. If your capacity plan assumes the stack scales replicas on its own, it does not, based on the README.

The second limitation is the router itself. Prefix-aware routing is WIP, so cache reuse across replicas is only as good as round-robin or session-ID routing allows. The examples directory does contain disaggregated_prefill and disaggregated_prefill_orchestrated folders, which suggests work in that area, but the README lists disaggregated prefill as a roadmap item, and the two statements do not resolve into a supported feature in the documentation shown here.

Third, the README documents install and uninstall but says nothing about upgrading between chart versions or rolling back a failed release. Helm has its own mechanisms for that; the project documentation does not describe how a running stack behaves across a version bump, so treat upgrade behaviour as something you verify yourself before trusting it.

## vLLM production stack vs llm-d and vs KServe: what actually differs

The comparison that comes up most is against llm-d, another Kubernetes-oriented serving effort in the vLLM orbit. The distinction visible here is scope and packaging: the Production Stack ships as a Helm chart whose components are a vLLM serving engine, a Python router (the package is named vllm-router, requires Python 3.12 or newer, and installs a vllm-router console script), and a Prometheus and Grafana pair. The router's routing keys are session IDs and, eventually, prefixes. Whether llm-d makes the same choices is not something this repository's documentation answers, so do not take a comparison from here at face value.

Against KServe, the difference is the layer you own. KServe is a general model-serving control plane with its own inference service abstractions. The Production Stack does not introduce a new serving CRD; it wires a router in front of vLLM engines and leaves the OpenAI API as the contract. If you already standardized on KServe's abstractions, adopting this stack means running a second, overlapping control plane. If you only want vLLM behind a router, the smaller surface is the point.

## Licence, maintenance and the cost of keeping the stack current

The project is Apache-2.0, stated in the README and in the pyproject.toml licence field. Apache-2.0 permits commercial use and modification and includes a patent grant; it also requires that you preserve notices. That is a summary of the licence text, not legal advice, and your own counsel should review how it interacts with the rest of your distribution.

Maintenance signals are mixed in a specific way. The last push was on 2026-09-22, days before this writing, and the most recent tagged release is vllm-stack-0.1.12 from 2026-07-24. Releases are spaced months apart: vllm-stack-0.1.11 came on 2026-05-07 and vllm-stack-0.1.10 on 2026-02-27. So the repository moves, but the chart you install changes on a quarterly cadence. The router package pins its dependencies tightly (kubernetes 36.0.3, uvicorn 0.34.0, numpy 1.26.4, and so on), which reduces surprise but also means an upgrade of any single component is a coordinated change. The optional extras, lmcache and semantic_cache, pull in vllm 0.13.0 and sentence-transformers 2.2.2 respectively, so the KV offloading path carries its own version constraints. Budget for reading the values file diff at each chart release rather than assuming a drop-in upgrade.

## Conclusion

Adopt it if you already run Kubernetes with GPUs and want vLLM's OpenAI-compatible API behind a router plus Prometheus and Grafana, installed from the vllm Helm repository. Do not adopt it if you have no Kubernetes cluster or expect autoscaling and disaggregated prefill today, since both are listed as roadmap items. Before committing, verify that the router algorithm you need is implemented rather than work in progress, and check the helm/values.yaml keys against your cluster.

## FAQ

### What is the vLLM Production Stack?

It is a reference implementation for building an inference stack on top of vLLM, deployed with Helm and made of a serving engine, a request router, and a Prometheus plus Grafana observability stack.

### How does the vLLM Production Stack compare with llm-d?

The README does not document a comparison with llm-d. What it does describe is this project's own shape: a Helm chart, a Python router that routes by session ID or routing key, and Prometheus and Grafana for metrics.

### How does the vLLM Production Stack compare with KServe?

The repository does not document a KServe comparison. The visible difference is that this stack keeps the OpenAI API as the contract and adds a router in front of vLLM engines, rather than introducing a separate serving abstraction.

## Sources

- [License: Apache-2.0](https://github.com/vllm-project/production-stack/blob/main/LICENSE)
- [Project website](https://docs.vllm.ai/projects/production-stack)
- [README](https://github.com/vllm-project/production-stack/blob/main/README.md)
- [Releases](https://github.com/vllm-project/production-stack/releases)
- [vllm-project/production-stack on GitHub](https://github.com/vllm-project/production-stack)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/vllm-project-production-stack
