# AIBrix: Kubernetes building blocks for vLLM inference, from gateway to KV cache

> AIBrix is a Go control plane that adds LLM-aware routing, autoscaling, LoRA management and distributed KV cache to vLLM on Kubernetes. It is a platform team's tool, not a single binary, and the quickstart reflects that.

**vllm-project/aibrix** — Cost-efficient and pluggable Infrastructure components for GenAI inference.

- Repository: https://github.com/vllm-project/aibrix
- Stars: 5,113 · Forks: 701
- Language: Go
- License: Apache-2.0
- Published: 2026-08-04 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/vllm-project-aibrix

## What AIBrix adds on top of a plain vLLM deployment

A vLLM server exposes an OpenAI-compatible HTTP API. Put it behind a Kubernetes Service and you get round-robin routing, which is a poor fit for LLM traffic: two replicas can hold very different KV cache pressure, and a request that lands on the loaded one pays for it in latency. AIBrix exists to fill that gap. The README describes it as "essential building blocks to construct scalable GenAI inference infrastructure", and the key feature list reads like a checklist of the things a vLLM deployment lacks by default: an LLM gateway and routing layer, an app-tailored autoscaler, high-density LoRA management, a unified runtime sidecar, distributed inference, distributed KV cache, heterogeneous serving, and GPU hardware failure detection.

The intended user is a platform or infrastructure engineer who already runs Kubernetes and already runs vLLM. Nothing in the repository suggests AIBrix is useful without both. The Go module is a Kubernetes operator and control plane, the install instructions are kubectl apply commands, and the samples directory is organised by cluster concern (autoscaling, kvcache, disaggregation, heterogeneous, modelclaim). If you are serving one model on one GPU box, this is more machinery than the problem needs.

## The architecture: a controller manager, gateway plugins and a runtime sidecar

The repository layout tells most of the story. cmd/ holds the entry points, pkg/ holds the shared logic, config/ holds the Kustomize bundles that the quickstart applies, and api/ holds the CRD types. The Makefile names the images it builds: controller-manager, gateway-plugins, runtime and metadata-service. Those four names map onto the four jobs the project does.

The controller-manager reconciles the custom resources. The gateway-plugins image is the LLM-aware routing layer, and the go.mod dependency on github.com/envoyproxy/go-control-plane indicates an Envoy-based data plane rather than a bespoke proxy. The runtime image is the sidecar the README calls a "unified AI runtime", handling metric standardization and model downloading next to the inference server. The metadata-service backs the shared state, and the go.mod dependency list (redis/go-redis, miniredis for tests, consistent hashing via buraksezer/consistent, xxhash) points at a distributed cache and routing table rather than a single-process design.

Two dependencies are worth flagging because they shape what you are buying into. The go.mod requires github.com/ray-project/kuberay/ray-operator, so Ray-based distributed inference is a first-class path, not an afterthought. It also requires github.com/openai/openai-go, which confirms the API surface is OpenAI-compatible rather than a custom protocol. The pkoukk/tiktoken-go dependency implies token counting happens inside the control plane, which is what makes token-aware routing and autoscaling possible at all.

## Installing AIBrix v0.7.0 and running the quickstart sample

The README gives two install paths: a nightly path from a cloned repository, and a stable path from release manifests. The stable path is three kubectl apply commands against the v0.7.0 release assets. Note the --server-side flag on the first two, and that the CRDs are deliberately shipped as a separate manifest from the operator.

```bash
# Install component dependencies
kubectl apply -f "https://github.com/vllm-project/aibrix/releases/download/v0.7.0/aibrix-dependency-v0.7.0.yaml" --server-side

# Install AIBrix CRDs (separate from the operator so uninstalls don't wipe user CRs)
kubectl apply -f "https://github.com/vllm-project/aibrix/releases/download/v0.7.0/aibrix-core-crds-v0.7.0.yaml" --server-side

# Install aibrix components
kubectl apply -f "https://github.com/vllm-project/aibrix/releases/download/v0.7.0/aibrix-core-v0.7.0.yaml"
```

After the third command, the four image names from the Makefile should appear as workloads in the cluster. The README does not state which namespace they land in or what a healthy rollout looks like, so check the pods the apply created rather than assuming a fixed namespace.

If you want to build from source instead, the nightly path clones the repository and applies the Kustomize overlays directly. The dependency bundle comes first, then the CRDs, then the components.

```bash
git clone https://github.com/vllm-project/aibrix.git
cd aibrix

# Install nightly aibrix dependencies
kubectl apply -k config/dependency --server-side

# Install nightly AIBrix CRDs (separate from the operator so uninstalls don't wipe user CRs)
kubectl apply -k config/crd --server-side

# Install nightly aibrix components
kubectl apply -k config/default
```

For a first real use, the repository ships samples/quickstart/, and the README points at the documentation site at aibrix.readthedocs.io/latest/ for configuration detail. The samples directory is the honest starting point: samples/autoscaling/, samples/kvcache/, samples/heterogeneous/ and samples/modelclaim/ each correspond to one advertised feature, so you can apply a single sample instead of the whole feature set. The README itself does not walk through a sample end to end, so expect to read the manifests before applying them.

## Where AIBrix is the wrong tool

The install procedure is the first limitation. Three manifests, a dependency bundle that pulls in third-party components, and a CRD bundle that must be applied server-side is a real operational commitment. The README explains why CRDs are separate ("so uninstalls don't wipe user CRs"), which is a sensible design, but it also means uninstalling AIBrix is a two-step process the README does not document. There is no rollback section and no upgrade section in the README; version-to-version migration is left to the release notes and the documentation site.

The second limitation is scope. Every feature in the list assumes a Kubernetes cluster and a vLLM-compatible inference server. The unified runtime is a sidecar, so it needs a pod to attach to. The distributed KV cache needs multiple engines to share across. The heterogeneous serving feature needs more than one class of GPU. On a single node, or with a non-vLLM server, most of the value disappears and you are left carrying the control plane's weight for nothing.

The third is maturity signalling rather than maturity itself. The release cadence in the README is roughly quarterly: v0.5.0 in November 2025, v0.6.0 in March 2026, v0.7.0 in June 2026. The last push to main was on 2026-06-18, three months before this writing. That is a normal cadence for a project of this size, but it means you should not expect a fix for a bug you file today to land this week. Treat the pinned release manifest as the unit of adoption, not main.

## How AIBrix differs from llm-d and from plain vLLM

The comparison people search for is AIBrix versus llm-d. Both are Kubernetes control planes for LLM inference and both are associated with the vLLM ecosystem, but the README only describes AIBrix's side of the difference, so treat the contrast as directional. AIBrix is built around Envoy for the data plane (the go-control-plane dependency) and around Ray for distributed execution (the kuberay ray-operator dependency). Its distinguishing claim in the README is the breadth of the feature list rather than depth in any one area: LoRA management, routing, autoscaling, runtime, distributed inference, KV cache, heterogeneous serving and GPU failure detection all appear in the same initial release.

The comparison with plain vLLM is easier to state precisely. vLLM is the inference engine; AIBrix does not replace it and does not ship a model server. What AIBrix adds is everything around the engine: who receives a request, how many replicas exist, whether KV cache is reused across them, and whether a failing GPU is detected before it takes down a replica. If your vLLM deployment is a single replica behind a Service, AIBrix is solving problems you do not have yet. If it is twenty replicas across three GPU types with LoRA adapters loaded on some of them, that is the shape AIBrix is designed for.

## Licence, maintenance and upgrade cost

AIBrix is Apache-2.0. For most adopters that is the least restrictive of the common open source licences: it permits commercial use, modification and redistribution, and it includes a patent grant. The practical implication is that you can vendor the control plane into a proprietary platform without publishing your changes. That is a statement about the licence text, not legal advice; if you are redistributing a modified AIBrix, read the NOTICE and attribution requirements in the licence yourself.

The maintenance cost is dominated by the dependency bundle rather than the AIBrix code. The quickstart's first command installs third-party components, and the go.mod shows what the project itself leans on: Envoy's control plane, Ray's operator, Redis, OpenTelemetry, Prometheus client libraries and the OpenAI Go SDK. Upgrading AIBrix means upgrading that bundle, and the README does not describe a supported upgrade path between releases. The CRD separation helps here: because CRDs ship separately from the operator, an operator upgrade does not silently rewrite your custom resources. Plan upgrades as three applies in the same order as the install, and read the release notes for the version you are moving to before you run them.

## Conclusion

AIBrix fits platform teams already running vLLM on Kubernetes who need routing, autoscaling and KV cache behaviour that plain Kubernetes Services and HPAs do not provide. It does not fit a single-machine deployment, a team without cluster admin rights, or anyone who wants one binary: the quickstart applies a dependency bundle, a CRD bundle and a component bundle, and the README does not document rollback. Before adopting, verify three things in the v0.7.0 release manifests: which Kubernetes and vLLM versions the dependency bundle pins, whether the gateway plugin image matches your Envoy or gateway setup, and which autoscaler signals your metrics pipeline actually exports. Then run the samples/quickstart manifests on a non-production cluster and confirm the controller-manager and gateway-plugins pods reach Ready before pointing real traffic at them.

## FAQ

### How does AIBrix compare with Dynamo?

The README does not describe Dynamo, so a feature-by-feature comparison is not possible from what is documented here. What can be said is that AIBrix is an Apache-2.0 Kubernetes control plane for vLLM inference, built around Envoy for routing and Ray for distributed execution, and its README frames it as a set of building blocks rather than a single serving stack.

### How do I install AIBrix?

The README gives a stable path of three kubectl apply commands against the v0.7.0 release manifests: a dependency bundle, a CRD bundle and the core components, with --server-side on the first two. A nightly path clones the repository and applies config/dependency, config/crd and config/default with kubectl apply -k instead.

### Does AIBrix replace vLLM?

No. AIBrix does not ship a model server; it provides the infrastructure around one. The README describes it as building blocks for GenAI inference infrastructure, and the feature list covers routing, autoscaling, LoRA management, a runtime sidecar, distributed inference, KV cache, heterogeneous serving and GPU failure detection, all of which sit alongside an inference engine such as vLLM.

### What does the AIBrix runtime sidecar do?

The README describes the unified AI runtime as a sidecar that handles metric standardization, model downloading and model management. Because it is a sidecar, it attaches to an inference pod rather than running standalone, which is why AIBrix assumes a Kubernetes deployment.

## Sources

- [Official README](https://github.com/vllm-project/aibrix#readme)
- [Project repository](https://github.com/vllm-project/aibrix)
- [Release notes](https://github.com/vllm-project/aibrix/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/vllm-project-aibrix
