# KubeAI: an inference operator that puts vLLM and Ollama behind one OpenAI-compatible endpoint

> KubeAI is a Kubernetes operator plus proxy that manages vLLM, Ollama, Infinity and FasterWhisper backends, scales them from zero, and routes requests with prefix-aware load balancing. It is Apache-2.0 and aims to work without Istio, Knative or the Prometheus adapter.

**kubeai-project/kubeai** — AI Inference Operator for Kubernetes. The easiest way to serve ML models in production. Supports VLMs, LLMs, embeddings, and speech-to-text.

- Repository: https://github.com/kubeai-project/kubeai
- Website: https://www.kubeai.org
- Stars: 1,270 · Forks: 137
- Language: Go
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/kubeai-project-kubeai

## The problem KubeAI solves: vLLM is stateful, and Kubernetes load balancing does not know that

A plain Kubernetes Service in front of several vLLM replicas spreads connections with kube-proxy's default strategy. The README argues this performs poorly on time-to-first-token and throughput, because vLLM is not stateless: its speed depends heavily on the state of its KV cache. Two replicas can hold completely different prefix caches, and a random pick throws away the work one of them already did.

KubeAI's answer is a proxy that sits in front of the serving engines and applies a prefix-aware load balancing strategy, so requests that share a prefix tend to reach a replica that already has that prefix cached. The project points to a paper in its blog directory for the measurements rather than quoting a single number here. Treat the claim as directional until you reproduce it on your own traffic shape.

The second problem is operational. Standing up vLLM in production usually means writing Deployments, Services, PVCs, model download jobs and autoscaling glue yourself. KubeAI replaces that with a Model custom resource that the operator turns into Pods, volumes and downloads. It is aimed at platform teams running Kubernetes who want to serve text generation, embeddings, reranking or transcription without adopting a larger serving stack.

## Two components, one Deployment: the proxy and the model operator

The README describes two sub-components. The model proxy exposes an OpenAI-compatible API and implements prefix-aware load balancing, request queueing while the system scales from zero, and retries against bad backends. The model operator manages backend server Pods directly, automating model downloads, volume mounting and dynamic LoRA adapter loading through the Model CRD.

Both run co-located in the same Deployment. The README notes they could be deployed independently, linking to issue 430, which is an open design discussion rather than a documented deployment mode. If you were hoping to run only the proxy against externally managed vLLM servers, that split is not something the documentation currently promises.

The proxy exposes these endpoints according to the README: /v1/chat/completions, /v1/completions, /v1/embeddings, /v1/rerank, /v1/models and /v1/audio/transcriptions. That list is the practical contract. Anything an OpenAI client library sends outside it will not be handled.

Backends are chosen per model. The topics list names vLLM, Ollama, FasterWhisper and Infinity, and the model catalog entries carry an engine field. The operator, not the proxy, is what creates and mounts the storage for cached weights, so a model that fails to download shows up as a Pod that never becomes ready rather than as a proxy error.

## Installing KubeAI with Helm and serving your first model

The quickstart assumes a local cluster. The README suggests kind or minikube, and includes a tip for Podman users whose machine defaults to 2G of memory: stop and remove the existing machine, then re-init with a larger memory and disk size before starting it.

```bash
kind create cluster # OR: minikube start
```

Add the Helm repository and install the chart. The --wait flag with a 10 minute timeout means the install blocks until components are ready; the README warns this may take a minute.

```bash
helm repo add kubeai https://www.kubeai.org
helm repo update
helm install kubeai kubeai/kubeai --wait --timeout 10m
```

Models come from a separate chart, kubeai/models, driven by a values file. The README's example enables a DeepSeek R1 1.5B CPU model backed by Ollama, plus Qwen2 500M and Nomic Embed Text, all on CPU resource profiles.

```yaml
catalog:
  deepseek-r1-1.5b-cpu:
    enabled: true
    features: [TextGeneration]
    url: 'ollama://deepseek-r1:1.5b'
    engine: OLlama
    minReplicas: 1
    resourceProfile: 'cpu:1'
  qwen2-500m-cpu:
    enabled: true
  nomic-embed-text-cpu:
    enabled: true
```

Install that file into the cluster with the models chart, then confirm the endpoint is live by asking the proxy what it serves. The README does not spell out the port-forward command, so check the chart's Service name in your release before assuming one.

```bash
helm install kubeai-models kubeai/models -f ./kubeai-models.yaml
curl http://localhost:8000/v1/models
```

If /v1/models lists your catalog entries, the operator reconciled the Model resources and the proxy can see them. If it returns an empty list, the models chart was installed before the operator was ready, or the catalog keys do not match what you enabled.

## Where KubeAI gets in the way: statefulness, split deployments and hardware assumptions

The prefix-aware routing that gives KubeAI its reason to exist also makes the proxy a required hop. You cannot point clients at the vLLM Service directly and keep the benefit, and the README does not document a mode where the proxy runs without the operator.

Scale-from-zero has a cost the README acknowledges indirectly: the proxy implements request queueing while the system scales from zero replicas. A queued request is still a request your client is waiting on. For interactive chat, a cold model means the first user after an idle period pays the full model load, which for a large checkpoint is minutes, not seconds. The catalog's minReplicas field exists precisely to avoid that, at the price of keeping a GPU warm.

Model caching is the other sharp edge. The README says the operator automates downloading and mounting, naming EFS as an example. That implies your cluster needs a storage class the operator can use, and the examples directory has a storage-classes folder, which suggests this is a common place to get stuck. A slow or misconfigured volume turns every scale-up into a re-download.

The hardware story is CPU, GPU or TPU, but the pre-configured catalog is described as tuned for common GPU types. Running a 1.5B model on cpu:1 works for a smoke test; it is not a substitute for sizing a production deployment, and the README does not publish throughput figures for CPU profiles.

## KubeAI compared with KServe and with running vLLM directly

The most direct comparison in the search data is KubeAI against KServe. The architectural difference is dependency count. KubeAI's README states it does not require Istio or Knative for scale-from-zero, nor the Prometheus metrics adapter for autoscaling, and calls out day-two operations as the reason: fewer inter-project version and configuration mismatches. KServe's documented installation path involves a larger set of components, which is the trade you are making. If your platform already runs Istio and Knative for other services, that argument weakens considerably, and KServe's broader model-format support may matter more than a smaller dependency graph.

The comparison with running vLLM on Kubernetes by hand is about what you give up in control. A hand-written Deployment lets you set every vLLM flag exactly as you want. KubeAI's catalog hides those flags behind resource profiles and engine settings, which the README frames as saving time on vLLM-specific tuning. The flip side is that an unusual flag combination may not be expressible, and you would be fighting the abstraction.

Against Ollama alone, KubeAI is a different layer. Ollama is a serving engine; KubeAI operates Ollama servers as one of its engines, alongside vLLM, and adds routing, queueing and autoscaling around them. If you only ever run one model on one machine, Ollama by itself is less machinery.

## Licence, upgrade surface and what a version bump actually touches

KubeAI is Apache-2.0, which permits commercial use and modification with the usual attribution and notice requirements. That is a permissive licence, and nothing in the repository suggests a dual-licensing or open-core split; the operator, proxy and charts all live in the same repository. This is a description of the licence file, not legal advice.

The upgrade surface is wider than the version number suggests. There are two charts, helm-chart-kubeai and helm-chart-models, released together at 0.23.4 in July 2026, with the application release v0.23.3 slightly earlier. Upgrading the operator without the models chart, or the reverse, leaves you running mismatched catalog definitions against a newer CRD. Check both chart versions before a helm upgrade.

Because the operator creates Pods and volumes on your behalf, a CRD change can alter what gets created without any change to your values file. The repository has a CHANGELOG.md and a proposals directory; the proposals are the place to look when a behaviour change is planned rather than shipped. There is no documented rollback procedure in the README, so plan upgrades around a cluster you can rebuild.

Finally, the operator needs permissions to manage Pods and PersistentVolumeClaims cluster-wide. That is a real security boundary, and the README does not describe a namespace-scoped installation mode.

## Conclusion

Adopt KubeAI if you already run Kubernetes and want vLLM or Ollama behind one OpenAI-compatible endpoint without installing Istio, Knative or the Prometheus adapter. Do not adopt it if you need a multi-framework serving layer with a Python SDK, or if you cannot give the operator cluster-scoped access to create Pods and PersistentVolumeClaims. Before rolling it out, install the Helm chart on a throwaway cluster, confirm the /v1/models response lists your catalog entries, and check what your cluster does when a node is drained while a model pod is still loading weights.

## FAQ

### Is kube the same as Kubernetes?

No. Kube is only a prefix used in tool names such as kubeadm and kubectl; Kubernetes is the cluster system itself. KubeAI is an operator that runs on Kubernetes and manages model serving Pods, so it is a workload on top of the cluster, not the cluster.

### Which AI is best for Kubernetes?

No ranking of AI systems for Kubernetes is available here. What can be said is that KubeAI serves LLMs, embeddings, reranking and speech-to-text by operating vLLM, Ollama, Infinity and FasterWhisper backends, and that it does not require Istio, Knative or the Prometheus metrics adapter.

### What is kubeadm used for?

kubeadm is a cluster bootstrap tool. It is unrelated to KubeAI, which assumes you already have a working cluster, such as one created with kind or minikube, and then installs its operator and proxy into it.

### What exactly is Kubernetes used for?

Kubernetes is the container orchestration system KubeAI runs on. KubeAI uses it to schedule model server Pods, mount cached weights from volumes, scale replicas from zero, and expose an OpenAI-compatible API through its proxy.

## Sources

- [kubeai-project/kubeai on GitHub](https://github.com/kubeai-project/kubeai)
- [License: Apache-2.0](https://github.com/kubeai-project/kubeai/blob/main/LICENSE)
- [Project website](https://www.kubeai.org)
- [README](https://github.com/kubeai-project/kubeai/blob/main/README.md)
- [Releases](https://github.com/kubeai-project/kubeai/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/kubeai-project-kubeai
