# KServe: a Kubernetes control plane for both LLM and predictive model serving

> KServe puts generative and classic model inference behind one Kubernetes CRD, the InferenceService. The design pays off in multi-tenant clusters and costs you a heavier control plane than a plain Deployment.

**kserve/kserve** — Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes

- Repository: https://github.com/kserve/kserve
- Website: https://kserve.github.io/website/
- Stars: 5,952 · Forks: 1,679
- Language: Go
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/kserve-kserve

## The problem KServe solves, and who ends up using it

Serving a model on Kubernetes is easy once. Serving forty of them, in several frameworks, with different scaling rules and a shared ingress, is not. Each team ends up writing its own Deployment, Service, HPA and routing config, and the platform team inherits a pile of YAML nobody can review.

KServe's answer is a single custom resource, the InferenceService, that describes a model and its serving requirements. The description is framework-neutral: the README lists TensorFlow, PyTorch, scikit-learn, XGBoost and ONNX under predictive AI, and vLLM plus llm-d under generative AI. The controller turns that one object into the underlying Kubernetes resources.

The audience is narrower than the tagline suggests. You need a Kubernetes cluster you control, permission to install cluster-scoped CRDs, and someone who understands ingress and service meshes. Data scientists who only want to call a hosted endpoint are not the target. Platform engineers who are tired of hand-rolled inference Deployments are.

## How the InferenceService becomes running pods

The repository is a Go operator. The Dockerfile builds ./cmd/manager into a static binary and copies it into a distroless image, so the controller itself carries no shell and no package manager. That binary watches InferenceService objects and reconciles the rest.

The README describes a split between a predictor, an optional transformer and an optional explainer, with request routing between them and automatic traffic management. That is the core data flow: a request arrives at the router, the transformer can preprocess or postprocess, and the predictor runs the model. The explainer is a parallel path for feature attribution, not part of the hot path.

Where the pods actually come from depends on the installation mode. With Knative, an InferenceService maps to a Knative Service, which is what makes request-based autoscaling and scale-to-zero possible. On the standard Kubernetes path the README says the install is more lightweight but does not support canary deployment and request-based autoscaling with scale-to-zero. Those two sentences describe the real architectural fork in the project, and choosing wrong is expensive to undo.

For generative workloads the router also fronts OpenAI-compatible endpoints, which is why vLLM and llm-d appear as backends rather than as replacements for the controller.

## Installing KServe and deploying a first InferenceService

The README does not inline install commands. It points at the KServe website for four standalone options (standard Kubernetes, Knative, ModelMesh and a local quick install) and at the Kubeflow documentation for the addon path. The repository ships charts/, config/ and install/ directories, which is where the manifests live, but the README treats the website as the source of truth. Follow that page rather than assembling manifests by hand.

The serverless path is the default one, so the Knative installation guide is the normal starting point. Once the control plane is up, an InferenceService is a small YAML object. The README links a dedicated page for creating the first one. The repository's own manifests under config/ are the only place in the repository where the object's shape can be read, so treat the website's InferenceService API reference as the field-level authority; the top-level structure the README describes is a predictor with a model, an optional transformer and an optional explainer.

For an LLM, the same object changes backend. The README lists vLLM and llm-d as optimized backends and describes an OpenAI-compatible protocol, so a generative InferenceService is reached with an OpenAI-style request rather than a KServe-specific one. The exact field names for the generative path are on the website, not in the README.

If the object stays not ready, the failure is usually the model storage location: the controller needs credentials to read that bucket, and the repository's go.mod lists SDKs for AWS S3, Azure Blob and Google Cloud Storage, so those are the storage backends the code is built against.

## Where KServe is the wrong tool

The clearest limitation is stated by the project itself: the lightweight standard Kubernetes install gives up canary deployment and request-based autoscaling with scale-to-zero. If those two features are why you looked at KServe, the simple install is not an option, and you are taking on Knative.

Knative is a large dependency. It brings its own networking and revision model, and it interacts with whatever ingress or service mesh you already run. The topics list includes istio and knative, and go.mod pins istio.io/api, so an Istio-based mesh is a supported configuration rather than an accident. In a cluster with an existing, differently-opinionated ingress stack, that is a real integration project, not a single apply.

There is also a scale mismatch. A single model with steady traffic behind one Deployment and one Service needs none of this. A controller, CRDs, a router and a Knative install are overhead you will pay for in debugging time the first time a revision fails to become ready.

Finally, the README is a signpost, not a manual. Every substantial instruction is a link to the website. Offline or air-gapped installation is not described in the README at all, and neither is rollback. If your environment cannot reach the website, budget time to read charts/ and config/ directly.

## KServe compared with Ray Serve and plain vLLM

Ray Serve is the closest alternative in spirit. It is a Python serving library: you write a deployment class, and Ray handles replicas and routing inside a Ray cluster. KServe inverts that. The unit of configuration is a Kubernetes object, not Python code, and the scaling machinery is Kubernetes-native. If your team already thinks in Kubernetes and wants GitOps-friendly declarative model definitions, KServe fits better. If your team is Python-first and would rather not touch CRDs, Ray Serve asks less of you.

vLLM is not really a competitor, though the search data treats it as one. vLLM is an inference engine; KServe is the control plane that can deploy it. The README lists vLLM as an optimized backend, and the same applies to llm-d. Choosing between them is a question of what runs inside the pod, not of which platform to adopt.

Against Kubeflow the relationship is different again. The README states that KServe is an addon component of Kubeflow, so this is a component comparison rather than an either-or. Kubeflow covers pipelines and training; KServe covers the serving end.

## Maintenance, releases and the Apache-2.0 licence

The repository is not archived, and the last push was on 2026-09-10. The most recent release is v0.20.0, published on 2026-08-06, preceded by two release candidates in July and early August. That cadence, a tagged release roughly monthly with candidates before it, is visible in the release list and is the main signal about how fast the surface moves.

Upgrade cost is the part to plan for. KServe installs cluster-scoped CRDs, and CRD schema changes are not automatically applied by a controller upgrade, so the install path matters on every version bump. The Makefile exposes CRD generation options and the repository carries kserve-deps.env and kserve-images.env to pin dependency and image versions, which is the mechanism the project uses to keep a release coherent. Read RELEASES.md and the v0.20.0 notes before upgrading a production cluster.

The licence is Apache-2.0, which is permissive and includes an explicit patent grant. The Dockerfile runs go-licenses save and copies the collected third-party licence texts into /third_party in the image, so the distributed container carries its dependency licences with it. That is a convenience, not a legal opinion: if you redistribute a modified KServe image, review the notices yourself.

## What the KServe documentation does not settle

Two things are worth checking before you commit, because the README leaves them open.

First, GPU support is listed as a feature with no detail on how GPU resources are requested or scheduled. The generative features mention GPU acceleration and optimized memory management for large models, plus KV cache offloading to CPU or disk, but the README does not say which node labels or resource requests you configure. That is in the website documentation.

Second, model caching is described as reducing loading times for frequently used models. The repository has a localmodel-agent.Dockerfile and a localmodel.Dockerfile at the top level, which suggests a node-local caching component, but the README does not explain how it is enabled or where the cache lives on disk. If cold-start latency drives your design, confirm the caching path before assuming it applies to your workload.

## Conclusion

Adopt KServe if you already run Kubernetes and need one API for both predictive models and LLM backends such as vLLM, with canary rollouts and scale-to-zero. Do not adopt it if a single model behind a Deployment is enough, because the serverless path pulls in Knative and the raw path drops autoscaling and canary support. Before committing, verify which installation mode your cluster can support, check that the vLLM or llm-d backend you intend to use is available for your GPU nodes, and read the v0.20.0 release notes for upgrade steps from your current version.

## FAQ

### What is KServe used for?

It is a platform for deploying and serving AI models on Kubernetes. The README describes it as covering both generative AI, with vLLM and llm-d backends and an OpenAI-compatible protocol, and predictive AI, with TensorFlow, PyTorch, scikit-learn, XGBoost and ONNX support.

### Is KServe part of Kubeflow?

Yes, in the sense that the README calls KServe an important addon component of Kubeflow and links to the Kubeflow KServe documentation. It is also a CNCF incubating project in its own right, with its own releases and website.

### Is KServe open source?

Yes. The repository is public under the Apache-2.0 licence, and the README links to the LICENSE file. The Dockerfile also bundles third-party licence texts into the image via go-licenses.

### How do I install KServe?

The README does not list install commands inline. It points to the KServe website for standalone options: standard Kubernetes, Knative, ModelMesh and a local quick install, plus Kubeflow installation guides for AWS and OpenShift Container Platform.

### How do I use KServe?

You install the control plane, then create an InferenceService object describing the model format, the storage location and an optional transformer or explainer. The README links a dedicated page for creating your first InferenceService.

### What is the difference between KServe and vLLM?

They operate at different layers. vLLM is an inference engine that KServe lists as an optimized backend for serving LLMs, while KServe is the Kubernetes control plane that deploys and routes to it.

## Sources

- [kserve/kserve on GitHub](https://github.com/kserve/kserve)
- [License: Apache-2.0](https://github.com/kserve/kserve/blob/master/LICENSE)
- [Project website](https://kserve.github.io/website/)
- [README](https://github.com/kserve/kserve/blob/master/README.md)
- [Releases](https://github.com/kserve/kserve/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/kserve-kserve
