# Knative Serving: scale-to-zero serverless containers on your own Kubernetes cluster

> Knative Serving is a Kubernetes middleware layer that turns a container image into a request-driven service with autoscaling down to zero. It is for platform teams who want serverless semantics without handing the cluster to a cloud vendor.

**knative/serving** — Kubernetes-based, scale-to-zero, request-driven compute

- Repository: https://github.com/knative/serving
- Website: https://knative.dev/docs/serving/
- Stars: 6,099 · Forks: 1,238
- Language: Go
- License: Apache-2.0
- Published: 2026-09-22 · Updated: 2026-09-22 · Language: en
- Canonical page: https://hysenlabs.com/projects/knative-serving

## What Knative Serving adds to a plain Kubernetes Deployment

A Kubernetes Deployment keeps a fixed number of pods running whether or not anyone is calling them. Knative Serving replaces that model with a request-driven one. The README describes the project as middleware primitives for "deploying and serving of applications and functions as serverless containers", with four named capabilities: rapid deployment of serverless containers, automatic scaling up and down to zero, routing and network programming, and point-in-time snapshots of deployed code and configurations.

The audience is platform and infrastructure engineers who already operate a cluster. If you do not have Kubernetes, nothing here applies to you, because Serving is not a standalone runtime. It installs into a cluster and adds custom resources on top of it. The payoff for that dependency is that an idle service costs no pod capacity, and a busy one scales without you writing a HorizontalPodAutoscaler and a Service and an Ingress for every workload.

## Revisions, Routes and the scale-to-zero autoscaler

The mechanism has three moving parts. A Service is the top-level object you create. Each time its configuration changes, the controller produces a new Revision, which the README calls a "point-in-time snapshot of deployed code and configurations". A Revision is immutable; you do not edit it, you create a new one. A Route then maps incoming host and path traffic onto one or more Revisions, which is what makes gradual rollouts and traffic splitting possible without a separate service mesh.

The autoscaler sits between the request path and the pods. When no requests arrive, it scales the Revision to zero replicas; when traffic returns, it starts pods and the request is held until one is ready. That cold start is the cost of the model, and the README does not quantify it. The networking layer is a separate concern: go.mod pins knative.dev/networking and knative.dev/pkg as dependencies, so the controller image you build is tied to a specific networking revision. Upgrading Serving without matching the networking layer is the most common way an install breaks.

## Installing Knative Serving and deploying a first Service

The repository is the controller implementation, not the install bundle. The README points readers at the serving section of the Knative documentation site for usage, and at the docs folder in this repository for the specification. The repository's own files do not contain install commands, so the place to get them is the documentation site it links to, and the release assets published alongside each knative-vX.Y.Z tag.

Once the controller is running in the knative-serving namespace, a Service is defined with the serving.knative.dev/v1 API group. The repository's sample directory carries its own README with examples, and the specification in docs/ describes the resource shape: a Service whose spec.template.spec.containers list holds the image to run. Applying that manifest is what causes the controller to create a Configuration, a Revision and a Route for you.

Serving does not ship its own ingress. A networking layer has to be installed separately before traffic can reach a Service, and the documentation covers which options are supported. Until that layer is in place, the Service will exist and reconcile, but the Route will not have a working address.

## Where Knative Serving is the wrong tool

Scale-to-zero is a bet that your workload is bursty and tolerant of a cold start. A queue consumer that polls continuously, a WebSocket server holding long-lived connections, or a batch job that runs for an hour will not benefit, and the request-driven model can actively get in the way. The activator has to proxy traffic for a scaled-to-zero Revision, which adds a hop that a permanently warm Deployment does not have.

The second limitation is operational surface. Serving brings its own controller, webhook, autoscaler and activator into your cluster, and go.mod shows the dependency list is large: client-go, apiextensions-apiserver, OpenTelemetry, gRPC and an InfluxDB client all appear as direct requirements. Every one of those is something you now patch. The README also does not document rollback of a Serving upgrade, so a version mismatch between the controller and the networking layer is a situation you have to plan for rather than read about.

## Knative Serving compared with a plain HorizontalPodAutoscaler

The obvious alternative is what Kubernetes already gives you: a Deployment plus an HPA plus a Service and an Ingress. The difference in approach is where the scaling decision lives. An HPA reacts to metrics such as CPU or a custom metric over a polling interval, and it has a minimum replica count of at least one because it cannot scale a Deployment to zero. Knative Serving reacts to request concurrency and can reach zero.

That difference decides the choice. If your service must never have a cold start, an HPA with a floor of two replicas is simpler and has fewer components to maintain. If your traffic is spiky and mostly idle, the HPA's floor is exactly the waste you are trying to remove. A second distinction is revisioning: an HPA rolls a Deployment forward and the old ReplicaSet stops serving. Knative keeps each Revision addressable so a Route can send a percentage of traffic to an older one, which the README lists as a first-class capability.

## Maintenance cadence, licensing and upgrade cost

The repository is not archived, and the last push was on 2026-09-21. Releases are cut on a roughly monthly cadence: knative-v1.23.0 on 2026-07-29, and knative-v1.22.1 and knative-v1.21.3 both on 2026-06-02. The parallel patch releases for 1.22 and 1.21 indicate that older lines still receive fixes, which matters if you cannot move quickly.

The licence is Apache-2.0, declared in the LICENSE file at the repository root. That is a permissive licence with an explicit patent grant, and it imposes no copyleft obligation on the services you build with it. It says nothing about the licences of the container images you deploy, which are your responsibility.

Upgrade cost is the real maintenance line item. Because go.mod pins knative.dev/networking and knative.dev/pkg to dated pseudo-versions, the controller and the networking layer move together. The release notes for each knative-vX.Y.Z tag are the place to check for breaking changes before you apply a new controller manifest over a running install.

## Conclusion

Adopt Knative Serving if you already run Kubernetes and want request-driven scaling, revision snapshots and traffic splitting without a vendor's serverless product. Do not adopt it if you have no cluster to install it on, or if your workloads are long-running batch jobs that never go idle, because scale-to-zero buys you nothing there. Before committing, verify that the networking layer in the current release is compatible with your ingress choice, and check the revision retention setting, since every deploy creates a new Revision that stays addressable until you configure otherwise.

## FAQ

### What is Knative Serving?

It is a Kubernetes-based middleware layer for deploying and serving applications and functions as serverless containers. The README lists rapid deployment, automatic scaling up and down to zero, routing and network programming, and point-in-time snapshots of deployed code and configurations as its capabilities.

### How do I install Knative Serving?

The README points to the serving section of the Knative documentation site for usage instructions, and the repository's docs folder for the specification. The repository itself does not carry install commands, so the documentation site and the release assets for each tag are where the manifests come from.

### Does Knative Serving scale to zero?

Yes. Automatic scaling up and down to zero is one of the capabilities the README lists, and it is what distinguishes the request-driven model from a Kubernetes Deployment with a HorizontalPodAutoscaler, which cannot reach zero replicas.

### What is a Revision in Knative Serving?

The README describes a Revision as a point-in-time snapshot of deployed code and configurations. A Revision is immutable, so changing a Service's configuration produces a new one, and a Route decides how traffic is distributed across the Revisions that exist.

### What licence does Knative Serving use?

Apache-2.0, declared in the LICENSE file at the root of the repository.

## Sources

- [knative/serving on GitHub](https://github.com/knative/serving)
- [License: Apache-2.0](https://github.com/knative/serving/blob/main/LICENSE)
- [Project website](https://knative.dev/docs/serving/)
- [README](https://github.com/knative/serving/blob/main/README.md)
- [Releases](https://github.com/knative/serving/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/knative-serving
