Self-hosted service
kubernetes/kops avatar
kubernetes/kops

kOps: a cluster lifecycle tool for Kubernetes on AWS and GCP

Kubernetes Operations (kOps) - Production Grade k8s Installation, Upgrades and Management

16,684 stars4,720 forksGoApache-2.0

At a glance

What is it?
kOps provisions the cloud infrastructure and the control plane for a Kubernetes cluster, then keeps managing it. It fits teams that want a self-managed cluster on AWS or GCP, and it is the wrong tool if you do not want to own the control plane.
Who is it for?
Adopt kOps if you want a self-managed, highly available Kubernetes cluster on AWS or GCP and you are prepared to own the control plane, the etcd state and the upgrade path. Do not adopt it if a managed control plane is a requirement, or if you need a cloud that is only in alpha support.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Go, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What kOps takes over, and who ends up owning the control plane

kOps is not a client that talks to a cluster somebody else built. The README describes it as "the easiest way to get a production grade Kubernetes cluster up and running" and compares it to `kubectl` for clusters. That comparison sets the boundary. Where `kubectl` operates objects inside a running cluster, kOps operates the cluster itself: creation, destruction, upgrades and ongoing maintenance.

The second half of the job is the part that distinguishes it from a bootstrap script. kOps provisions the necessary cloud infrastructure alongside the cluster. That means the network, the instances and the load balancers are part of the same declarative model as the Kubernetes control plane, so a change to the cluster spec can imply a change to cloud resources rather than only to Kubernetes objects.

This is aimed at platform teams that want a highly available, self-managed cluster and are willing to be responsible for it. Cloud support is tiered, and the tiers matter when you are choosing a target. AWS and GCP are officially supported. DigitalOcean, Hetzner and OpenStack are in beta. Azure is in alpha. If your cloud is in the alpha tier, you are working ahead of the documented support level, and that is a decision about risk rather than a configuration detail. The README states the tiers plainly and does not soften them, which is useful: you can read the support level before you invest in a proof of concept.

The word "production" in the project description carries weight here. A tool that creates a cluster is easy to write. A tool that also owns the upgrade path and the cloud resources underneath has to keep working across versions of two things at once, and that is where the maintenance burden actually lands on the team running it.

How the cluster spec and the state store drive everything

The repository layout shows the split clearly. There is a `cmd/` directory for the CLI entry points, `pkg/` for the shared logic, `upup/` for the cluster creation and update models, `nodeup/` for the node-side bootstrap agent, and `dns-controller/` and `dnsprovider/` for DNS record management. `cloudmock/` exists for testing cloud interactions without real cloud accounts.

The workflow those pieces support is spec-first. You define a cluster, and kOps renders and applies the result. The rendered output is not only Kubernetes manifests. It includes the cloud resources needed to run them, which is why `upup/` sits next to `pkg/` rather than inside it: the update model is a first-class part of the tool, not a script bolted onto the side.

The `nodeup/` directory is worth understanding separately. It is the agent that runs on the node side of the bootstrap, which means the node is not a blank machine waiting for a join command. It receives instructions and configures itself. That design is what lets kOps treat a node group as a declarative unit rather than a list of machines you join one at a time.

The `clusterapi/` directory indicates an integration path with the Cluster API model, which is a different way of expressing the same intent. The `channels/` directory at the top level relates to how kOps resolves which images and versions to use for a given cluster, and it is part of why the release compatibility page exists as separate documentation.

Two consequences follow. First, the cluster spec is the source of truth, so anything you change by hand in the cloud console is drift that the next reconciliation may fight. Second, because the spec is stored and read back, the state store is part of your operational surface. Losing it does not lose the running cluster, but it complicates the next change. The README does not document rollback behaviour for a failed apply, so treat that as something to establish from the docs before you rely on it.

Installing kOps and creating a first cluster

The README does not inline installation commands. It points to the Getting Started page at https://kops.sigs.k8s.io/getting_started/install/, and the repository builds from source through the Makefile. That file exposes several variables you will need to set for a source build, including `S3_BUCKET`, `GCS_LOCATION` and `DOCKER_REGISTRY`.

The Makefile also pins its tooling. `CODEGEN_VERSION=v0.37.0` and `KO_VERSION=v0.18.0` are fixed in the file, and `KO` is invoked as `go run github.com/google/ko@$(KO_VERSION)`. The `go.mod` declares `go 1.27.1`. Read those three numbers before you start, because they define the toolchain the build expects and a mismatch will surface as a build failure rather than a clear error message.

For a source build, the entry point is the default Make target:

bash
make

Once the binary is available, the documented flow is to create a cluster definition, then apply it. The README gives no command examples for this, so follow the Getting Started page for the exact flags rather than guessing at them. That page is also where the cloud-specific prerequisites live, and those differ enough between AWS and GCP that copying a command from one to the other is not safe.

For image builds, the Makefile defines the registry and upload destinations as overridable variables. The defaults are deliberately unusable, so the values have to be supplied when you invoke the target:

bash
DOCKER_REGISTRY?=gcr.io/must-override
S3_BUCKET?=s3://must-override/
GCS_LOCATION?=gs://must-override

Those three lines are the defaults as written in the Makefile. The `?=` form means the value can be overridden from the environment or the command line, so a CI pipeline can inject it without editing the file. The `DOCKER_REGISTRY` default is a guard rail, not a working value: if you see `gcr.io/must-override` in an error, you have not set the variable.

One practical note on the state store: the presence of both `S3_BUCKET` and `GCS_LOCATION` in the Makefile reflects that the store is cloud-specific. Whichever you choose becomes a dependency of your cluster's lifecycle, and it should outlive any individual cluster you create with it.

Where kOps is the wrong tool

The clearest limitation is structural. kOps gives you a self-managed control plane. If your organisation has decided that running etcd and the API server is not its problem, kOps is solving a problem you have chosen not to have. A managed Kubernetes service removes that entire class of operational work, and no amount of tooling in kOps changes the fact that you are the one on call for the control plane.

Cloud tiering is the second constraint. Beta and alpha support are stated as such in the README, and the maturity of the integration follows from that label. A team standardised on a cloud in the alpha tier is running ahead of the documented support level. That does not mean it will fail. It means that when it does, you are the one finding the bug, and the fix may not be in a release you can consume quickly.

The third is drift. Because kOps manages cloud infrastructure as well as Kubernetes objects, a cluster that has been hand-edited in the console is a cluster whose spec no longer describes reality. That is a normal consequence of declarative tooling, but it bites harder here because the blast radius includes load balancers, instances and DNS records, not just pods. A reconciliation that decides to correct a load balancer is a different order of event from a controller restarting a deployment.

Finally, the README does not document rollback. For a tool whose job includes upgrades, that silence is worth taking seriously. Establish your recovery path from the documentation before an upgrade, not during one. If the documentation does not give you one, that is information about the tool, and it belongs in your adoption decision rather than in a post-incident review.

There is also a scope question that has nothing to do with quality. If you need one cluster, once, and you are comfortable assembling it from parts, a smaller tool will get you there with less to learn. kOps earns its keep when the cluster is a long-lived thing you will change repeatedly, and when the cloud resources around it need to change with it.

kOps compared with kubeadm

kubeadm is the natural point of comparison, and the difference is scope rather than quality. kubeadm initialises a control plane and joins nodes to it. It expects you to have already produced the machines, the network and the load balancing. It is a building block, and it assumes the surrounding infrastructure is somebody's job.

kOps takes that surrounding infrastructure as part of its own job. It provisions the cloud resources and the cluster together, which is why the README can describe a production-grade, highly available cluster as a single outcome rather than a sequence of tasks you assemble. The `nodeup/` agent is the mechanism behind that difference: nodes configure themselves from instructions the tool generates, instead of waiting for you to run a join command on each one.

The practical difference shows up at upgrade time. With kubeadm, the upgrade procedure is something you run and sequence yourself, machine by machine. With kOps, the cluster spec and the tool's own upgrade path are the mechanism. That is less work when it goes well and a larger single point of failure when it does not, because the tool is touching cloud resources as well as the control plane. The README's separate Releases and versioning page exists precisely because that upgrade path is coupled to Kubernetes versions, and the coupling is not optional.

A second comparison is with Cluster API, which the repository accommodates through the `clusterapi/` directory. Both express cluster intent declaratively. Cluster API is a broader Kubernetes-native interface across providers, while kOps is a tool with its own model and its own CLI. If your organisation already standardises on Cluster API for other infrastructure, the `clusterapi/` directory is the seam to examine before committing to the CLI workflow.

The honest summary is that kubeadm gives you a smaller tool with a larger manual surface, and kOps gives you a larger tool with a smaller manual surface but more that can move without your direct instruction.

Maintenance, releases and the Apache-2.0 licence

The repository is not archived, and the last push was on 2026-09-17. Recent releases are v1.37.0-beta.1 on 2026-08-29, v1.36.2 on 2026-08-08, and v1.36.1 on 2026-07-28. The presence of a beta release alongside stable patch releases indicates that a new minor line is in testing while the previous line receives fixes.

The README treats release compatibility as its own topic, linking to a Releases and versioning page. That is the right place to look, because kOps and Kubernetes versions are not independent choices: the kOps release you run constrains the Kubernetes versions you can target. Upgrading kOps and upgrading Kubernetes are two decisions, and the documentation treats them as such. A team that upgrades Kubernetes first and kOps second, or the other way round without checking the matrix, is guessing.

The upgrade cost is not only the binary. Because kOps manages cloud infrastructure, an upgrade can imply changes to cloud resources, and those changes are where the risk concentrates. Budget for a staging cluster that mirrors your production spec. A staging cluster that differs in instance types or network layout will not exercise the same cloud changes.

The licence is Apache-2.0. That is a permissive licence with an explicit patent grant. It does not impose copyleft obligations on your own code. This is not legal advice, and if you redistribute kOps or embed it in a product, have counsel read the LICENSE file rather than this paragraph.

Governance is worth noting too. The project sits under the Kubernetes organisation, and the README describes public office hours held every other week with a maintained agenda, open to both developers and users. The README is explicit that bullet form is fine and that items are covered even when they arrive late. For a tool that owns your control plane, that public cadence is a real input when you assess how quickly a problem gets attention.

Editorial conclusion

Adopt kOps if you want a self-managed, highly available Kubernetes cluster on AWS or GCP and you are prepared to own the control plane, the etcd state and the upgrade path. Do not adopt it if a managed control plane is a requirement, or if you need a cloud that is only in alpha support. Before you commit, verify which kOps release matches the Kubernetes version you intend to run, confirm the state store you will use for the cluster spec, and check that your cloud is in the officially supported tier rather than beta or alpha.

Frequently asked questions

What are the key differences between kOps and EKS?

kOps creates and manages a self-managed, highly available Kubernetes cluster and provisions the cloud infrastructure for it. The README frames kOps as the tool for teams that want to run the cluster themselves, including the control plane, rather than consuming a managed control plane.

How do I install kOps?

The README does not inline installation commands. It points to the Getting Started page at https://kops.sigs.k8s.io/getting_started/install/, and the repository can also be built from source through the Makefile, which declares go 1.27.1 in go.mod.

What is Kubernetes kOps?

kOps is Kubernetes Operations, a Go tool that creates, destroys, upgrades and maintains production-grade, highly available Kubernetes clusters and provisions the necessary cloud infrastructure. The README describes it as kubectl for clusters.

Official sources

  1. kubernetes/kops on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/kubernetes-kops.svg)](https://hysenlabs.com/projects/kubernetes-kops)