# kube-monkey: opt-in pod deletion for Kubernetes clusters

> kube-monkey brings Netflix Chaos Monkey style random pod termination to Kubernetes, with dry run as the default and apps opting in through labels. It is for teams that want scheduled failure injection without a full experiment platform.

**asobti/kube-monkey** — An implementation of Netflix's Chaos Monkey for Kubernetes clusters

- Repository: https://github.com/asobti/kube-monkey
- Stars: 3,082 · Forks: 258
- Language: Go
- License: Apache-2.0
- Published: 2026-09-24 · Updated: 2026-09-24 · Language: en
- Canonical page: https://hysenlabs.com/projects/asobti-kube-monkey

## The problem kube-monkey solves, and who it is for

A deployment with three replicas looks resilient on paper. Whether it survives losing one of those replicas at 3am is a different question, and most teams never find out until it happens for real. kube-monkey exists to answer that question on a schedule you control. It is an implementation of Netflix's Chaos Monkey for Kubernetes clusters, and the README states its purpose plainly: it randomly deletes pods in the cluster, "encouraging and validating the development of failure-resilient services."

The audience is narrow on purpose. This is for platform or SRE teams running Kubernetes who want a low-ceremony way to inject pod-level failures into services that have already agreed to be disrupted. It is not a general fault-injection framework, and it is not aimed at teams who want to model network partitions or disk failures. The unit of chaos here is one thing: a pod goes away.

## How kube-monkey decides what to kill

The design is opt-in, which is the most important architectural fact about it. Nothing is a candidate for termination unless its workload carries the kube-monkey/enabled label. An app that has not been labelled is invisible to the scheduler, so the blast radius is bounded by your own labelling decisions rather than by a namespace-wide policy.

Two more labels shape the behaviour. kube-monkey/identifier gives the workload a name that shows up in logs and notifications, and kube-monkey/mtbf sets the mean time between failures, expressed in run days. The README's own example uses a value of "2", which it describes as an app that "expects to lose a pod on about one run day in two." That is a probability per run day, not a fixed interval, so a victim can be spared several days in a row and then hit twice in one week.

The scheduling model separates two timestamps. The documentation distinguishes scheduling time from termination time, and terminations are restricted to the hours and days you configure. That separation matters operationally: the tool decides during one window which pods are marked, and the actual deletion lands inside the configured attack window rather than whenever the process happens to wake up.

The binary itself is a small Go program. The go.mod file shows the dependency set is essentially client-go, viper for configuration, fsnotify, glog and the Prometheus client library. There is no custom resource definition in the top-level layout, and no operator loop. Configuration arrives through a ConfigMap, which is why the examples directory contains files like examples/configmap.yaml, examples/custom-resources-configmap.yaml, examples/debug-mode-configmap.yaml, examples/metrics-configmap.yaml and examples/notifications-configmap.yaml. Metrics are exposed on a Prometheus endpoint, and notifications can be posted to Slack or to an API you run yourself.

## Installing kube-monkey and opting in your first app

The README's quick start uses the project's Helm repository. The chart installs into kube-system in the example, and the command sequence is three lines.

```bash
helm repo add kubemonkey https://asobti.github.io/kube-monkey/charts/repo
helm repo update
helm install kube-monkey kubemonkey/kube-monkey --namespace kube-system
```

After the install completes, kube-monkey is running in dry run mode. The README is explicit that dry run is the default, so nothing dies until you say so. If you want to confirm the deployment exists before going further, kubectl get pods -n kube-system will show the kube-monkey pod alongside the other system components.

Opting an app in is a label change on the workload's pod template. The README gives this exact set:

```yaml
metadata:
  labels:
    kube-monkey/enabled: enabled
    kube-monkey/identifier: monkey-victim
    kube-monkey/mtbf: "2"
```

With those three labels applied, the workload becomes a candidate. The mtbf value of "2" means roughly one run day in two. The identifier monkey-victim is what you will see in the tool's output. The README points to the Getting started page for the walk through, including how to watch a real termination before you trust it with a live namespace, which is the step worth taking before you disable dry run on anything that serves traffic.

## Where kube-monkey is the wrong tool

The limitation is in the name of the action. kube-monkey deletes pods. It does not inject latency, corrupt packets, fill disks, exhaust CPU, or fail a dependency. If your service degrades because a downstream API gets slow, kube-monkey will never reproduce that, because a deleted pod is a binary event with a clean recovery path through the ReplicaSet controller.

That matters more than it first appears. Pod deletion is the easiest failure mode for Kubernetes to absorb. The control plane notices the missing pod and starts a replacement, and for many stateless services the user impact is a few seconds of reduced capacity. A tool that only does this will confirm that your readiness probes and replica counts work, and will tell you almost nothing about how your service behaves when a dependency is unreachable.

The second constraint is the opt-in model itself. Because nothing is a candidate unless it is labelled, kube-monkey cannot tell you that you forgot to label something. An unlabelled deployment is silently exempt, and a team that labels three of its twenty services has a chaos program covering fifteen percent of its surface area while looking, in dashboards, fully deployed. The blast radius is bounded by your labelling hygiene, and so is the value.

There is also the question of state. A pod that owns a PersistentVolumeClaim is not a stateless replica. Deleting it may be fine, or it may leave a volume attached to a node that no longer runs anything, depending on the storage driver. The README does not document a guard against this, so the labelling decision carries that weight.

## kube-monkey compared with Chaos Mesh

Chaos Mesh is the alternative that shows up most often in searches around Kubernetes chaos tooling, and the difference in approach is structural rather than a matter of feature count. Chaos Mesh is a full experiment platform built around custom resources: you declare a chaos experiment as a Kubernetes object, and a controller reconciles it. It covers pod kill, network faults, I/O faults, stress and time skew, and it can target by selector across namespaces.

kube-monkey takes the opposite path. There is no experiment object and no controller to reconcile one. Configuration lives in a ConfigMap, victims are selected by label, and the timing model is a run day with an attack window rather than an explicit experiment lifecycle. That means less to learn and less to run, and it also means less control. You cannot express "kill exactly two of these pods and then stop" as a declarative object, and you cannot compose a network fault with a pod kill in a single experiment.

For a team whose only unanswered question is whether a replica loss is survivable, the smaller tool is the better fit, because the operational cost of the platform is real. For a team that needs to test partial network failure between two services, Chaos Mesh is the tool that can express it and kube-monkey is not. Choosing kube-monkey is choosing to answer one question well rather than many questions adequately.

## Maintenance, upgrades and the Apache-2.0 licence

The repository is not archived, and the last push was on 2026-09-20. Three releases landed that same day: v0.6.0, v0.6.1 and v0.7.0, with v0.7.0 published at 18:00 UTC. That is an active release cadence, and it is recent enough that the project is not coasting.

The upgrade story has a specific wrinkle that comes from the deployment model. Because configuration is delivered through a ConfigMap and the examples directory ships several variants (base, custom resources, debug mode, metrics, notifications), an upgrade is not only a new container image. The binary is built from Go with a pinned toolchain, and the Dockerfile shows the image is built FROM scratch with only the CA certificates and zoneinfo copied in. A scratch image means no shell, so debugging a running pod in place is not an option; you read logs and metrics instead.

On licensing, kube-monkey is Apache-2.0, and the README points to the LICENSE file for details. Apache-2.0 is a permissive licence that includes an explicit patent grant, which is generally the reason organisations prefer it over MIT for infrastructure components. Nothing in the repository indicates a dual-licence arrangement or a commercial tier. This is a description of what the licence file says, not legal advice; if your organisation has a policy on which licences are acceptable for infrastructure tooling, that policy is the thing to check.

## Conclusion

Adopt kube-monkey if you already run Kubernetes workloads that declare their own disruption tolerance and you want a small, label-driven killer rather than a full experiment framework. Do not adopt it if you need probes, network faults or dependency-level failures, because the project only deletes pods. Before trusting it with a live namespace, verify three things: that the default dry run setting is still in effect after your Helm install, that your victim deployments actually carry the kube-monkey/enabled label, and that your PodDisruptionBudgets and replica counts can absorb a single pod loss during the configured run window.

## FAQ

### Does kube-monkey delete pods by default when I install it?

No. The README states that dry run is the default, so nothing dies until you explicitly turn it off. You should confirm the setting after your Helm install and watch a termination in dry run before trusting it with a live namespace.

### How do I opt an application in to kube-monkey?

Add the kube-monkey/enabled label set to enabled on the workload, along with kube-monkey/identifier and kube-monkey/mtbf. The README's example uses an identifier of monkey-victim and an mtbf of "2", which it describes as losing a pod on about one run day in two.

### Can kube-monkey inject network latency or other faults?

No. The project randomly deletes pods and nothing else, so latency, packet loss, I/O faults and dependency failures are outside what it does. A platform such as Chaos Mesh is the alternative when you need those fault types.

### Does Netflix use Chaos Monkey?

The README describes kube-monkey as an implementation of Netflix's Chaos Monkey for Kubernetes clusters and links to the original Netflix project. It does not document how Netflix itself runs Chaos Monkey today.

## Sources

- [asobti/kube-monkey on GitHub](https://github.com/asobti/kube-monkey)
- [Issues](https://github.com/asobti/kube-monkey/issues)
- [License: Apache-2.0](https://github.com/asobti/kube-monkey/blob/master/LICENSE)
- [README](https://github.com/asobti/kube-monkey/blob/master/README.md)
- [Releases](https://github.com/asobti/kube-monkey/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/asobti-kube-monkey
