Open-source project
chaos-mesh/chaos-mesh avatar
chaos-mesh/chaos-mesh

Chaos Mesh: a Kubernetes-native chaos engineering platform with a dashboard and CRD-driven fault injection

Project brief: A Chaos Engineering Platform for Kubernetes. To lower the threshold for a Chaos Engineering project, Chaos Mesh provides you with a visualization operation.

7,903 stars1,036 forksGoApache-2.0

At a glance

What is it?
Chaos Mesh defines fault injection as Kubernetes custom resources and runs node-level work through a privileged DaemonSet. It is a good fit for teams already fluent in kubectl, and a poor fit for anyone who wants chaos experiments without cluster-wide privileges.
Who is it for?
Adopt Chaos Mesh if your workloads already run on Kubernetes and your team is comfortable writing and reviewing custom resources, because the experiment definition, the scheduler and the health checks all live in the same API you already use. Do not adopt it if you need fault injection outside Kubernetes, if you cannot grant a DaemonSet privileged access to every node, or if you want a hosted service that owns the blast radius for you.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 20 days ago.
What is it written in?
Mainly Go, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Chaos Mesh actually solves for Kubernetes teams

Most teams can describe what should happen when a pod dies. Far fewer can demonstrate it. Chaos Mesh exists to close that gap inside Kubernetes: it lets you declare a fault as a custom resource, apply it through the Kubernetes API, and watch the cluster react. The README describes it as an open source, cloud-native chaos engineering platform that uses Kubernetes custom resources to define, orchestrate, and observe controlled fault injection against workloads, infrastructure, cloud services, and applications.

The audience is narrow and specific. You need a Kubernetes cluster you control, RBAC to create custom resources, and the willingness to grant a DaemonSet privileged access on every node. If your services run on virtual machines with no Kubernetes in front of them, the core mechanism does not apply to you. The README lists physical machine faults among the covered types, but the platform's own framing is Kubernetes-native, and the install path it points to is a Helm chart.

The project is a CNCF incubating project, which is a governance fact rather than a quality claim. The last push to the default branch was on 2026-08-18, and v2.8.4 was released the same day, with v2.8.3 in June 2026 and v2.8.2 in March 2026. That is a steady release rhythm rather than a burst, and the repository is not archived.

Three runtime components and the path a fault takes

Chaos Mesh splits responsibilities across three pieces, and understanding the split explains most of its operational constraints.

Chaos Controller Manager watches Chaos Mesh resources, validates requests through admission webhooks, schedules workflows and experiments, and coordinates injection and recovery. Chaos Daemon runs on Kubernetes nodes as a DaemonSet and performs privileged node- and container-level operations for faults involving runtimes, processes, networks, filesystems, clocks, and kernels. Chaos Dashboard provides the HTTP API and web interface, and the README states it is optional when experiments are managed directly through the Kubernetes API.

The data flow is the standard Kubernetes reconciliation loop. A user creates or updates a Chaos Mesh resource, either through the Kubernetes API directly or through the Dashboard. The Controller Manager reconciles the desired state, and when a fault needs node-level work it delegates to Chaos Daemon. Recovery follows the same path in reverse: the resource is updated or removed, the controller reconciles, and the daemon undoes the injection.

Two design consequences follow. First, the admission webhooks sit in the write path for Chaos Mesh resources, so a webhook that is unavailable affects experiment creation. Second, because the daemon is a DaemonSet, the fault surface is exactly as wide as your node count, and the privileges it holds are the privileges the platform needs. That is not a flaw, but it is the reason security review takes longer than the install.

Installing Chaos Mesh with Helm and running a first experiment

The README points to the Helm installation guide for production and to a separate page for running a chaos experiment. It does not inline the commands, so the exact chart values belong to the chart documentation. The repository does contain a Helm chart under helm/chaos-mesh with its own README.

If you want to see the platform before committing a cluster, the README offers an interactive Killercoda playground: it installs Chaos Mesh on a real 2-node Kubernetes cluster, runs PodChaos and NetworkChaos experiments, walks through the Dashboard, and chains a Workflow, in about 20 minutes. That is the lowest-cost way to judge whether the model fits how your team works.

Once installed, experiments are ordinary Kubernetes objects. The repository ships example manifests under examples/, including burn-cpu.yaml, container-kill-example.yaml, dns-chaos-example.yaml, io-delay-example.yaml, network-delay-example.yaml, network-loss-example.yaml and network-partition-example.yaml. The README's own getting-started path is the run-a-chaos-experiment guide, which walks through creating the resource and observing the result.

What you should see is the resource appear in the cluster, the affected pods change behaviour, and then normal behaviour return once the resource is removed. The README does not document a rollback procedure beyond deleting or updating the resource, so treat cleanup as a normal Kubernetes deletion rather than a platform feature. If the pods do not recover, the problem is usually in the daemon rather than the controller, and the daemon logs are the place to look.

Workflows, schedules and status checks

Single faults are the easy part. The features that make Chaos Mesh more than a fault injector are Schedule, Workflow and StatusCheck, which the README groups under experiment orchestration. Schedule handles recurring experiments. Workflow supports serial or parallel sequences of experiments. StatusCheck lets a workflow consult application health before or during execution.

This is where the platform earns its keep for teams doing release validation. A workflow that injects network delay, waits for a status check, then injects a pod kill, is a repeatable test of a specific failure hypothesis rather than a one-off experiment someone ran during an incident review.

The trade-off is that a workflow is a graph of custom resources, and debugging a workflow that stalls means reading controller logs and resource status rather than a single error message. The README does not describe workflow timeout semantics or what happens to a partially completed workflow when the controller restarts. If you plan to run workflows in CI, verify those behaviours against your own cluster before you depend on them.

Multi-cluster execution and the privilege boundary

The README lists multi-cluster execution as a feature: manage remote clusters and dispatch supported chaos experiments from a management cluster. Note the word supported. Not every fault type necessarily travels across a cluster boundary, and the README defers the details to the documentation rather than enumerating them.

The privilege boundary is the more important limitation. Chaos Daemon performs privileged node- and container-level operations. On managed Kubernetes services, or in clusters where a platform team owns node configuration, that requirement is often the point at which adoption stops. It is not something you can configure away; faults involving kernels, clocks and filesystems have to touch the node.

There is also a scope limitation worth stating plainly. Chaos Mesh is a Kubernetes platform. If your architecture spans Kubernetes and a substantial non-Kubernetes estate, you will end up running a second chaos tool for the other half, and correlating results across the two. The README lists AWS, Azure and GCP faults, which address cloud services, but the orchestration and the daemon still live in the cluster.

Chaos Mesh versus LitmusChaos and Chaos Monkey

The comparison people search for most often is Chaos Mesh versus LitmusChaos, and the difference is in where the experiment definition lives. Chaos Mesh puts it in Kubernetes custom resources reconciled by its own controller and executed by its own daemon, with the Dashboard as an optional front end. LitmusChaos takes a different route, packaging experiments as a hub of reusable chaos experiments that run as Kubernetes jobs, with the experiment logic carried in the job rather than in a reconciler owned by the platform.

The practical consequence is what you debug when something goes wrong. With Chaos Mesh, you read the status of a custom resource and the logs of a controller and a daemon. With a job-based model, you read job pods and their logs. Neither is better in the abstract; the reconciler model gives you schedules, workflows and status checks as first-class resources, while the job model makes each experiment a self-contained unit that is easier to reason about in isolation.

Chaos Monkey is a different kind of comparison. It comes from Netflix and is associated with randomly terminating instances in production to push teams toward resilient design. Chaos Mesh is not random by default and is not aimed at continuous production disruption; it is aimed at controlled, declared experiments with a defined start and a defined recovery. If what you want is background randomness, Chaos Mesh is the wrong shape of tool, and if what you want is a reproducible hypothesis test, Chaos Monkey is.

Licence, upgrade cost and what to check before adopting

Chaos Mesh is licensed under Apache-2.0, which permits commercial use, modification and redistribution, and includes an explicit patent grant. It does not carry the copyleft obligations of a GPL-family licence. This is a description of the licence text, not legal advice; if your organisation has a policy on the CNCF projects it consumes, route the decision through whoever owns that policy.

The upgrade cost is the part teams underestimate. Chaos Mesh installs CRDs, and CRD upgrades are the step that most often goes wrong in Kubernetes operators. The supported releases page is the authoritative source for which Kubernetes versions a given Chaos Mesh release supports, and the release history shows roughly quarterly minor releases with patch releases in between. Pinning the chart version and reading the changelog before each upgrade is cheaper than discovering a CRD schema change during a release window.

Before adopting, verify three things on your own cluster. Confirm the Helm chart's default values match the namespaces and node selectors you intend to target. Confirm your security posture allows a privileged DaemonSet on every node. And run one PodChaos experiment in a disposable namespace to observe recovery, because the README describes the architecture but does not describe a rollback procedure, and the recovery path is the part you will depend on.

Editorial conclusion

Adopt Chaos Mesh if your workloads already run on Kubernetes and your team is comfortable writing and reviewing custom resources, because the experiment definition, the scheduler and the health checks all live in the same API you already use. Do not adopt it if you need fault injection outside Kubernetes, if you cannot grant a DaemonSet privileged access to every node, or if you want a hosted service that owns the blast radius for you. Before you commit, read the supported releases page for your Kubernetes version, confirm the Helm chart's default values match the namespaces you intend to target, and run one PodChaos experiment against a disposable namespace so you can see the recovery path rather than reading about it.

Frequently asked questions

What is Chaos Mesh?

It is an open source, cloud-native chaos engineering platform for Kubernetes that uses custom resources to define, orchestrate and observe controlled fault injection against workloads, infrastructure, cloud services and applications. It is a CNCF incubating project.

Is Chaos Mesh open source and free?

Yes. The repository is licensed under Apache-2.0, which permits commercial use, modification and redistribution. There is no paid tier described in the README.

How do I install Chaos Mesh?

The README points to the production installation guide for Helm, and the repository contains a Helm chart under helm/chaos-mesh with its own README. For a trial without setting up a cluster, the README links an interactive Killercoda playground that installs Chaos Mesh on a real 2-node cluster.

How do I use Chaos Mesh?

You create or update Chaos Mesh resources through the Kubernetes API, either directly or through the Dashboard. The Controller Manager reconciles the desired state and delegates node-level operations to Chaos Daemon. The repository ships example manifests under examples/, including network-delay-example.yaml and container-kill-example.yaml.

How does Chaos Mesh compare with LitmusChaos?

Chaos Mesh defines experiments as Kubernetes custom resources reconciled by its own controller and executed by a DaemonSet, with the Dashboard optional. LitmusChaos packages experiments as reusable units that run as Kubernetes jobs, so what you inspect when an experiment fails differs between the two.

What are the alternatives to Chaos Mesh?

LitmusChaos and Chaos Monkey are the two alternatives the README's material covers. Chaos Monkey is associated with randomly terminating instances in production, while Chaos Mesh is aimed at controlled, declared experiments with a defined start and recovery.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/chaos-mesh-chaos-mesh.svg)](https://hysenlabs.com/projects/chaos-mesh-chaos-mesh)