# LitmusChaos: chaos engineering for Kubernetes workloads, and who should not adopt it

> LitmusChaos is a CNCF chaos engineering platform that runs as Kubernetes microservices and defines faults through custom resources. It fits teams already operating Kubernetes who want controlled fault injection in pipelines; it is the wrong tool for anyone without a cluster to break.

**litmuschaos/litmus** — Litmus helps  SREs and developers practice chaos engineering in a Cloud-native way. Chaos experiments are published at the ChaosHub  (https://hub.litmuschaos.io). Community notes is at https://hackmd.io/a4Zu_sH4TZGeih-xCimi3Q

- Repository: https://github.com/litmuschaos/litmus
- Website: https://litmuschaos.io
- Stars: 5,624 · Forks: 899
- Language: Go
- License: Apache-2.0
- Published: 2026-09-22 · Updated: 2026-09-22 · Language: en
- Canonical page: https://hysenlabs.com/projects/litmuschaos-litmus

## The problem LitmusChaos addresses for Kubernetes operators

Most reliability work happens before production: unit tests, integration tests, load tests. None of them answer the question a Kubernetes operator actually loses sleep over, which is what happens to this deployment when a node dies, a pod is killed, or a dependency becomes slow. LitmusChaos exists to induce those failures on purpose, in a controlled way, against a named workload. The README frames the goal as identifying weaknesses and potential outages in infrastructures by inducing chaos tests in a controlled way.

The audience is stated plainly. Developers run experiments as an extension of unit or integration testing. CI/CD pipeline builders run chaos as a pipeline stage to find bugs when an application is subjected to fail paths. SREs plan and schedule experiments against the application and the infrastructure around it. All three groups share one precondition: they run Kubernetes. LitmusChaos is not a general fault injection library you link into an application. It is a platform that lives in a cluster and manipulates other things in that cluster.

## Control plane, execution plane and the three custom resources

The architecture splits into two halves. The chaos control plane is a centralized management tool called chaos-center, which the README describes as the place to construct, schedule and visualize Litmus chaos workflows. The chaos execution plane is a chaos agent plus multiple operators that execute and monitor experiments inside a target Kubernetes environment. That split matters operationally: the control plane can sit in one cluster or namespace while the execution plane reaches into another.

The resources carry the semantics. A ChaosExperiment groups the configuration parameters of one fault; the README calls these installable templates that describe the library carrying out the fault, the permissions needed to run it, and the defaults it operates with. A ChaosEngine links a workload, node or infra component to a fault described by a ChaosExperiment, and it is where you tune run properties and specify steady state validation constraints using probes. The Chaos-Operator watches ChaosEngine objects and reconciles them by triggering execution through runners. A ChaosResult holds the outcome: whether each validation constraint succeeded, the revert or rollback status of the fault, and a verdict. The chaos-exporter reads ChaosResults and exposes them as Prometheus metrics.

ChaosExperiment and ChaosEngine are embedded inside a Workflow object, which strings one or more experiments together in a chosen order. That is the difference between a single fault and a scenario: kill a pod, wait, confirm recovery, then degrade the network path.

## Installing LitmusChaos and running a first experiment

The repository README does not contain installation commands. It points to the Litmus Docs and specifically to the Installation section of the Getting Started with Litmus page, which lists prerequisites. Treat that page as the source of truth; anything below describes the shape of the workflow rather than a copy of the docs.

The platform installs into a Kubernetes cluster, and the operator image is published as litmuschaos/chaos-operator on Docker Hub, which the README's badge links to. Once the operator and control plane are running, the usual path is to install an experiment from the ChaosHub at hub.litmuschaos.io, which hosts ChaosExperiment CRs that application developers and vendors share so their users can increase the resilience of their applications in production.

The workflow is driven by custom resources rather than a command line. The README names the ChaosExperiment, ChaosEngine and ChaosResult CRs and the Workflow object that embeds the first two. What the README does not give is a manifest to copy, so read the docs page for the experiment you pick before writing one.

After the Chaos-Operator reconciles a ChaosEngine, a runner executes the fault and a ChaosResult appears in the same namespace. Inspecting that object is the first real use: it tells you whether the probe constraints held and whether the fault was reverted. In automated runs the ChaosResult is the artifact you read, since the exporter turns it into Prometheus metrics rather than requiring a human to watch the cluster.

## Where LitmusChaos is the wrong tool

The most direct limitation is the one the README implies rather than states: everything runs on Kubernetes. If your production system is a single VM, a managed database, or a serverless function outside a cluster you control, LitmusChaos has nothing to attach to. The ChaosEngine targets a Kubernetes workload, node or infra component. There is no documented path for injecting faults into a system that is not represented in the cluster.

A second constraint is the blast radius of the platform itself. The execution plane runs inside the environment it is breaking. An operator that kills pods, deletes containers or degrades network paths needs permissions to do exactly that, and the README notes that ChaosExperiment CRs indicate the permissions needed to run them. Granting those permissions in a production cluster is a real decision, not a formality.

Third, the repository separates experiments by maturity. CHAOS_EXPERIMENT_MATURITY.md exists at the top level, which tells you the project distinguishes experiments that are ready from those that are not. Picking an experiment for a production pipeline without reading that file is a mistake the repository is trying to prevent. Finally, the README does not document rollback semantics beyond the revert status recorded in ChaosResult. If your fault cannot be safely reverted, the result object will tell you what happened, but the README does not promise that every fault undoes itself cleanly.

## How LitmusChaos differs from Gremlin and from writing your own fault scripts

Gremlin is the commercial alternative most teams compare against. The difference in approach is where the fault logic lives. LitmusChaos defines faults as Kubernetes custom resources that you install from a public hub, version alongside your manifests, and reconcile with an operator; the experiment definition is an object in your cluster. Gremlin runs as an agent you install on hosts and drive from a hosted control plane, with faults configured in that external service. If your team already treats cluster state as the source of truth and wants chaos definitions reviewed in the same pull request as the deployment, the custom resource model fits. If your estate is mostly virtual machines and you want a vendor to manage the fault library, the agent model fits better.

The other alternative is the one most teams actually start with: a script that kills pods on a schedule. That works until you need a steady state hypothesis. LitmusChaos makes the hypothesis explicit through probes attached to the ChaosEngine, and records the verdict in ChaosResult. A shell script has no equivalent object, so there is nothing to assert against in a pipeline and nothing to export as a metric.

## Release cadence, maintenance and licence

The project is not archived, and the last push to master was on 2026-09-22. Releases are frequent and versioned: 3.32.0 on 2026-09-17, 3.31.0 on 2026-07-15, and 3.30.0 on 2026-06-17. Three minor releases in roughly three months means upgrade cost is a recurring line item, not a one-off. The upgrade path itself is not documented in the README; RELEASE_GUIDELINES.md at the top level is where the project describes how releases are cut, and the Litmus Docs are the place to look for upgrade steps for the control plane and the execution plane. Budget for the fact that the operator, the runners and the ChaosExperiment CRs can move independently.

The licence is Apache-2.0, which permits commercial use and modification with the usual attribution and notice obligations; the repository ships LICENSE and NOTICE.md for that purpose. Apache-2.0 does not include support. COMMERCIAL_SUPPORT.md exists at the top level, which tells you the project draws a line between community use and paid support. Nothing here is legal advice, and if you redistribute LitmusChaos inside a product, read LICENSE and NOTICE.md yourself.

## Conclusion

Adopt LitmusChaos if you already run Kubernetes and want fault injection expressed as custom resources you can schedule, probe and export as Prometheus metrics; skip it if you have no cluster, no steady state hypothesis, or no appetite for running an operator that deliberately breaks workloads. Before installing, read the Installation section of the Litmus Docs rather than the repository README, check the maturity level listed for your chosen experiment in CHAOS_EXPERIMENT_MATURITY.md, and confirm COMMERCIAL_SUPPORT.md covers the support path you need.

## FAQ

### How do I install LitmusChaos on Kubernetes?

The repository README does not give install commands. It directs readers to the Litmus Docs and specifically to the Installation section of the Getting Started with Litmus page, which lists the prerequisites.

### What is LitmusChaos?

It is an open source chaos engineering platform, and a CNCF project, that lets teams induce chaos tests in a controlled way to find weaknesses in infrastructure. It runs as microservices on Kubernetes and defines faults through custom resources.

### What does a ChaosEngine do in LitmusChaos?

A ChaosEngine links a Kubernetes workload, node or infra component to a fault described by a ChaosExperiment, and lets you tune run properties and set steady state validation constraints using probes. The Chaos-Operator watches it and reconciles it by triggering execution through runners.

### Where do LitmusChaos experiments come from?

ChaosExperiment CRs are hosted on hub.litmuschaos.io, described in the README as a central hub where application developers or vendors share their chaos experiments. The repository also carries CHAOS_EXPERIMENT_MATURITY.md, which records the maturity level of experiments.

### How do I see the result of a LitmusChaos experiment?

The ChaosResult resource holds the outcome: the success of each validation constraint, the revert or rollback status of the fault, and a verdict. The chaos-exporter reads ChaosResults and exposes the information as Prometheus metrics, which the README notes is useful during automated runs.

## Sources

- [License: Apache-2.0](https://github.com/litmuschaos/litmus/blob/master/LICENSE)
- [litmuschaos/litmus on GitHub](https://github.com/litmuschaos/litmus)
- [Project website](https://litmuschaos.io)
- [README](https://github.com/litmuschaos/litmus/blob/master/README.md)
- [Releases](https://github.com/litmuschaos/litmus/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/litmuschaos-litmus
