Library / SDK
jepsen-io/jepsen avatar
jepsen-io/jepsen

Jepsen: a Clojure framework for fault-injection testing of distributed systems

A framework for distributed systems verification, with fault injection

7,509 stars757 forksClojureLicense varies

At a glance

What is it?
Jepsen is a Clojure library for verifying distributed systems under injected faults, and its README is explicit that it will break machines you point it at. This article covers the mechanism, the setup paths, and the limits.
Who is it for?
Adopt Jepsen if you maintain a distributed datastore, coordination service or scheduler and you can dedicate a disposable cluster of machines with SSH and sudo to it. Do not adopt it if you need a CI-friendly unit testing library, if you cannot give the test boxes root, or if you are not prepared to write Clojure.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly Clojure, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Jepsen is for, and who writes tests with it

Jepsen answers a question that ordinary integration tests cannot: does this system still behave correctly when the network drops, when clocks move, or when a node is killed mid-operation? The README describes it as a Clojure library where a test is a Clojure program that sets up a distributed system, runs operations against it, and verifies that the resulting history makes sense. The stated scope runs from eventually-consistent commutative databases to linearizable coordination systems to distributed task schedulers.

The audience is narrower than the phrase distributed systems suggests. You need to be comfortable writing Clojure, because there is no declarative test format here: the test is code that pulls in the Jepsen library. You also need machines you are willing to have broken. The README's own warning is blunt: tests may mess with clocks, add apt repos, run killall -9 on processes, and generally break things, so you should not point Jepsen at production machines unless you wrote the test and know exactly what it is doing. That sentence should be treated as the entry requirement, not as fine print.

What you get in return is evidence rather than opinion. The README says Jepsen can generate graphs of performance and availability, helping you characterize how a system responds to different faults, and points to jepsen.io/analyses for examples of the analyses you can carry out.

Control node, DB nodes, generator, nemesis, checker

The architecture is spelled out in the README's design overview and it is worth following the data flow in order. A test runs as a Clojure program on a control node. That program uses SSH to log into a set of db nodes, where it sets up the system under test using the test's pluggable os and db components. Those two pluggable pieces are where the portability lives: the os abstraction handles machine-level concerns and the db abstraction handles the datastore's lifecycle.

Once the system is running, the control node spins up a set of logically single-threaded processes, each with its own client for the distributed system. A generator produces new operations for each process to perform, and the processes apply them through their clients. The start and end of every operation is recorded in a history. In parallel, a special nemesis process introduces faults into the system, and the generator schedules those faults too. So the same scheduling machinery drives both normal load and disruption, which is what makes the fault timeline reproducible alongside the operation history.

Tear-down comes next: the DB and OS are torn down. Then a checker analyzes the test's history for correctness and generates reports and graphs. Everything the test produced, the history, the analysis and any supplementary results, is written under store/<test-name>/<date>/ for later review, with symlinks to the latest results maintained at each level. That directory layout is the practical interface between a test run and whatever you do with the results afterwards.

Installing Jepsen and running a first test against a cluster

Jepsen is distributed as a Clojure library, published on Clojars as jepsen/jepsen, and the README's own workflow assumes you run tests from a checkout with Leiningen. There is no single install command in the README; the library is a dependency you add to a test project, and the README's setup instructions are about getting machines, not about installing a binary.

The README gives several ways to get a cluster. The AWS Marketplace listing for Jepsen LLC, for example, offers either a single EC2 VM with several DB nodes in containers, or a cluster of separate EC2 VMs. The README notes the separate-VM option is a little more expensive but significantly faster, and that it lets you test clock skew. The marketplace VMs carry an hourly fee, generally $1/hour/node according to the README, which it says helps fund Jepsen development.

After launching the CloudFormation stack, the README says to SSH to the control node with the key you chose:

bash
ssh -i <your-ec2-ssh-key.pem> admin@ec2-<whatever>.compute-1.amazonaws.com

On a multi-VM cluster the DB nodes are named n1, n2 and so on, and the README's example invocation is:

bash
lein run test --username admin --nodes-file ~/nodes ...

For the single-VM container option, you start the containers on the instance and the DB nodes are named c1, c2, and so on, listed in ~/container-nodes. The README gives the corresponding invocation:

bash
lein run test --nodes-file ~/container-nodes ...

Note that both examples end in an ellipsis: the README does not spell out the remaining arguments, so expect to read your specific test's code for the flags it needs. If you would rather not use AWS, the README says you can run Jepsen against almost any machines with a TCP network, an SSH server, and sudo or root access. DB nodes default to the names n1 through n5, and the SSH username, password and identity files are all definable in the test or at the CLI. LXC containers are also supported for local hacking, with the caveat that containers do not have real clocks, so you generally cannot use them to test clock skew. For Docker, the README is explicit that it is unsupported, and points to a community-maintained project instead.

Where Jepsen is the wrong tool

The most important limitation is operational, not technical. Jepsen assumes it owns the machines. The README states that the account you use on the DB nodes needs sudo access to set up DBs and control firewalls, and that tests may manipulate clocks, add apt repositories, and run killall -9. If your environment cannot hand a test runner root on a set of disposable hosts, Jepsen's model does not fit, and there is no sandboxed mode described in the README that would change that.

The second limitation is portability of tests between environments. The README warns that most Jepsen tests are written with more specific requirements in mind, like running on Debian and using iptables for network manipulation, and tells you to see the specific test code for details. So the framework is general, but the tests usually are not. A test written against one distribution and firewall toolchain is not a portable artifact you can hand to another team unchanged.

The third is clock fidelity. LXC containers are described as lacking real clocks, which rules them out for clock-skew testing, and the README frames the multi-VM AWS option as the one that lets you test clock skew. If skew is part of what you need to verify, the cheap container setup is the wrong choice.

Finally, the README does not document rollback of the changes a test makes to DB nodes. Given that tests can add apt repos and kill processes, anyone expecting the framework to restore a machine to its prior state should assume it does not, and use hosts that can be discarded.

Jepsen versus a chaos-engineering tool

The natural comparison is with fault-injection tools built around containers and orchestration, such as Chaos Mesh or LitmusChaos. The difference is what the two approaches treat as the output. A chaos tool typically injects a fault into a running system and leaves you to observe whether your monitoring noticed. Jepsen's README describes something different: it sets up the system itself through pluggable os and db components, drives it with generated operations from logically single-threaded processes, records a history with the start and end of every operation, and then runs a checker over that history for correctness. Faults are scheduled by the same generator that schedules load, so the disruption is part of the test definition rather than an external event.

That changes the skill requirement. Chaos tooling is configured, usually in YAML, and fits teams that already run Kubernetes. Jepsen is programmed in Clojure, and correctness checking means reasoning about the history your operations produced. The payoff is a verdict about consistency, not just an observation that something restarted. If your question is whether a system stayed linearizable while a node was partitioned, Jepsen is built for that question. If your question is whether your alerting fires when a pod is killed, a chaos tool answers it with far less ceremony.

Maintenance, releases and licensing

The repository is not archived, and the last push was on 2026-09-03, so the project is being worked on. Recent releases are v0.3.13 on 2026-07-31, v0.3.12 on 2026-07-20, and v0.3.11 on 2026-03-10. The release cadence over that window is steady rather than dramatic, and the version numbers stay in the 0.3.x line, which is worth noting if your dependency policy treats pre-1.0 versions differently.

Upgrade cost is mostly a matter of how much of your test code touches internals. Because a Jepsen test is a Clojure program that calls the library directly, rather than a configuration file, library changes can reach into your test code in ways a config schema would not. The README does not describe a compatibility policy or a deprecation process, so pinning a known-good version in your project and reading the release notes before moving is the conservative path.

On licensing: the repository metadata available here does not state a license, and the README does not name one either. Clojars and the README point to the library under the jepsen/jepsen coordinate, but no license identifier appears in the README or the repository metadata. Anyone planning to depend on it, especially in a commercial setting, should read the LICENSE file in the repository, or its absence, before adopting it. That is a factual gap in the project's documentation, and it is the kind of gap that procurement teams will ask about.

Editorial conclusion

Adopt Jepsen if you maintain a distributed datastore, coordination service or scheduler and you can dedicate a disposable cluster of machines with SSH and sudo to it. Do not adopt it if you need a CI-friendly unit testing library, if you cannot give the test boxes root, or if you are not prepared to write Clojure. Before writing a single test, verify three things: that your control node can reach every DB node over SSH, that the account on those boxes has sudo, and that the nodes are not production. The README's warning about killall -9 and clock manipulation is the boundary that decides whether this tool is for you.

Frequently asked questions

What is Jepsen?

Jepsen is a Clojure library for distributed systems verification with fault injection. A test is a Clojure program that uses the library to set up a distributed system, run operations against it, and verify that the history of those operations makes sense.

What is a Jepsen test?

A Jepsen test is a Clojure program that runs on a control node, uses SSH to set up a distributed system on DB nodes, drives it with generated operations, injects faults through a nemesis process, and then runs a checker over the recorded history for correctness.

Can I run Jepsen in Docker?

The README lists Docker as unsupported by Jepsen and points to a community-maintained project, Jepsen in Docker, Community Edition. LXC is documented as a supported local option, with the caveat that containers do not have real clocks.

What machines does Jepsen need?

The README says almost any machines with a TCP network, an SSH server, and sudo or root access will work. Each DB node must be reachable from the control node over SSH, and the account used needs sudo to set up DBs and control firewalls.

Where does Jepsen write test results?

The test, history, analysis and any supplementary results are written to the filesystem under store/<test-name>/<date>/. Symlinks to the latest results are maintained at each level for convenience.

Is Jepsen safe to run against production machines?

No. The README warns that tests may mess with clocks, add apt repos, run killall -9 on processes, and generally break things, and advises against pointing Jepsen at production machines unless you wrote the test and know exactly what it is doing.

Official sources

  1. Issues
  2. jepsen-io/jepsen on GitHub
  3. README
  4. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/jepsen-io-jepsen.svg)](https://hysenlabs.com/projects/jepsen-io-jepsen)