jepsen-io/maelstrom: a workbench for toy distributed systems in any language
A workbench for writing toy implementations of distributed systems.
At a glance
- What is it?
- Maelstrom runs your binaries as nodes on a simulated network, feeds them JSON over STDIN and STDOUT, and checks the recorded history with Jepsen. It is a teaching harness, not a production test rig, and its documentation is the product.
- Who is it for?
- Adopt Maelstrom if you are learning distributed algorithms or teaching them, and you want a checker that can flag serializability violations rather than a pass/fail smoke test. Do not adopt it for production validation: the README calls the target systems toy implementations, and the simulated network replaces real process management, real sockets and real serialization, which is exactly the busywork you would need to test.
- Can I use it commercially?
- Yes, with conditions. EPL-1.0 is a weak copyleft licence: you can use it inside commercial and closed-source software, but if you distribute changes to its own files, you must publish those changes under the same licence.
- Is it still maintained?
- Yes. The repository last received commits 82 days ago.
- What is it written in?
- Mainly Clojure, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The busywork Maelstrom removes, and who still has to do it
The README is direct about the problem: writing real distributed systems involves process management, networking and message serialization, and running a cluster of virtual machines on a real IP network is awkward for many people. Maelstrom deletes that layer. Your program is a plain binary. It reads JSON messages from STDIN, writes JSON messages to STDOUT, and logs to STDERR. Maelstrom spawns instances of that binary as processes on your local machine and connects them through a simulated network. There is no service discovery, no daemonization, no socket code, and no test harness to write.
The audience follows from that. This is for people learning distributed algorithms, and for instructors running a workshop; the README says Maelstrom is used as part of a distributed systems workshop by Jepsen. It is also usable by engineers who already know the theory and want to see how a protocol behaves under latency and message loss without provisioning anything. It is not aimed at teams validating a production database. The README calls the targets toy implementations, and that framing is honest rather than modest: the simulated network is the whole point, and it is also the boundary of what the results mean.
How a Maelstrom run is wired: CLI to Jepsen to simulated network
The design overview in the README traces the data flow. Maelstrom starts in maelstrom.core, which parses CLI operations and builds a test map. That map goes to Jepsen, which sets up servers and Maelstrom services through maelstrom.db. Spawning binaries and handling their IO lives in maelstrom.process. Built-in services such as lin-kv are defined in maelstrom.service. Jepsen then spawns clients according to the workload (maelstrom.workload.*) plus a nemesis (maelstrom.nemesis) that injects faults.
During the run, messages between nodes are routed by maelstrom.net and recorded in maelstrom.journal. Clients send requests and parse responses through maelstrom.client, which also defines the RPC protocol. After the run, Jepsen checks the accumulated history using a checker built in maelstrom.core, with workload-specific checkers in their own namespaces. Network statistics are computed in maelstrom.net.journal, and Lamport diagrams come from maelstrom.net.viz. The README also notes that maelstrom.doc generates documentation from maelstrom.client's registry of RPC types and workloads.
That split matters when something goes wrong. If your node never receives a message, the fault is in your STDIN loop or in the protocol shape. If the run completes but the checker fails, the fault is in your algorithm. The journal and the visualizations exist to make that distinction visible rather than guessed.
Installing Maelstrom and running a first workload
Maelstrom is a Clojure program on the JVM, packaged as a jar. The README points to the guide's first chapter for setup, and the CLI is invoked as java -jar maelstrom.jar with a subcommand. The three options the README singles out are the workload, the binary to spawn, and the node count:
java -jar maelstrom.jar test \
--workload echo \
--bin ./demo/python/echo.py \
--node-count 3The repository ships demo implementations in several languages under demo/, including c++, clojure, go, java, js, python, ruby and rust. Running one of those first is the cheapest way to confirm the jar, the JVM and the process spawning all work before you write any code. A full option list is available from the jar itself:
java -jar maelstrom.jar test --helpWhen a node misbehaves and the summary is not enough, two flags in the README surface more detail: --log-stderr shows STDERR output from each node inside the Maelstrom log, and --log-net-send logs outgoing network messages. STDERR is where your node's own logging should go, because STDOUT is reserved for protocol messages. After a run, the artifacts described in doc/results.md are the place to look: plots, statistics and log files that explain what the checker saw.
What the checkers actually prove, and what they do not
The checking is the part that separates Maelstrom from a stress script. The README states that Maelstrom's checkers can verify sophisticated safety properties up to strict serializability, using Elle, and generate minimal examples of consistency anomalies. The promotional image in doc/ is a Lamport diagram alongside a serialization anomaly drawn as a cycle of dependency edges between transactions. A cycle is a concrete artifact you can read, not a red X.
What that does not give you is a statement about your real system. The history being checked was produced by processes on one machine talking over a simulated network that Maelstrom controls. Latency and message loss are injected by Maelstrom, not observed from a kernel. A protocol that passes here has survived the fault model the workload and nemesis define, and nothing more. The README's performance note is worth reading in the same spirit: Maelstrom is described as reasonably fast, able to handle simulated clusters of 25+ nodes, and on a 48-way Xeon able to use 94% of cores while pushing upwards of 60,000 network messages per second. That is a statement about the harness's capacity, not a claim about the systems it tests.
Where Maelstrom is the wrong tool
The simulated network is the limitation. Because nodes are local processes exchanging JSON over pipes, Maelstrom cannot exercise the failure modes that live below the protocol: TCP backpressure, partial writes, slow disks, clock skew on different hosts, or a process that dies in a way the parent does not notice. If your question is whether your Raft implementation survives a real network partition between real machines, Maelstrom answers a nearby but different question.
The second boundary is language and runtime. The protocol is JSON over STDIN and STDOUT, and the README invites you to write Plumtree in Bash or Byzantine Paxos in Intercal. That openness is real, but it means the harness knows nothing about your runtime. A garbage collection pause, a thread pool starvation or a runtime-level deadlock appears to Maelstrom only as a node that stopped replying. You will be reading your own logs, not Maelstrom's diagnostics.
Finally, the documentation is the interface. The README routes almost everything to the guide chapters and the reference documents on protocol, workloads, results and services. If you want a tool that tells you what to do from a single --help, this is not that tool.
Maelstrom against a hand-rolled harness
The obvious alternative is to write your own test driver: spawn N processes, give them addresses, send requests, collect responses, and assert something at the end. That approach is fully general and it is the only way to test against real sockets on real hosts. The difference in approach is where the fault injection and the checking live. In a hand-rolled harness, both are yours: you decide which messages to drop, when to partition, and what counts as a violation, and you write the checker. In Maelstrom, the nemesis injects faults and Jepsen checks the history, and the workload definitions in doc/workloads.md fix what a commutative set or a transactional key-value store is supposed to mean.
The trade is control for comparability. A hand-rolled harness can model anything, including things Maelstrom's protocol cannot express, but two hand-rolled harnesses are not comparable and neither is reusable by the next person. Maelstrom's workloads and checkers are shared, which is why an anomaly it reports is meaningful to someone else. If your goal is to learn an algorithm, the shared checker is worth more than the extra fidelity. If your goal is to ship, the extra fidelity is the point.
Maintenance, upgrades and the EPL-1.0 licence
The repository is not archived. Its last push was on 2026-07-10, which is recent enough to treat the codebase as moving. The release history is sparser: v0.2.4 landed on 2024-12-04, after v0.2.3 on 2023-02-27 and v0.2.2 on 2023-01-13. So commits and tagged releases do not move at the same pace, and pinning a version means pinning something that may be older than the branch you read on GitHub.
The upgrade cost is mostly the guide, not the jar. Maelstrom's surface for you is the JSON protocol and the workload definitions, both documented under doc/. A change to the protocol document or to a workload's message shape is what would force you to edit your node, and those documents are versioned alongside the code in the same repository. The pragmatic move is to read doc/protocol.md and doc/workloads.md at the tag you run rather than at main, since the README does not describe a compatibility policy between releases.
The project is under EPL-1.0, a file-level copyleft licence. The LICENSE file is at the repository root. This is not legal advice; if you plan to redistribute Maelstrom or embed it in a commercial product, read the licence text and, where it matters, ask a lawyer. For the common case of running the jar locally to test your own code, the licence question is unlikely to be the blocker.
Editorial conclusion
Adopt Maelstrom if you are learning distributed algorithms or teaching them, and you want a checker that can flag serializability violations rather than a pass/fail smoke test. Do not adopt it for production validation: the README calls the target systems toy implementations, and the simulated network replaces real process management, real sockets and real serialization, which is exactly the busywork you would need to test. Before committing, verify that your language has a demo under demo/ so the JSON protocol wiring is already solved, and confirm the workload you care about appears in doc/workloads.md.
Frequently asked questions
What is jepsen-io/maelstrom?
It is a workbench for learning distributed systems by writing your own, built on the Jepsen testing library. It runs your binaries as nodes on a simulated network and checks the recorded history against standardized workloads.
Which languages can I write a Maelstrom node in?
Any language. Nodes read JSON from STDIN and write JSON to STDOUT, and the repository ships demo implementations under demo/ for c++, clojure, go, java, js, python, ruby and rust.
Does Maelstrom test real distributed systems in production?
No. The README describes it as a workbench for toy implementations, and nodes run as local processes connected by a simulated network rather than over real sockets between real hosts.
How do I see what a node is printing while a test runs?
Pass --log-stderr to show STDERR output from each node in the Maelstrom log, and --log-net-send to log outgoing network messages. Node logging belongs on STDERR because STDOUT carries protocol messages.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/jepsen-io-maelstrom)