# Mini-LSM: a Rust course that builds an LSM storage engine from memtable to MVCC

> Mini-LSM is a three-week guided course that has you implement an LSM-tree key-value engine in Rust, covering compaction, crash recovery and transactions. It teaches the storage layer only, and the reference solution ships in the same repository as the starter code.

**skyzh/mini-lsm** — learn database internals by building a storage engine in Rust

- Repository: https://github.com/skyzh/mini-lsm
- Website: https://skyzh.github.io/mini-lsm/
- Stars: 4,172 · Forks: 648
- Language: Rust
- License: Apache-2.0
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/skyzh-mini-lsm

## What Mini-LSM solves for engineers who use databases but have never written one

Most backend engineers can describe what an LSM tree does at a whiteboard level and still have no idea how a memtable, an SST file and a manifest interact when a write lands. Mini-LSM targets exactly that gap. It is a course, not a library: the README describes it as a hands-on course in database internals for systems and backend engineers, and the deliverable is your own engine, built up chapter by chapter. The prerequisites are stated plainly. You need basic Rust. You do not need prior knowledge of LSM trees, compaction, MVCC or transaction isolation. The README says the course suits people who have used PostgreSQL, MySQL, Redis or RocksDB and want to understand what happens below those APIs. That framing matters, because it tells you the value is in the implementation work, not in reading. If you only want to read about leveled compaction, the book chapters are online, but the repository is organized around code you write and tests you make pass.

## The mechanism: memtable, SST, compaction, manifest and WAL in one engine

The course builds a single storage engine in layers. Week 1 starts with an ordered in-memory write buffer, the memtable, then a merge iterator, then the block format, then the sorted string table, and only then the read and write paths. Week 1.7 adds prefix key encoding and Bloom filters to the SST. Week 2 moves to disk-side maintenance: compaction implementation, a simple leveled strategy, a tiered strategy described as RocksDB universal compaction, and a leveled strategy described as RocksDB leveled compaction. The same week adds a manifest and a write-ahead log for crash recovery, plus batch writes and checksums. Week 3 layers multi-version concurrency control on top: timestamp key encoding, snapshot reads through memtables and a transaction API, a watermark with garbage collection, optimistic concurrency control, serializable snapshot isolation, and compaction filters. The data flow is the classic one. Writes go to the memtable and the WAL; the memtable is frozen into an immutable SST; background compaction merges SSTs; the manifest records which SSTs exist; reads merge the memtable and the SSTs through iterators. What is unusual here is that you write each of those pieces yourself, in order, with the reference solution sitting next to your starter code.

## Installing the toolchain and running your first Mini-LSM command

The repository is a Cargo workspace with four members: mini-lsm, mini-lsm-mvcc, mini-lsm-starter and xtask. The README directs students to modify code in mini-lsm-starter. The first step is installing the project's own tools through the xtask alias, which is what makes the cargo x subcommands available. Run it from the repository root.

```bash
cargo x install-tools
```

After that, the README gives the copy-test command, which pulls the tests for a given week and day into your working tree. This is the loop you repeat as you move through the chapters.

```bash
cargo x copy-test --week 1 --day 1
cargo x scheck
```

scheck is the check command for student code. When you want to poke at a running engine rather than run tests, the README offers two reference binaries plus a compaction simulator. These run the final solution, so you can see the intended behaviour before writing it.

```bash
cargo run --bin mini-lsm-cli-ref
cargo run --bin mini-lsm-cli-mvcc-ref
cargo run --bin compaction-simulator-ref
```

The workspace Cargo.toml pins rust-version to 1.97.1 and edition to 2024, so a toolchain older than that will not build the crates. A rust-toolchain.toml file in the repository root handles the version automatically if you use rustup. Course developers, as opposed to students, work in mini-lsm and mini-lsm-mvcc and use cargo x check, cargo x book and cargo x sync instead.

## Where Mini-LSM stops: no SQL, no replication, no distributed consensus

The README is explicit that Mini-LSM focuses on the storage layer of a database and does not cover SQL parsing, query optimization, replication or distributed consensus. That is a real boundary, not a footnote. If your goal is to understand how a query planner picks an index, or how a replicated log commits across nodes, this course will not get you there. The engine you build is a key-value store with ordered keys, snapshots and transactions, and nothing above or beside it. A second limitation is the shape of the work. This is a course with a starter crate and hidden tests, so your implementation has to match the expected API. The README notes that if you change the public API in the reference solution you may need to synchronize it to the starter crate with cargo x sync, which tells you the two are kept in step deliberately. You are not free to redesign the interface as an exercise. Third, the repository layout puts the final solution for weeks 1 and 2 in mini-lsm and the week 3 MVCC solution in mini-lsm-mvcc, both in the same tree as your starter code. That is convenient for checking your work and a temptation to read ahead, which defeats the point of the exercise.

## Mini-LSM compared with reading RocksDB or following a shorter LSM tutorial

The obvious alternative is to read the RocksDB source or its wiki, since Mini-LSM borrows its vocabulary directly: the course names RocksDB universal compaction and RocksDB leveled compaction as the models for chapters 2.3 and 2.4. The difference in approach is the direction of the work. Reading RocksDB means navigating a large production C++ codebase where compaction, column families, rate limiting and statistics are interleaved, and where you cannot easily tell which parts are essential to the LSM idea and which are operational machinery. Mini-LSM inverts that: you implement a minimal version of each idea first, then meet the production variant by name. The cost is fidelity. A course engine does not carry the tuning surface, the background thread pool sizing or the years of bug fixes that a production engine has. A second alternative is a shorter tutorial that ends at a working memtable and SST, which is roughly the first half of week 1 here. Mini-LSM keeps going through compaction strategies, the manifest, the WAL and then three weeks of MVCC, which is why the README frames it as three weeks rather than an afternoon. The repository also points at two projects it says it inspired, SlateDB and Tonbo, both of which put LSM structures over object storage. Those are useful to look at after the course, not instead of it, because the object-storage layer replaces the local file assumptions the course builds on.

## Maintenance, releases and what the Apache 2.0 split actually covers

The last push to the default branch was on 2026-09-23, and the repository is not archived, so it is reasonable to treat it as current. Releases are dated rather than semantic: MiniLSM v202608 on 2026-08-08, Mini-LSM v202501 on 2025-01-20 and Mini-LSM v202401 on 2024-01-30. The workspace version field is 0.4.0-alpha.1, an alpha, so the crates are not published as a stable dependency you would pin in a production Cargo.toml. The upgrade cost for a student is low, because you are working inside a checkout and the course is versioned by release rather than by API stability. The cost that does exist is toolchain drift: the workspace pins rust-version 1.97.1 and edition 2024, so an environment on an older Rust will fail to build until you update. On licensing, the README states that the starter code and solution are under Apache 2.0, and that the author reserves full copyright of the course materials, meaning the markdown files and figures. That is a split licence, and it is the detail people miss. The LICENSE file at the root is Apache 2.0, and licensesnip.config.jsonc plus the .licensesnip entry suggest header management across the crates. If you plan to reuse the book text or diagrams rather than the code, the Apache grant does not obviously cover them, and you should read the README's licence paragraph and the LICENSE file yourself rather than assume.

## Conclusion

Adopt Mini-LSM if you already write Rust and want the storage layer of a database to stop being a black box; the three-week guided course is the path to take, and the coding-agent track only covers days 1 to 3. Skip it if you need SQL parsing, query optimization, replication or distributed consensus, because the README states those are out of scope, and skip it if you want a library to depend on rather than a course to finish. Before starting, run cargo x install-tools and confirm the toolchain matches the rust-version field of 1.97.1 in the workspace Cargo.toml, then decide whether you will work in mini-lsm-starter or read the reference crates first.

## FAQ

### How does an LSM tree work in Mini-LSM?

Writes land in an ordered in-memory memtable and are flushed into immutable sorted string table files; reads merge the memtable and the SSTs through iterators, and background compaction merges SSTs while a manifest tracks which files exist and a write-ahead log supports crash recovery. Mini-LSM has you implement each of those pieces in order, starting with the memtable in chapter 1.1.

### What is the difference between a B tree and an LSM tree in this course?

Mini-LSM only builds the LSM side, so it does not compare the two structures directly. The course covers an ordered memtable, SST files, multiple compaction strategies, a manifest and a WAL, and then MVCC and transactions in week 3.

### Does Mini-LSM use RocksDB?

It does not depend on RocksDB, but it borrows its vocabulary: chapters 2.3 and 2.4 are described as tiered compaction in the style of RocksDB universal compaction and leveled compaction in the style of RocksDB leveled compaction. The README also names RocksDB as one of the systems whose internals the course helps you understand.

## Sources

- [License: Apache-2.0](https://github.com/skyzh/mini-lsm/blob/main/LICENSE)
- [Project website](https://skyzh.github.io/mini-lsm/)
- [README](https://github.com/skyzh/mini-lsm/blob/main/README.md)
- [Releases](https://github.com/skyzh/mini-lsm/releases)
- [skyzh/mini-lsm on GitHub](https://github.com/skyzh/mini-lsm)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/skyzh-mini-lsm
