Open-source project
erikgrinaker/toydb avatar
erikgrinaker/toydb

toyDB: A Distributed SQL Database in Rust Built to Teach Database Internals

Distributed SQL database in Rust, written as an educational project

7,288 stars623 forksRustApache-2.0

At a glance

What is it?
toyDB is an educational distributed SQL database written in Rust from scratch, covering Raft consensus, MVCC snapshot isolation, BitCask storage, and an iterator-based SQL query engine. It is built to be readable and correct, not performant, and its author states that performance, scalability, and availability are explicit non-goals.
Who is it for?
toyDB is the right reference for a developer or student who wants to trace through a complete distributed SQL database implementation from consensus layer to query execution in a single readable codebase.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 66 days ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What toyDB Is Built to Teach

toyDB is a distributed SQL database written from scratch in Rust, designed to be simple and understandable at the cost of performance and production readiness. The author, Erik Grinaker, wrote the first version in 2020 to learn database internals, then rewrote it after spending several years building real distributed SQL databases at CockroachDB and Neon. The README describes the rewrite as a simple illustration of the architecture and concepts behind distributed SQL databases.

The educational value comes from the scope. toyDB covers Raft distributed consensus for replication, MVCC-based snapshot isolation for transactions, a pluggable storage engine with BitCask and in-memory backends, an iterator-based query engine with heuristic optimization, and a SQL interface including joins, aggregates, and transactions. Each component is implemented from scratch, which means the code shows the mechanism directly rather than delegating to external libraries.

The project explicitly states shortcuts: aspects like performance, scalability, and availability introduce major complexity in production-grade databases and obscure the underlying concepts, so they are not addressed. This makes toyDB a cleaner study resource than a production codebase but rules it out for any real workload.

Architecture: Raft, MVCC, and the SQL Query Engine

The README describes toyDB's architecture as fairly typical for a distributed SQL database: a transactional key/value store managed by a Raft cluster with a SQL query engine on top. The architecture guide at `docs/architecture/index.md` provides the detailed tour.

The Raft implementation handles linearizable state machine replication. Each node runs on two ports: one for SQL clients and one for Raft peer communication. The `./cluster/run.sh` script starts a five-node cluster on ports 9601 to 9605 for SQL and ports 9701 to 9705 for Raft.

On top of Raft sits an MVCC layer in `src/storage/mvcc.rs`, providing snapshot isolation. The storage engine is pluggable: BitCask is the durable default, and an in-memory backend is available for testing. The query engine in `src/sql/execution/executor.rs` uses an iterator model where each node in the query plan pulls rows from its child, and the optimizer in `src/sql/planner/optimizer.rs` applies heuristic rules. The SQL parser supports joins, aggregates, transactions, and time-travel queries.

The test suite uses the Goldenscript framework, which scripts scenarios, captures events and output, and later asserts that behavior remains identical. The Raft cluster tests, MVCC transaction tests, and SQL execution tests each live in dedicated `testscripts/` directories.

Running a Five-Node Cluster and Connecting a SQL Client

With a Rust compiler installed, the cluster script handles the full setup:

bash
./cluster/run.sh

The script starts five nodes on ports 9601 to 9605 for SQL and 9701 to 9705 for Raft, with data stored under `cluster/*/data/`. The output shows each node announcing its SQL and Raft listening addresses, followed by Raft leader election logs.

Connecting the command-line SQL client to node 1:

bash
cargo run --release --bin toysql

The client connects to `localhost:9601` by default. From the prompt, standard SQL works:

code
toydb> CREATE TABLE movies (id INTEGER PRIMARY KEY, title VARCHAR NOT NULL);
toydb> INSERT INTO movies VALUES (1, 'Sicario'), (2, 'Stalker'), (3, 'Her');
toydb> SELECT * FROM movies;

The `EXPLAIN` command displays the query plan. The README shows a plan for a multi-join query that includes Remap, Order, Projection, and Aggregate nodes, giving a readable view of how the heuristic optimizer structures execution.

The test suite runs with:

bash
cargo test

The Write Performance Problem and the Benchmark Tool

toyDB ships a `workload` benchmark tool for measuring throughput against a running cluster. The README gives this example:

bash
cargo run --release --bin workload read

Three workloads are available: `read` for single-row primary key lookups, `write` for sequential single-row inserts, and `bank` for multi-table bank transfers with joins, secondary indexes, and conflicts.

The benchmark results in the README make the write performance problem concrete. With the BitCask storage engine and fsync enabled, write throughput is 35 transactions per second. Disabling fsync raises it to 4,719 transactions per second. Using the in-memory engine reaches 7,781 transactions per second. Read performance is roughly the same across all three configurations, around 13,900 to 14,200 transactions per second.

The README explains the cause directly: fsync and a lack of write batching in the Raft layer. Each write transaction goes through the full Raft commit cycle, and each disk write includes an fsync. Fixing this would require implementing write batching in the Raft layer, which the project treats as an optimization outside its educational scope.

The `workload` tool accepts parameters for rows, batch size, and worker count. Run `cargo run --bin workload -- --help` for the full list.

What toyDB Does Not Cover

Several topics central to production distributed databases are absent by design. The README names performance, scalability, and availability as explicit non-goals, each dismissed as a source of complexity that obscures the basic underlying concepts.

On the availability side, the Raft implementation handles leadership election and log replication, but more advanced availability patterns such as read replicas, follower reads, or multi-region routing are not present. The five-node cluster from `./cluster/run.sh` is a development setup, not a high-availability deployment.

On the SQL side, the feature set covers joins, aggregates, and transactions, but the README does not document rollback behavior in detail beyond the MVCC foundation, and the SQL reference at `docs/sql.md` is the place to check what dialect is and is not supported. There is no documented extension mechanism for custom functions or types.

The Cargo.toml marks the crate as `publish = false` and lists no GitHub releases, confirming that toyDB is a study resource rather than a packaged library. A developer who wants to import distributed SQL capabilities into their own application should look at production databases such as CockroachDB or Neon, which inspired this project.

License, Dependencies, and Maintenance

toyDB is released under the Apache-2.0 license. The Cargo.toml lists the author as Erik Grinaker. The project has no external database or consensus library dependencies: Raft, MVCC, and the query engine are all written from scratch. The dependency set covers serialization (bincode, serde), CLI parsing (clap), logging (log, simplelog), and histograms for the benchmark tool (hdrhistogram).

The last push to the repository was on 2026-07-27. The project is not archived. Given that it is explicitly described as a rewrite done after its author gained production experience, and that it aims to be a clean reference rather than an evolving product, updates are likely to remain infrequent and focused on correctness and documentation rather than new features.

The architecture guide and SQL examples at `docs/examples.md` are part of the repository and provide structured reading alongside the code. The references document at `docs/references.md` lists the research materials used while building the project, which is itself a useful starting point for anyone who wants to read the primary sources behind the design decisions.

Editorial conclusion

toyDB is the right reference for a developer or student who wants to trace through a complete distributed SQL database implementation from consensus layer to query execution in a single readable codebase. It is not a fit for any production workload: the README explicitly names performance, scalability, and availability as non-goals, and write throughput with the default BitCask storage drops to 35 transactions per second due to fsync and absent write batching in the Raft layer. Before using it as a study resource, read the architecture guide at docs/architecture/index.md alongside the source, since the README points there for the full design explanation.

Frequently asked questions

Can toyDB be used in a production application?

No. The README explicitly states that performance, scalability, and availability are non-goals. Write throughput with the default storage engine is 35 transactions per second due to fsync and absent write batching. The project is intended as an educational reference for understanding distributed SQL database internals.

What SQL features does toyDB support?

The README lists joins, aggregates, and transactions as supported features. The `EXPLAIN` command displays query plans. The full SQL dialect is documented in the SQL reference at `docs/sql.md` in the repository.

How does toyDB differ from a production distributed database like CockroachDB?

The README states directly that toyDB takes shortcuts wherever possible on performance, scalability, and availability, because those are major sources of complexity that obscure the basic concepts. CockroachDB, which the author worked on before rewriting toyDB, is a production system where those concerns are primary. toyDB is a simplified illustration of the same architectural pattern.

Official sources

  1. erikgrinaker/toydb on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/erikgrinaker-toydb.svg)](https://hysenlabs.com/projects/erikgrinaker-toydb)