skyzh/type-exercise-in-rust: a Rust course that builds a vectorized expression engine
Learn advanced Rust techniques by building an expression evaluation framework for a database system.
At a glance
- What is it?
- The repository is an mdBook course plus a Cargo workspace. You implement generic types, erased enums and scalar functions inside type-exercise-starter while supplied chapter tests decide whether you got the boundaries right.
- Who is it for?
- Adopt it if you already write ordinary Cargo and want to see how logical types, physical arrays and erased enums fit together in one execution path, and if you are willing to work only in type-exercise-starter and rerun cargo test -p type-exercise-starter chapter_N --locked after each step.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 7 days ago.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What the course actually asks you to build
The problem it names is specific. A hand-written loop for one integer function is easy, and it stops scaling once the engine has to borrow strings without copying, preserve nulls, read several column encodings, coerce types and pick functions at runtime. That is the situation the course starts from, and the fix it teaches is to move those decisions out of each row loop into type families and checked boundaries.
The audience is a Rust programmer who knows ordinary Cargo use, enums, traits, references and Option. Nothing in the README suggests prior database work is required, but the vocabulary is database vocabulary: logical types, owned values, borrowed values, physical arrays, nullable arrays, constants, Indexed views, one-level Lists. If those words are new to you, the chapters will read as a type-system exercise with unfamiliar nouns attached.
How the type families and checked boundaries fit together
The README describes five layers that the published Chapters 1 to 13 connect: logical types, owned values, borrowed values, physical arrays, and checked erased enums. The point of separating them is that an expression can be described once at the logical level and executed against several physical layouts, including nullable arrays, constants and Indexed views that are read without materializing a new column.
On top of that sit three function interfaces. The README lists generic, erased and bound interfaces, applied to unary, binary and ternary scalar functions. Numeric pairs are promoted for +, -, *, /, comparisons and string contains. Lists are one level deep, with checked offsets and independent outer and child nullability. The same results and errors are supposed to survive one representative fast path, and one ready future is exposed per batch through the static, erased and already-bound expression paths. The course also says explicitly where it stops: no Decimal arithmetic, no casts or rounding, no implicit narrowing or lossy casts, no nested or list-producing functions, no concrete four- and five-input builtins, no exhaustive fast paths, no aggregate engine, no per-row futures.
Installing the workspace and running your first chapter test
There is no package to install. The README points at the published course, which you read in chapter order, and the work happens in the repository. Start by creating a branch and checking that the starter baseline compiles:
git fetch origin
git switch --create course-work --track origin/main
cargo check -p type-exercise-starter --lib --lockedChoose another branch name if course-work already exists. The README says the starter baseline should compile, so a failure here is an environment problem, not an exercise.
When you are ready for a chapter, copy the cumulative supplied contract and run the focused test. The first run is expected to fail, because the behavior you are about to implement does not exist yet:
cargo x copy-test --chapter 1
cargo test -p type-exercise-starter chapter_1 --lockedThe copied destination lands under type-exercise-starter/src/tests/. Implement the named API in the other starter files, then rerun the same command until it passes. The README is firm about two rules: never edit copied tests or src/tests.rs, and keep all earlier copied chapters green. It also says not to inspect type-exercise/ or archived/ while solving an exercise, which is the difference between working the problem and reading the answer.
Where the course is deliberately incomplete
The omission list is the most useful part of the README, because it tells you what you will not be able to build on top of the finished exercises. Decimal arithmetic is out. So are casts and rounding, and so is implicit narrowing or lossy casting. Nested lists and list-producing functions are out. There are no concrete four- and five-input builtins, no exhaustive fast paths, no aggregate engine and no per-row futures. If your goal is a query engine you can extend into a real workload, the course leaves you at the point where those pieces begin.
A second limit is structural. The workspace sets publish = false, so this is not a crate to add to Cargo.toml. The Apache 2.0 licence covers the source code, but the README states the mdBook text is under CC BY-NC-SA 4.0, which is a different licence with a non-commercial restriction. If you plan to reuse the prose, that split matters and is worth checking with someone qualified rather than assuming the code licence covers everything in the repository.
How it differs from reading an engine's source
The obvious alternative is to read a real engine's expression module instead of building one. Bustub, which appears in the related searches, is a teaching database in C++, so the comparison is not just language: you would be reading someone else's C++ for the physical layout decisions rather than writing the Rust type families yourself. Reading scales better if you want breadth, and it teaches you to recognize patterns. It does not force you to make the nullability and offset choices compile.
The closer alternative is a from-scratch tutorial that builds a vector database, which the related searches also mention. That path tends to start at storage and indexing and reach expressions late, if at all. This course inverts the order: it starts at the expression layer, gives you supplied tests as a contract, and never touches an aggregate engine. Pick this one when the type-level part is the part you keep getting wrong.
Maintenance, chapter order and upgrade cost
There are no releases in the repository, and the workspace version is 0.2.0-alpha.1, which tells you the project does not present itself as stable. The repository is not archived, but no last push date is recorded here, so there is no basis for calling it actively maintained. Judge it as course material: the value is in the chapters and the supplied tests, not in a dependency you upgrade.
Two things drive the cost of working through it. The edition is 2024 and the workspace uses resolver = "3", so you need a toolchain new enough for both; the rust-toolchain file in the repository is the authority on which one. The second is the checkpoint/ directory, which holds chapter-*/expr and chapter-*/supplied-tests as workspace members. Those checkpoints make it possible to compare your implementation at a given chapter, but the README asks you not to consult the reference while solving, so the upgrade path is really a discipline question rather than a technical one. Every chapter extends the same starter and keeps earlier supplied tests green, which means a wrong decision in Chapter 4 shows up again later rather than being isolated.
Editorial conclusion
Adopt it if you already write ordinary Cargo and want to see how logical types, physical arrays and erased enums fit together in one execution path, and if you are willing to work only in type-exercise-starter and rerun cargo test -p type-exercise-starter chapter_N --locked after each step. Do not adopt it if you want a library to depend on: the workspace sets publish = false and the README says the course deliberately stops short of Decimal arithmetic, casts, rounding, aggregates and per-row futures. Before you start, confirm that cargo check -p type-exercise-starter --lib --locked passes on your machine, because every later chapter assumes that baseline compiles.
Frequently asked questions
Do I need to install anything to use type-exercise-in-rust?
No package is installed. You clone the repository, create a branch with git switch --create course-work --track origin/main, and verify the baseline with cargo check -p type-exercise-starter --lib --locked.
How do I run the tests for a chapter in type-exercise-in-rust?
Copy the cumulative contract with cargo x copy-test --chapter N, then run cargo test -p type-exercise-starter chapter_N --locked. The first run is expected to fail until you implement the named API.
Can I edit the supplied tests in type-exercise-in-rust?
No. The README says never to edit copied tests or src/tests.rs, and to keep all earlier copied chapters green. The tests act as the contract for each chapter.
What licence covers type-exercise-in-rust?
The source code is Apache 2.0, while the README states the mdBook text is licensed under CC BY-NC-SA 4.0. The two parts carry different terms.
What does type-exercise-in-rust deliberately leave out?
The README lists Decimal arithmetic, casts, rounding, implicit narrowing or lossy casts, nested or list-producing functions, concrete four- and five-input builtins, exhaustive fast paths, an aggregate engine and per-row futures.