# skyzh/tiny-llm: a four-week course that builds a Qwen3 inference path on Apple Silicon

> skyzh/tiny-llm is a hands-on course for systems engineers who want to implement LLM inference themselves, from matmul to a mini vLLM and a coding agent. It runs on MLX arrays, not high-level neural-network layers, and the reference solution doubles as the correctness oracle.

**skyzh/tiny-llm** — learn LLM inference system on Apple Silicon for systems engineers: build a tiny vLLM + Qwen

- Repository: https://github.com/skyzh/tiny-llm
- Website: https://skyzh.github.io/tiny-llm/
- Stars: 4,733 · Forks: 399
- Language: Python
- License: Apache-2.0
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/skyzh-tiny-llm

## What skyzh/tiny-llm solves, and for whom

Most material about LLM serving explains the concepts and then hands you a framework. You learn that paged attention exists, but you never see where the page table is consulted or why the scheduler stops rebuilding dense history on every decode step. skyzh/tiny-llm takes the opposite route. The README describes it as a hands-on course for systems engineers who want to understand LLM inference end to end, and compares it to CMU's Needle project: an LLM-serving counterpart in which you build the path that loads a Qwen3 model, turns tokens into logits and generates text.

The audience is narrow on purpose. You need to be comfortable with array programming, memory traffic and kernel occupancy, because the course connects equations to those quantities. It is not an introduction to Python, and it is not a tutorial on prompting a hosted model. The stated goal is that the implementation stays small enough to read end to end. That constraint is what makes the course different from a framework contribution: every operator you meet is one you wrote.

## MLX as the floor, your code as the implementation

The course is built on MLX arrays and the MLX extension runtime, deliberately without high-level neural-network layers. When a chapter teaches an operator, your solution implements that operator in Python, C++ or Metal rather than calling the corresponding optimized MLX operation. MLX remains the correctness oracle and performance baseline, so you always have a known-good answer to compare against and a known-fast implementation to measure yourself against.

That design has a direct consequence for the exercises. You cannot finish a chapter by importing the thing the chapter is about. Week 1 builds a Qwen3 model directly from mlx.core array operations: attention, RoPE, GQA, RMSNorm, the MLP, sampling and the autoregressive loop. Week 2 adds a KV cache, establishes a synchronized MLX baseline, and lets matched benchmarks choose each optimization, moving from quantized decode matvec to fused model kernels, tiled prefill and split-K where the measured Qwen shapes need it. Week 3 introduces continuous batching and chunked admission, then makes paged KV the canonical serving layout, so decode attention and FlashAttention learn to read pages directly.

The repository reflects this split. The tiny_llm package is where students implement the exercises; tiny_llm_ref contains the reference solution used by the tests and benchmark appendix. Two parallel trees exist because the tests need a correct implementation to compare against, and the extension build scripts mirror that split with build-ext and build-ext-ref.

## Installing the course and running your first check

The book is published at skyzh.github.io/tiny-llm, and the README points to an environment setup chapter as the starting point. The project uses pdm, and pyproject.toml pins requires-python to >=3.10, <3.13, so Python 3.13 is outside the supported range. The dependency list is heavier than a typical tutorial: mlx, torch, torchtune, torchao, mlx-lm, numpy, pytest, ruff, nanobind and pytest-benchmark. The comment in pyproject.toml notes that setuptools is included to simplify setup even though it would not usually appear in a project dependency list.

If you already have a checkout, the README gives this verification sequence:

```bash
pdm install -v
pdm run check-installation
pdm run test-refsol -- -- -k week_1
```

The first command resolves and installs the environment. The second checks that the installation is usable. The third runs the reference-solution tests filtered to Week 1, which is the quickest signal that your MLX setup actually executes the course code. If those pass, the toolchain is sound and you can start implementing in tiny_llm.

Running the course itself goes through the pdm scripts defined in pyproject.toml. Each week has its own loader:

```bash
pdm run main-week1
pdm run main-week2
pdm run main-week3
```

The loader selects which stage of the model path main.py assembles, so week1 exercises the plain Qwen3 implementation while later loaders pull in the cache and serving machinery you built in earlier weeks. Benchmarking has its own entry points, including pdm run bench and pdm run bench-week2-operators.

## Week 4 is a course in progress, and the roadmap says so

The roadmap table tracks four columns per chapter: Code, Test, Doc and Audit, where Audit is described as Chi's personal editorial pass on the published course content, independent of code and test readiness. Week 1 is complete across all four columns. Every Week 2 and Week 3 chapter, including the optional MoE and speculative decoding chapters, is marked complete for code, test and doc but still shows a construction marker in Audit. Week 4 is publishing one reviewed day at a time, with Days 1 through 9 currently available.

That distinction matters if you are planning a study schedule. A chapter can have a working implementation and passing tests while the learner-facing text has not had its editorial pass, so the prose may lag the code. The Week 4 overview is explicitly called out as required reading before running the agent loop, and the README warns that Day 3 can send file contents to the model, modify files after approval, and run one exact configured command. It advises using a disposable workspace without secrets. That is not boilerplate caution: an agent that reads your files and executes a configured command deserves an isolated directory, and the course says so rather than assuming you will figure it out.

The later Week 4 days are unusually specific. Day 4 checkpoints a complete tool-observation boundary with the scripted model's fake cache metadata and restores a fresh model without replaying the completed edit or command. Day 5 compacts older completed effects in the model-visible transcript while their exact receipts retain the full action, result and changed artifacts. Day 8 reuses one real tokenizer and KV checkpoint for two differently steered, effect-isolated continuations and makes one explicit passing selection without pretending completed effects were rewound. This is checkpoint and steering semantics, not prompt engineering, and the fact that the course names the boundary explicitly is the most interesting thing in the repository.

## Where the course is the wrong tool

The most obvious limitation is hardware. The README frames Apple silicon as the practical local environment: one shared memory space, direct access to Metal kernels, and the ability to inspect the complete path on one machine instead of depending on an expensive CUDA GPU setup. That is a real advantage for a student, and also a hard boundary. If your work happens on NVIDIA hardware, or you need multi-GPU tensor parallelism, the course does not target it. Search interest in running tiny models on an ESP32, a Raspberry Pi or in Ollama has nothing to do with what this repository does, and the README lists other topics as not covered.

The second limitation is scope. This is a course, not a serving library. The version in pyproject.toml is 0.1.0 and there are no retrieved releases, so there is no upgrade path to reason about. You would not put the tiny_llm package in front of traffic; the point is that you wrote it. The dependency on torch, torchtune and torchao alongside MLX also means the environment is not lightweight, which sits awkwardly with the small-implementation goal even though the comment in pyproject.toml explains the setuptools inclusion as a setup convenience.

The third is the audit gap. Week 2 through Week 4 chapters are code and test complete but not editorially reviewed, so a learner hitting an unclear paragraph should check the reference solution and the tests rather than assume the chapter text is final.

## How it differs from vLLM and from a from-scratch transformer tutorial

The name invites the comparison, and the course leans into it: Week 3 is titled Build a Mini vLLM. The difference in approach is that vLLM is a production serving system you install and configure, while this course is a sequence of exercises in which continuous batching, chunked admission and paged KV are things you implement and then benchmark. The README is explicit that matched benchmarks choose each optimization, so the optimizations are justified by measurement on Qwen shapes rather than adopted because a paper recommends them.

The comparison to a from-scratch transformer tutorial is sharper. Such tutorials typically stop at a forward pass and a sampling loop, often using high-level layers. Here the first week alone covers attention, RoPE, GQA, RMSNorm, the MLP, sampling and the autoregressive loop directly on mlx.core, and the later weeks keep going into kernels, cache layout and scheduling. The trade-off is time: four weeks of work, with Week 4 still being published day by day, against a from-scratch tutorial you can finish in an afternoon. If you want a working chatbot quickly, the shorter path wins. If you want to know why a decode step reads the pages it reads, this is the one that gets there.

## Licence and the cost of keeping a checkout current

The repository is Apache-2.0, which permits commercial and private use and includes an explicit patent grant. The licence file is at the repository root. Nothing in the published repository suggests a separate licence for the course text or the reference solution, but the book content and the code live in different parts of the tree, so if you plan to reuse the written chapters rather than the code, read the LICENSE file yourself rather than assuming one licence covers both. That is a factual observation about the layout, not legal advice.

Upgrade cost is low by design. There are no retrieved releases, the package version is 0.1.0, and the last push was on 2026-09-09, so the practical maintenance question is whether the MLX, torch and mlx-lm version floors in pyproject.toml still resolve together on your machine when you next run pdm install -v. Because Week 4 is being published one reviewed day at a time, a learner working through it should expect the later days to appear after they start, and the roadmap table is the place to check what is currently available.

## Conclusion

Adopt skyzh/tiny-llm if you already write systems code and want to implement attention, KV caching, paged attention and continuous batching yourself on a Mac instead of reading about them. Skip it if you want a production serving stack, a CUDA or multi-GPU target, or a project that runs on Windows and Linux hardware. Before committing a week, run pdm install -v, pdm run check-installation and pdm run test-refsol -- -- -k week_1 on your machine and confirm the Week 1 tests pass; then check the roadmap table for whether the chapter you care about has reached the Audit column, because Week 2 through Week 4 are code and test complete but still marked as under review.

## FAQ

### What is skyzh/tiny-llm?

It is a hands-on course for systems engineers who want to understand LLM inference end to end, described in the README as an LLM-serving counterpart to CMU's Needle project. You build the path that loads a Qwen3 model, turns tokens into logits and generates text, following a four-week learning path from matmul to a mini vLLM and a coding agent.

### What does the skyzh/tiny-llm course cover in each week?

Week 1 builds a Qwen3 model from mlx.core array operations including attention, RoPE, GQA, RMSNorm, the MLP, sampling and the autoregressive loop. Week 2 adds a KV cache and benchmark-driven optimizations, Week 3 introduces continuous batching, chunked admission and paged KV, and Week 4 builds a bounded, validated agent loop connected to a small workspace.

### How do I install skyzh/tiny-llm?

The book points to an environment setup chapter at skyzh.github.io/tiny-llm/setup.html, and the README gives a verification sequence for an existing checkout: pdm install -v, then pdm run check-installation, then pdm run test-refsol -- -- -k week_1. The project requires Python >=3.10, <3.13 and uses pdm as its package manager.

### What hardware does skyzh/tiny-llm need?

It is built for Apple silicon, which the README describes as a practical local environment with one shared memory space and direct access to Metal kernels, so students can inspect the complete path on one machine instead of depending on an expensive CUDA GPU setup. The course is built on MLX arrays and the MLX extension runtime.

## Sources

- [Issues](https://github.com/skyzh/tiny-llm/issues)
- [License: Apache-2.0](https://github.com/skyzh/tiny-llm/blob/main/LICENSE)
- [Project website](https://skyzh.github.io/tiny-llm/)
- [README](https://github.com/skyzh/tiny-llm/blob/main/README.md)
- [skyzh/tiny-llm on GitHub](https://github.com/skyzh/tiny-llm)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/skyzh-tiny-llm
