# zero-to-sglang: A Bilingual Course for Building LLM Inference Engines from Scratch

> zero-to-sglang is a hands-on course from Datawhale and RadixArk, the company behind SGLang, that guides Python developers from the core concepts of LLM inference through building a mini SGLang engine to reading the real SGLang source code. Part 0, Part I, and Chapters 1 and 2 of Part II are complete in both English and Chinese; the remaining chapters are still being written.

**datawhalechina/zero-to-sglang** — Official SGLang × Datawhale course on LLM inference (中英双语): understand inference, build a mini-sglang from scratch, then read the real SGLang source and land your first PR. 《从零手搓SGLang》：读懂推理，手搓 mini-sglang，吃透 SGLang 源码。

- Repository: https://github.com/datawhalechina/zero-to-sglang
- Website: https://datawhalechina.github.io/zero-to-sglang/
- Stars: 1,337 · Forks: 103
- Language: Python
- License: not declared
- Published: 2026-09-11 · Updated: 2026-09-11 · Language: en
- Canonical page: https://hysenlabs.com/projects/datawhalechina-zero-to-sglang

## The Gap This Course Fills

Most tutorials on large language models stop at the API level or at prompt engineering. The few that go deeper tend to either stay conceptual without touching code, or drop the reader directly into a production codebase with no explanation of the underlying design decisions. zero-to-sglang occupies the space between those two extremes.

The course starts from first principles: why KV Cache exists, what the difference between prefill and decode actually is, and what it means in practice for a system to be compute-bound versus memory-bound. Those concepts form the foundation for Part II, which walks you through building a mini SGLang implementation from scratch, adding one component at a time: forward pass, generation loop, KV Cache, HTTP serving, Continuous Batching, Paged KV Cache, RadixAttention, and multi-process tensor parallelism.

This approach is suited to engineers who already know PyTorch and want to move from using inference endpoints to reasoning about why those endpoints behave the way they do under load.

## Course Structure and What Is Complete

The course is organized into four parts. Part 0 covers setup: coding ethics, open-source workflow, and deploying a first SGLang server running Qwen3-0.6B on a GPU. Part I is purely conceptual, with no code and no GPU required; it covers LLM architecture, inference basics, GPU architecture, KV Cache derivation, and benchmark metrics like TTFT, TPOT, ITL, and Goodput. Part II builds the mini engine step by step. Part III goes into the real SGLang source to explain advanced optimizations. Part IV covers contributing a PR to SGLang.

As of the latest push on 2026-09-27, the completion status is:
- Part 0: all chapters done in both English and Chinese
- Part I: all five chapters done in both English and Chinese
- Part II: Chapters 1 and 2 done; Chapters 3 through 10 are planned but not yet written
- Part III and Part IV: planned, not yet available

Chapters written by SGLang core members are marked in bold in the syllabus table and include Chapter 2 of Part II, which traces the full lifecycle of a request through SGLang.

The VitePress site at datawhalechina.github.io/zero-to-sglang serves the course content with both English and Chinese versions. The repository itself holds the source markdown, a docs folder for VitePress, and a course-material folder organized by part and language.

## Prerequisites and Hardware Requirements

The README lists explicit prerequisites. Python fluency and basic software engineering habits are required. Familiarity with PyTorch and neural network fundamentals is required. Linear algebra and probability are required; knowing roughly how matrix multiplication and attention are computed is the stated threshold.

GPU programming knowledge is listed as optional: Part I needs no GPU at all, and most of Part II can be debugged on CPU. The full implementation and performance tests in Part II are best done on a GPU, and the README notes that cloud GPU instances are acceptable. The course was developed with CUDA as the GPU backend, but the README does not specify a minimum GPU memory requirement.

For developers who want to run the VitePress documentation site locally, the repository includes a package.json with three scripts:

```bash
pnpm install
pnpm run dev
```

This starts a local VitePress development server. The project uses pnpm 10.21.0 as its package manager, with VitePress 1.6.4 and markdown-it-mathjax3 for math rendering.

The course is affiliated with Datawhale, a Chinese open-source learning community, and with RadixArk, the company founded by the SGLang team. This affiliation means the SGLang-related content is written by people with direct knowledge of the production codebase, which is a meaningful difference from tutorials written by outside observers.

## What the Mini Engine Teaches

Part II is the core of the course. The stated goal is to build a mini SGLang implementation by adding one component at a time, so that each addition has a clear motivation.

Chapter 1 covers the overall architecture of an inference engine: what modules exist, what each does, and how they connect. Chapter 2, written by an SGLang core member, traces a single request from arrival to response through the real SGLang codebase. This is useful context before you write the mini version, because you know what the production system looks like before you build the simplified one.

The remaining Part II chapters are planned but not yet written. They would cover:
- The forward pass and autoregressive generation loop (Chapter 3)
- KV Cache implementation and the reduction from O(n squared) to O(n) attention cost (Chapter 4)
- HTTP serving and concurrent request handling (Chapter 5)
- Continuous Batching and scheduler design (Chapter 6)
- Paged KV Cache and GPU memory management (Chapter 7)
- RadixAttention and prefix caching (Chapter 8)
- Multi-process execution and tensor parallelism (Chapter 9)
- Speculative decoding (Chapter 10)

The choice to build these components incrementally rather than presenting a finished implementation is deliberate. Each chapter is designed to add exactly one new concept, making it possible to see the cost of each addition in isolation.

## Limitations and What the Course Does Not Cover

The most significant limitation at the time of writing is incompleteness. The bulk of Part II and all of Parts III and IV remain unwritten. Developers who need to understand Paged KV Cache, Continuous Batching, RadixAttention, speculative decoding, or how to submit a PR to SGLang cannot get that from this course yet.

The course focuses on SGLang specifically. Developers who want a similar deep dive into a different inference framework, such as vLLM or TensorRT-LLM, will not find that here. The conceptual material in Part I about KV Cache and prefill/decode applies broadly, but the code-level material in Part II is SGLang-specific.

The README does not document how often chapters are added or provide an estimated completion timeline. Readers who need a complete course now should look at the official SGLang documentation and the SGLang paper as supplementary sources, since zero-to-sglang explicitly targets the gap that those sources leave for readers without production inference experience.

The repository has no GitHub releases and no changelog, so tracking progress requires reading the status table in the README or watching the commit history.

## License and Ongoing Development

The repository does not list a license file in the top-level entries shown in the README. The README itself does not state a license. This is a practical concern for anyone who wants to adapt or redistribute the course material: without a license, the default copyright applies and redistribution is not permitted. Readers who need to use the material in a commercial setting should contact the maintainers.

The last push was on 2026-09-27, one day before the review date, which indicates active development. The course is a joint project of Datawhale and RadixArk, both of which have ongoing stakes in the SGLang ecosystem, so continued development is plausible. Part 0 and Part I are stable and can be used now. The decision of whether to start the course before Part II is complete depends on whether the conceptual foundation in Part I is sufficient for your current needs.

## Conclusion

zero-to-sglang is the right resource for Python developers with deep learning basics who want to understand how production inference engines work from the inside. It is not suitable for complete ML beginners: the prerequisites explicitly include PyTorch familiarity and basic neural network knowledge. Before starting, check the course status table in the README, since roughly half of Part II and all of Parts III and IV are marked as planned or in progress; readers who need complete coverage of topics like Paged KV Cache or speculative decoding will find those chapters unavailable at present.

## FAQ

### Does zero-to-sglang require a GPU to follow?

Part I requires no GPU and covers all the conceptual foundations. Most of Part II can be debugged on CPU, but the full implementation and performance tests need a GPU. The README states that a cloud GPU is acceptable.

### Is zero-to-sglang available in English?

Yes. The course is bilingual. Part 0, Part I, and the two completed chapters of Part II are available in both English and Chinese on the course website at datawhalechina.github.io/zero-to-sglang.

### What is the relationship between zero-to-sglang and the official SGLang project?

The course is a joint initiative of Datawhale and RadixArk, the company founded by the SGLang team. Some chapters in Part II are written by SGLang core members. The course explicitly aims to lead readers to the real SGLang source and through the SGLang PR contribution workflow.

## Sources

- [datawhalechina/zero-to-sglang on GitHub](https://github.com/datawhalechina/zero-to-sglang)
- [Issues](https://github.com/datawhalechina/zero-to-sglang/issues)
- [Project website](https://datawhalechina.github.io/zero-to-sglang/)
- [README](https://github.com/datawhalechina/zero-to-sglang/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/datawhalechina-zero-to-sglang
