OpenMLSys: an open textbook on machine learning systems, built with mdBook
《Machine Learning Systems: Design and Implementation》 (V2 is launching soon)
At a glance
- What is it?
- OpenMLSys is a nine-chapter Chinese-language textbook on the machine learning systems stack, from programming interfaces and computation graphs to GPU cluster scheduling. It is a book repository, not a library, and its build depends on Rust and mdBook.
- Who is it for?
- OpenMLSys suits students who already know machine learning theory and want the systems layer, and engineers who need vocabulary for compilers, training systems or GPU scheduling. It is the wrong choice if you want runnable code to import or an English-first text today, since the README describes the English version as just starting.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Activity is slowing. The repository last received commits 6 months ago.
- What is it written in?
- Mainly TeX, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What OpenMLSys is, and who the book is written for
OpenMLSys is a book repository, not a software library. The README describes it as an open project that explains the design principles and implementation experience of modern machine learning systems, covering the full stack from programming interfaces and computation graphs through compilers to distributed training. The primary language of the repository is TeX, and the online edition lives at openmlsys.github.io.
The README names three audiences. Students who have the basic theory of machine learning and want to go deeper into how systems are designed and implemented. Researchers who need to write custom operators or use distributed execution for large models. Engineers responsible for machine learning infrastructure, who need performance tuning and deep customization. That third group is the one most likely to notice a gap: a textbook chapter on operator generation is useful background, but it will not replace the source code of the framework you actually run.
The project also states that a second version is launching, and the repository layout reflects that. There are separate v1 and v2 directories, and the README's table of contents points at v2/zh_chapters paths for all nine chapters. Anyone arriving from an older link to v1 should check which version a chapter belongs to before assuming the content matches what they read elsewhere.
The nine-chapter structure and the stack it traces
The README lists nine chapters for the second edition, and the order is itself an argument about how the stack fits together. Chapter 1 is an introduction to machine learning system architecture and the technology stack. Chapter 2 covers programming interfaces and computation graphs: tensor abstraction, automatic differentiation, graph representation and execution. Chapter 3 moves to AI accelerators and programming, naming GPU architecture and the CUDA, Triton and CUTLASS programming models.
Chapter 4 is the compiler and runtime layer, with IR design, graph optimization, operator generation and runtime execution. Chapter 5 covers data processing: data loading, data pipelines and distributed data processing. Chapter 6 is training systems, including single-node and distributed training, parallelism strategies and training optimization. Chapter 7 is model serving, with inference optimization, online serving and model management. Chapter 8 covers reinforcement learning systems, and chapter 9 covers large-scale GPU cluster management, with GPU scheduling, resource management and large-scale training infrastructure.
The scope is broad, and that breadth is the main trade-off. A single book that starts at tensor abstraction and ends at cluster scheduling cannot go deep on every layer. The chapter list suggests the intended reading is sequential, building a mental model of how a tensor becomes a scheduled job on a GPU, rather than using any one chapter as a reference manual.
Installing the toolchain and building the book locally
The README's build guide lists three environment dependencies: curl, git and Python 3. The book is rendered with mdBook, which is a Rust program, so a Rust toolchain is installed first. The README gives these commands exactly.
# 克隆仓库
git clone https://github.com/openmlsys/openmlsys-zh.git
cd openmlsys-zh
# 安装rust toolchain
curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh
# 安装mdbook
cargo install mdbookThe clone step uses the openmlsys-zh repository name, which is what the README's commands reference even though the project page is openmlsys/openmlsys. After cargo install mdbook finishes, the mdbook binary is on your PATH and the build script can be run.
sh build_mdbook_v2.sh
# 英文版生成结果位于 .mdbook-v2/book
# 中文版生成结果位于 .mdbook-v2-zh/bookThe script produces two output trees: the English edition under .mdbook-v2/book and the Chinese edition under .mdbook-v2-zh/book. The README also states that more detail is in CONTRIBUTING/info_zh.md, which is where to look if the script fails. Note that the repository also carries build_mdbook_v1.sh for the first edition, and requirements.txt pins Python packages including setuptools<81 and sphinxcontrib-bibtex>=2.5.0, with comments explaining that the bibtex pin works around a d2lbook constraint and a Python 3.10 incompatibility. That file matters if you build the v1 pipeline; the README does not describe a pip install step for the v2 script.
Where the book stops being useful
The most obvious limitation is that this is prose and diagrams, not an installable artifact. There is no package to import, no API to call, and the README does not document a CLI for anything other than the build scripts. If your question is why a specific kernel is slow on your hardware, a chapter on operator generation gives you the concepts but not a profiler.
The language situation is the second constraint. The primary language is TeX and the README's chapter table points at zh_chapters paths, so the second edition's chapter content is Chinese. The changelog entry for 2026-03 records a bilingual build architecture refactor and the start of an English version, which means the English edition is in progress rather than complete. Readers who need English material should check the .mdbook-v2/book output before committing to the book.
The version split is a third practical issue. The repository contains both v1 and v2 trees, and the README's changelog shows the content has been adapted over time, including a 2023-05 entry for MindSpore 2.0. A chapter that names a specific framework version can age faster than the surrounding conceptual material. The maintenance picture is also worth stating plainly: the last push to the repository was on 2026-03-15, so treat this as a book that is updated in bursts around edition work rather than a continuously changing codebase.
OpenMLSys compared with a framework's own documentation
The natural alternative is the documentation shipped with a framework itself, for example the MindSpore docs the changelog references, or the CUDA and Triton documentation for accelerator programming. The difference in approach is structural. Framework documentation is organized around that framework's modules, APIs and version history, and it assumes you have already chosen the framework. OpenMLSys is organized around the layers of the stack and the design questions at each layer, so it can describe why an IR is shaped a certain way before showing any particular implementation.
That makes the two complementary rather than competing. A reader who wants to write a custom operator will find the conceptual framing in chapter 4 and the actual function signatures in the framework docs. A reader who wants to understand why data pipelines become the bottleneck in distributed training will find more in chapter 5 than in most API references, which tend to document configuration flags without explaining the pipeline they configure. The trade-off is currency: a book that surveys many systems cannot track every release, while vendor docs are updated with the software they describe.
Licence, citation and what reuse allows
The README states the project is licensed under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International licence, linking to the Chinese deed page. Two of those terms shape reuse. The NonCommercial term means you cannot build a paid course or a commercial training product on the text without separate permission. The ShareAlike term means adaptations must carry the same licence. The repository's LICENSE file is the authoritative text, and the README's badge points at the openmlsys-zh repository for it; the project page itself does not declare a licence identifier. This is a description of what the README says, not legal advice, and anyone planning to redistribute or translate the material should read the licence and, if the stakes are high, consult a lawyer.
The README also provides a citation format for academic use, both as plain text and as BibTeX. The BibTeX entry uses the key openmlsys2022, the title 机器学习系统:设计和实现, the author OpenMLSys Team, and the year 2022, with the URL openmlsys.github.io. Note that the citation year is 2022 while the repository is now on a second edition, so a paper citing this book should check whether the chapter it relies on is the one the 2022 citation refers to.
Editorial conclusion
OpenMLSys suits students who already know machine learning theory and want the systems layer, and engineers who need vocabulary for compilers, training systems or GPU scheduling. It is the wrong choice if you want runnable code to import or an English-first text today, since the README describes the English version as just starting. Verify the licence terms for your use, check whether the chapter you need is in v1 or v2, and confirm the mdBook toolchain installs before cloning the whole repository.
Frequently asked questions
Is there an open source project for machine learning?
OpenMLSys is an open source book project rather than a machine learning library. The README describes it as explaining the design principles and implementation experience of modern machine learning systems, and the online edition is published at openmlsys.github.io.
Can I learn ML in 3 months?
The README positions the book for readers who already have the basic theory of machine learning and want to understand system design and implementation, so it is not an introductory machine learning course. The nine chapters assume that foundation and move into computation graphs, compilers, training systems and GPU cluster management.
Is ML full of coding?
The README names researchers who develop custom operators and engineers who tune infrastructure performance as intended readers, and chapter 3 covers CUDA, Triton and CUTLASS programming models. The book itself is written in TeX and built with mdBook, so contributing to it involves writing prose and running a build script rather than training models.
Is TensorFlow an open source?
The OpenMLSys README does not discuss TensorFlow. It references MindSpore in its changelog and names CUDA, Triton and CUTLASS in the chapter list, so questions about TensorFlow licensing are outside what this material covers.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/openmlsys-openmlsys)