# bojieli/ai-infra-book: A Quantitative AI Infrastructure Book You Can Recompute

> Li Bojie's open manuscript derives LLM inference and training system design from hardware limits, and ships the Python tools that let you check every number yourself. It is a draft, and it is a book, not a framework.

**bojieli/ai-infra-book** — 《深入理解 AI Infra：量化分析与系统设计》（李博杰 著）开源书稿：从硬件约束和模型架构出发，量化推导 LLM 推理与训练系统设计。含全书正文、PDF、配套计算工具与实验

- Repository: https://github.com/bojieli/ai-infra-book
- Website: https://bojieli.github.io/ai-infra-book/
- Stars: 5,688 · Forks: 409
- Language: Python
- License: Apache-2.0
- Published: 2026-09-14 · Updated: 2026-09-14 · Language: en
- Canonical page: https://hysenlabs.com/projects/bojieli-ai-infra-book

## What bojieli/ai-infra-book is, and who it is written for

This is the manuscript of 《深入理解 AI Infra：量化分析与系统设计》 by Li Bojie, published under Apache-2.0 and built as a website and a PDF. The README frames it as the companion volume to the author's earlier AI Agent book and states the motivation plainly: after writing that book, he concluded that building good model-based applications requires understanding the infrastructure the model runs on. Most software engineers never write an operating system or a compiler, yet they still study operating systems, compilers and computer architecture, because allocating memory, reading a file and calling a function all carry resource and time costs. The same argument is applied here to latency and cost.

The intended reader is someone who has already called a model API or run a model locally and now wants to know why it is slow and how to make it cheaper. The README also lists systems, network and chip engineers, plus researchers and students. It asks for some Python, linear algebra and computer systems background, with the exact prerequisites in the preface. The reading paths are explicit rather than uniform: chapters 8 and 9 for inference and serving, chapters 5 to 7 for systems and networking, chapters 4 to 7 for chip and architecture work, with chapters 1 to 3 recommended first for everyone.

## From constraints to design: the method running through all twelve chapters

The organising method is derivation from constraints. You fix the task and the quality target, list the compute, storage, communication and dependencies it implies, then compare those against the capacity, bandwidth and throughput of the actual hardware to get an order-of-magnitude estimate. The README is candid about how easily this goes wrong: counting weight reads but forgetting the KV cache, extrapolating from peak compute without checking whether bandwidth can feed it, splitting work across cards while ignoring inter-card communication. Miss any one term and the answer can be off by several times or by orders of magnitude.

The book then asks five questions of every data movement: what is moved, how much, how many times, through where, and who has to wait for it. That framing is what connects the chapters. Chapter 2 covers how attention, historical state and expert structure change system requirements. Chapter 3 covers how task phase, arrival pattern and state lifetime change resource demand. Chapters 4 through 7 work through accelerators, operators and runtimes, supernodes, and datacenter networks. Chapters 8 to 12 apply the same arithmetic to inference optimisation, distributed inference, training systems, scheduling, and device-edge-cloud placement. A second habit is stated as a rule: compute the physical upper bound first, then measure how far the implementation is from it, and treat the gap as either a missing term in the model or removable system overhead.

## Installing the calculations CLI and running a first estimate

The static calculation tools need only Python 3.10 or newer and the standard library. No GPU and no model weights are required. Clone the repository and list the model configurations the tool knows about:

```bash
git clone https://github.com/bojieli/ai-infra-book.git
cd ai-infra-book

# 查看模型支持情况
python3 calculations/calc.py models

# 估算 Qwen3-8B 在 8192-token prefill 下的逐算子资源需求
python3 calculations/calc.py forward --model qwen3-8b --tokens 8192 --format md
```

The first command prints the supported models. The second runs a forward-pass estimate for Qwen3-8B at an 8192-token prefill and writes the result as Markdown, breaking the resource demand down per operator. The README points to calculations/README.md for the tool itself and to calculations/results/README.md for an index of recorded results, so you can compare your output against a stored run before trusting it.

One installation detail matters more than it looks. The repository stores papers and large input and measurement records in Git LFS, roughly 20 GB in total, and .lfsconfig skips all LFS files by default so a clone leaves only pointer files. The prose, figures and static calculations do not need them. To reproduce a specific experiment you fetch only that directory:

```bash
git lfs install
git lfs pull --include="experiments/ch05/05-01/**" --exclude=""
```

Building the reading site locally is a separate path, also Python 3.10+, using website/requirements.txt and scripts/build_site.py, with scripts/check_site.py as a validation step and a --serve flag that previews at http://127.0.0.1:8000. The PDF build is heavier: bash book/build_pdf.sh additionally requires Pandoc, XeLaTeX and fonts, documented in book/README.md.

## The PDF is the reading path, and the README says why

The README recommends downloading the PDF rather than reading the Markdown on GitHub. The stated reason is concrete: the book contains many formulas, tables, footnotes and cross-references, and GitHub's Markdown rendering often shows LaTeX formulas incompletely or misaligned. The PDF is typeset with XeLaTeX, rebuilt automatically after each update to main, and published to Releases, with a latest-download link that always points at the newest build. The table of contents in the README links to the per-chapter Markdown sources instead, which is the right entry point if you want to find a passage and file an erratum.

That split has a practical consequence for anyone citing the book. Releases are dated builds. The three most recent are build-20260914-145529, build-20260914-103601 and build-20260913-230516, all from September 2026, and the repository's last push was on 2026-09-14. Because the manuscript is still being revised, a section number or a figure you quote can move between builds. If you are citing this in your own work, record the release tag and not just the repository URL.

## Where it stops being useful: a draft, not a drop-in system

The README states that the manuscript is still a first draft and under continuous revision. That is the main limitation, and it is not a formality. A book whose whole argument is quantitative is only as good as its numbers, and the contribution guide treats errors as expected: wrong figures, wrong units or magnitudes, citations that do not match the original, faulty formulas and charts are all listed as things worth reporting, through an erratum issue template or a pull request. If you need a settled reference to hand to a team, this is not that yet.

There is a second boundary. This is a book with supporting tools, not a library you integrate. Nothing here serves a model for you. The calculations CLI produces estimates; the experiments under experiments/chXX/XX-YY/ come with run instructions, input conditions and result notes, and the README notes that GPU-based experiments state their hardware and dependency requirements, while readers without the hardware can read the existing records. That is useful for study and for sanity-checking your own capacity planning, and useless if what you actually wanted was a serving stack. The Git LFS setup is a third friction point: a default clone gives you pointer files, so reproducing an experiment means knowing which directory you need and pulling it explicitly.

## How it differs from a framework's documentation or a vendor tuning guide

The obvious alternative is the documentation of whatever serving or training framework you already use. That documentation tells you which flags exist and what they do, and it is authoritative for its own implementation. It generally does not derive why a batch size, a KV cache budget or a parallelism degree is the right one for your hardware, because that derivation depends on your model, your accelerator and your traffic, not on the framework. This book works in the opposite direction: it starts from the hardware envelope and the workload, and arrives at the design question.

A second comparison is a vendor performance guide. Those are typically tied to one accelerator family and one generation, and their numbers age with the hardware. The method here is meant to survive that, since the inputs are capacities, bandwidths and compute rates you supply yourself, and the README explicitly invites you to substitute the model and hardware you actually use and see whether the conclusion changes. The cost of that generality is that you have to do the substitution. Nothing in the repository knows your deployment. The case-studies/ directory, indexed by chapter, is where the book grounds the method in specific models, hardware and systems.

## Licence, third-party material and the cost of keeping up

The original prose, figures and companion code are Apache-2.0, copyright 2026 Bojie Li. The README is explicit that third-party code, fonts, templates and reference material in the repository keep their own copyright and licence terms and are not relicensed to Apache-2.0 by inclusion; the applicable terms are in the relevant directories. If you plan to reuse figures or code, check the directory you are pulling from rather than assuming the top-level licence covers it. This is a description of what the repository states, not legal advice.

On maintenance cost, the repository is not archived and the last push was on 2026-09-14, so the manuscript is being revised rather than frozen. Upgrading means re-reading changed chapters, and the automated PDF release means there is no long-term-support version to pin to. The contribution paths are listed and specific: erratum issues, pull requests for unclear explanations, missing mechanisms, bug fixes in experiments/ and the calculations CLI, improvements to the reading site's layout, navigation, search and mobile experience, and translations, which the README asks you to announce in an issue first so work is not duplicated. The single source of truth for the text is the Markdown under manuscripts/; both the website and the PDF are generated from it, so corrections belong there and not in the built output.

## Conclusion

Adopt this if you already call model APIs or run models locally and want the arithmetic behind latency and cost, or if you work on systems, networking or chips and want to see how hardware limits turn into serving capacity. Skip it if you need a stable reference today: the README states the manuscript is still a first draft and under revision, and the PDF is rebuilt automatically on every push to main, so chapter numbers and figures move between releases. Before quoting anything, pin the release tag you read, since the repository publishes dated builds such as build-20260914-145529, and verify the model configuration and hardware assumptions in calculations/ against your own numbers.

## FAQ

### What does bojieli/ai-infra-book cover?

It is the open manuscript of 《深入理解 AI Infra：量化分析与系统设计》 by Li Bojie, with twelve chapters running from model architecture and accelerator design through datacenter networks, inference optimisation, distributed inference, training systems, scheduling and device-edge-cloud placement. The method throughout is to derive system design from hardware constraints using order-of-magnitude estimates, then compare measured results against the physical upper bound.

### How should I study AI infrastructure with this book?

The README recommends reading chapters 1 to 3 first to understand models and workloads, then choosing a path by interest: chapters 8 and 9 for inference and serving, chapters 5 to 7 for systems and networking, chapters 4 to 7 for chips and architecture. It suggests estimating each worked example yourself before reading the derivation, and substituting your own model and hardware to see whether the conclusion changes.

### Is bojieli/ai-infra-book one of the best books for learning about AI and infrastructure?

That is a judgement the material cannot settle. What it does state is that the manuscript is still a first draft under revision, that it assumes some Python, linear algebra and computer systems background, and that it is the companion volume to the author's earlier AI Agent book. Readers who want a finished reference should weigh the draft status before starting.

## Sources

- [bojieli/ai-infra-book on GitHub](https://github.com/bojieli/ai-infra-book)
- [License: Apache-2.0](https://github.com/bojieli/ai-infra-book/blob/main/LICENSE)
- [Project website](https://bojieli.github.io/ai-infra-book/)
- [README](https://github.com/bojieli/ai-infra-book/blob/main/README.md)
- [Releases](https://github.com/bojieli/ai-infra-book/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/bojieli-ai-infra-book
