Model or dataset
jax-ml/scaling-book avatar
jax-ml/scaling-book

How To Scale Your Model: Google DeepMind's Open Textbook on LLM Scaling

Home for "How To Scale Your Model", a short blog-style textbook about scaling LLMs on TPUs

1,457 stars209 forksHTMLMIT

At a glance

What is it?
How To Scale Your Model is a free MIT-licensed online textbook from Google DeepMind researchers that explains how large language models run at scale on TPUs, covering parallelism strategies, training and inference arithmetic, and practical configuration trade-offs. The book is written as a Distill-style Jekyll site hosted at jax-ml.github.io/scaling-book and can be read without installing anything.
Who is it for?
The scaling-book is essential reading for ML engineers who work on or want to understand large-scale LLM training and inference on TPUs. The material is opinionated and based on Google DeepMind's working experience with JAX and TPU clusters rather than general hardware abstraction.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 7 days ago.
What is it written in?
Mainly HTML, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What this book is and what it covers

The scaling-book is a short blog-style textbook that the README describes as aiming to demystify the art of scaling LLMs on TPUs. It explains how TPUs work, how large language models actually run at scale, and how to choose parallelism schemes during training and inference that avoid communication bottlenecks.

The book was written by Jacob Austin, Sholto Douglas, Roy Frostig, Anselm Levskaya, Charlie Chen, Sharad Vikram, Federico Lebron, Peter Choy, Vinay Ramasesh, and Albert Webson at Google DeepMind. The README credits James Bradbury and Reiner Pope as the originators of many of the ideas.

The repository holds source for a multi-chapter website. Top-level Markdown files correspond to chapters: tpus.md, transformers.md, sharding.md, training.md, inference.md, applied-training.md, applied-inference.md, roofline.md, profiling.md, gpus.md, and jax-stuff.md. A conclusion.md wraps the book. The index.md file serves as the landing page at the hosted URL.

How the book is organised and what each chapter covers

The chapter structure follows a learning arc from hardware fundamentals to applied configuration.

The tpus.md chapter explains how TPUs work at a hardware level: memory bandwidth, matrix multiply units, and the pod topology that connects multiple devices. Understanding these constraints is the prerequisite for the parallelism chapters that follow.

The transformers.md chapter covers how a transformer model is represented at scale: how the weight matrices map to devices, what the dominant costs are, and how the attention mechanism scales with sequence length.

The sharding.md chapter is the core parallelism guide. It covers tensor parallelism, pipeline parallelism, and data parallelism, explaining when each is appropriate and what communication patterns each introduces. The README frames the central problem as picking parallelism schemes that avoid communication bottlenecks; this is the chapter that addresses that problem directly.

The training.md and inference.md chapters cover the arithmetic of training and inference runs: how to estimate compute requirements, memory usage, and throughput. applied-training.md and applied-inference.md take those foundations into practical configuration decisions.

The roofline.md chapter introduces the roofline model as a tool for identifying whether a given operation is compute-bound or memory-bound. profiling.md covers how to measure actual behaviour on hardware. The gpus.md and jax-stuff.md chapters discuss GPU differences and JAX-specific details.

Reading the book and building it locally

The book is available immediately at https://jax-ml.github.io/scaling-book without any installation. GitHub Pages serves the site automatically on each commit to the main branch through a GitHub Action.

To build and serve the book locally, Ruby, ImageMagick, and Jupyter are required. On macOS, Homebrew installs them:

bash
brew install imagemagick ruby
pip install jupyter

After ensuring Ruby 3.4.5 or higher is on the PATH, clone the repository and use Bundler:

bash
git clone https://github.com/jax-ml/scaling-book.git
cd scaling-book
bundle install
bundle exec jekyll serve

The local build is then accessible at http://127.0.0.1:4000/scaling-book. The Gemfile manages Ruby gem dependencies; bundle install installs them. The site uses a Distill-style Jekyll theme from the al-folio project.

For offline or document-format reading, a conversion script combines all chapters:

bash
python bin/convert_to_single_md.py

This produces scaling-book-combined.md in the repository root. The script strips Jekyll frontmatter, converts liquid tags for figures to standard Markdown images, resolves internal page links to anchor links, converts inline math notation, and strips unsupported LaTeX commands with warnings. The single Markdown file can then be converted to a Word document:

bash
pandoc scaling-book-combined.md -o scaling-book.docx

The parallelism content: what makes it specific to this book

The sharding chapter is the part of this book that is hardest to find explained at this level elsewhere. It covers three parallelism strategies and their interaction with the TPU topology.

Tensor parallelism splits individual weight matrices across devices, reducing per-device memory but introducing all-reduce communication at every layer boundary. Pipeline parallelism partitions layers across devices, reducing communication frequency but introducing pipeline bubbles where devices wait for earlier stages to finish. Data parallelism replicates the model across devices and averages gradients, scaling well to large batch sizes but requiring the full model to fit on a single device.

The book explains how these strategies compose, what the communication volume looks like for each, and what the trade-off is when model size, batch size, or sequence length pushes against a hardware limit. This is the reasoning behind the statement in the README that the book aims to help readers pick parallelism schemes that avoid communication bottlenecks.

The roofline model chapter provides a companion analytical framework. For any given kernel on any given hardware, the roofline tells the practitioner whether more flops or more memory bandwidth is the limiting factor. This shapes decisions about which operations to fuse, which precision to use, and whether a proposed change to the model will help or be bottlenecked by something else.

Contributing to the book and citation

Comments on the published website are powered by Giscus and backed by the GitHub Discussions feature. Pull requests are accepted; contributors must sign a Google Contributor Licence Agreement at cla.developers.google.com before their first contribution.

For academic attribution, the README provides both a plain citation and a BibTeX entry. The plain citation is: Austin et al., "How to Scale Your Model", Google DeepMind, online, 2025. The BibTeX cites the title, all ten named authors, Google DeepMind as publisher, and the hosting URL. The year in the BibTeX is 2025; the README notes the book was originally called "How To Scale Your Dragon" after the Dreamworks film, which explains the dragon imagery on the site.

The README does not describe a formal review process for pull requests beyond the CLA requirement. Corrections and improvements that stay within the existing chapter structure are the expected contribution type; major structural changes would need coordination through a GitHub discussion or issue before a pull request. The .pre-commit-config.yaml file in the repository configures pre-commit hooks, and the .prettierrc file configures Prettier formatting rules used through the prettier Shopify Liquid plugin listed in package.json.

What this book does not cover and when to look elsewhere

The book is TPU-specific in significant ways. The tpus.md and jax-stuff.md chapters discuss hardware that most practitioners do not have direct access to. Engineers working on NVIDIA GPU clusters will find the parallelism and roofline concepts broadly applicable, but the specific memory bandwidth numbers, inter-chip interconnect topology, and JAX compilation behaviour described throughout do not map directly to CUDA or ROCm environments. A GPU practitioner reading this book gets the conceptual framework; they still need GPU-specific references for the numerical details.

The book does not include runnable code for training or serving a model. It explains the arithmetic and the configuration choices, not how to implement them. For executable JAX examples, the JAX and Flax official documentation provides notebooks and scripts that can be run directly.

The book covers LLMs specifically: transformers trained on text at scale. It does not address diffusion models, CNNs, reinforcement learning, or other architectures. The scaling arithmetic and parallelism strategies apply most directly to dense transformer models.

For practitioners who want a broader view of scaling from the software and systems angle, not limited to TPUs or JAX, academic papers from NVIDIA on Megatron-LM and papers from MLCommons benchmarks provide complementary perspectives with GPU-focused implementation details.

Editorial conclusion

The scaling-book is essential reading for ML engineers who work on or want to understand large-scale LLM training and inference on TPUs. The material is opinionated and based on Google DeepMind's working experience with JAX and TPU clusters rather than general hardware abstraction. Engineers working primarily on NVIDIA GPU clusters will find the TPU-specific material less directly applicable, though the parallelism concepts and roofline analysis transfer. The book is not a step-by-step tutorial: it explains mechanisms and trade-offs rather than providing runnable code. For practitioners who want executable JAX examples alongside the theory, the JAX documentation and Flax examples fill that gap. The last push was on 2026-09-22.

Frequently asked questions

What is the Jax scaling book?

The Jax scaling book, officially titled How To Scale Your Model, is a free online textbook written by Google DeepMind researchers that explains how to scale large language models on TPUs. It covers TPU hardware, parallelism strategies, training and inference arithmetic, and the roofline model. It is hosted at jax-ml.github.io/scaling-book.

Can the scaling book be downloaded as a PDF or Word document?

A PDF is not directly provided, but the repository includes a script (python bin/convert_to_single_md.py) that combines all chapters into a single Markdown file. That file can then be converted to a Word document with the command `pandoc scaling-book-combined.md -o scaling-book.docx`.

Does the scaling book cover GPU clusters or only TPUs?

The book is primarily written for TPUs and JAX. It includes a gpus.md chapter, and the parallelism and roofline concepts apply to GPU clusters as well, but the specific hardware numbers and compilation details throughout the book are TPU-specific. The README describes the goal as explaining how LLMs run at scale on TPUs.

Official sources

  1. Issues
  2. jax-ml/scaling-book on GitHub
  3. License: MIT
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/jax-ml-scaling-book.svg)](https://hysenlabs.com/projects/jax-ml-scaling-book)