Library / SDK
mlc-ai/modern-gpu-programming-for-mlsys avatar
mlc-ai/modern-gpu-programming-for-mlsys

Modern GPU Programming for MLSys: a Blackwell kernel book taught through TIRx

A tutorial on modern GPU programming for machine learning systems

1,328 stars146 forksHTMLLicense varies

At a glance

What is it?
A free Sphinx book from the MLC AI team that teaches GPU kernel programming as hardware first, IR second, and state-of-the-art GEMM third. Every example targets sm_100a, so most readers can read it and only a few can run it.
Who is it for?
This is the rare GPU book that starts from the memory and compute engines rather than from a templated GEMM, and the Part IV Flash Attention chapter is a genuine payoff because it reuses the Part III techniques instead of introducing new ones. Read it even if you have no Blackwell card, because the hardware chapters transfer to any recent datacenter part.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository last received commits 8 days ago.
What is it written in?
Mainly HTML, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.

Editorial analysis

A book whose subject is the hardware, not the API

The organising claim is in the first line of the README, and it is a claim about pedagogy. This book teaches modern GPU kernel programming as a progression: understand the GPU hardware, then learn to program it, then write state-of-the-art kernels. Most kernel material in circulation starts somewhere else, from a working CUDA or Triton example that you then disassemble. This one treats the Blackwell-class GPU as the real subject, with its memory hierarchy and Tensor Memory, its tensor-core and asynchronous data-movement engines, warpgroups and clusters, and uses a compiler to reach it.

That ordering has a practical consequence for a reader. The early chapters explain things that experienced kernel authors use by reflex and have never had to name: why a copy engine and a tensor core contend for the same bandwidth, what a memory coalescing failure costs, why the roofline model predicts a ceiling and overlap is how you approach it. If you have written GEMMs without being able to explain those, this is the book that fills the gap. If you already can, Part I becomes a fast reference and the value is concentrated in Part III.

The deployment story is equally plain. Every push to `main` is built and published by GitHub Actions through `.github/workflows/build_deploy.yaml` to a public URL, with a parallel Chinese edition under a `zh/` path and its own Chinese chapter tree in the repository. Both are live. There is no paywall, no signup and no per-version hosting, so you always read whatever is on `main`, which is a genuine advantage for a book in a fast-moving area and a real problem for citing specific pages later.

Why a Python DSL at the IR level rather than another kernel template

The vehicle is TIRx, described in the README as Tensor IR next, a Python DSL for writing GPU kernels at the IR level. It ships as the `tvm.tirx` module of the Apache TVM wheel, which means the book does not ask you to build a compiler and does not ask you to leave Python either.

Programming at the IR level is a deliberate midpoint and worth being precise about what it buys. Writing CUDA means owning instruction selection, register allocation and scheduling, which is total control and a large amount of work. Writing a high-level DSL means owning a tile shape and hoping the compiler does the rest. TIRx asks the author to state loops, layouts, scopes and dispatch explicitly, then hands those to a compiler that makes the scheduling decisions. For the subjects this book covers, that is exactly the right altitude: warp specialization and 2-CTA clusters are scheduling choices, and a DSL where scheduling is expressed rather than implied lets the reader change those choices instead of reverse-engineering them from generated PTX.

One implementation detail has a real effect on how you follow along. The README states that TIRx parses kernel source through Python source inspection, so examples need to live in a file or a notebook cell rather than inside `python -c`. That is not a stylistic note. It means the one-liner you would normally use to smoke-test an import cannot be used to smoke-test a kernel, and it is why the book's verification step is an import check rather than an execution check.

Four parts that build on each other in order

The curriculum is unusually disciplined about dependencies, and the chapter directories in the repository mirror the parts closely enough that you can read the shape of the course without opening the site.

Part I is hardware. Execution and memory model, a performance model built on roofline and overlap, a deep dive into data layout, the memory and compute engines including TMA and Tensor Memory and Tensor Cores, asynchronous coordination, and advanced scheduling through CLC. The matching directories include `chapter_background`, `chapter_data_layout`, `chapter_layout_generations`, `chapter_tma`, `chapter_tmem`, `chapter_tensor_cores`, `chapter_async_barriers` and `chapter_clc`.

Part II is the language, and it is anchored to one runnable artefact rather than to syntax tours. A single-MMA GEMM, developed through scope, layout and dispatch, with a section on how compilation works and another on the tensor layout model covering `TileLayout`, named axes and swizzle. `chapter_intro_tirx` and `chapter_tirx_layout_api` are the directories. Learning a layout system on one working MMA is a better plan than learning it in the abstract, because swizzle only means something once you have watched a shared-memory access pattern go from banked to conflict-free.

Part III is where the book earns its keep. A tiled GEMM built up through TMA pipelining, persistent scheduling, warp specialization and 2-CTA clusters, in `chapter_gemm_basics`, `chapter_gemm_advanced` and `chapter_gemm_async`. Those four techniques are the ones that separate a correct GEMM from a fast one, and they are the hardest to pick up from a paper or a vendor blog because each one assumes the others.

Part IV is Flash Attention 4, described as a complete attention kernel built from the Part III techniques rather than from a new toolkit: two MMAs with softmax between them, online-softmax rescaling, causal masking and GQA, in `chapter_flash_attention`. A closing chapter that recombines rather than introduces is a good sign about the sequence. `chapter_performance` and `tirx_guide` cover measurement, and the appendices hold the language reference, reproducible benchmarking and profiling, compiler internals and asynchronous-kernel debugging, which is the appendix list a kernel author actually needs and rarely finds.

Everything targets sm_100a, and that shapes who can verify it

The README is explicit that the kernels target Blackwell with compute capability `sm_100a`, so running them needs a Blackwell GPU such as a B200, the TIRx compiler and a CUDA build of PyTorch. There is no fallback path and no CPU emulation mentioned.

The setup is three steps. Install the compiler from the Apache TVM wheel:

bash
pip install apache-tvm==0.26.0 cuda-bindings

Then check that the module is present:

bash
python -c "import tvm, tvm.tirx; print(tvm.__version__)"

PyTorch with a CUDA build matched to your GPU comes next, used for the example inputs and the reference checks. Optionally, the reference kernels come from a companion repository pinned to the revision tested against that TVM version:

bash
git clone https://github.com/mlc-ai/tirx-kernels.git
cd tirx-kernels
git checkout 5be39749e7dfd2c4bdae9b4d396f8ec35af07126
pip install -e .

which you then exercise with `python -m tirx_kernels.test --kernel fp16_bf16_gemm`.

Pin that commit hash. It is not incidental detail: the README says the companion revision is the one tested with Apache TVM 0.26.0, and the TVM dependency itself is pinned to exactly `0.26.0` rather than a range. So both halves of the verification chain are frozen, which means published numbers are reproducible only for that pairing. If your GPU is Hopper or Ampere, the compiler may well install and the book may well be readable, but nothing in the examples will produce the numbers, and the TMA, Tensor Memory and 2-CTA cluster chapters describe hardware your card does not have.

This is why the book is worth two different things to two different audiences. For someone with a B200 or an H100 and a need to write real kernels, it is a working path with a reference implementation and a test command. For everyone else it is a hardware and IR textbook with code you cannot execute, which is still worth reading and still cheap.

Building the site, and what the repository does not say

The book is a Sphinx site using Markdown with MyST alongside reStructuredText, and it builds with two commands:

bash
pip install -r requirements-docs.txt
sphinx-build -b html . _build/html

Previewing it locally is one more:

bash
python -m http.server -d _build/html 8000

The README handles the remote case properly, noting that the server runs on the remote machine so you should forward the port with `ssh -L 8000:localhost:8000 user@your-server` and that VS Code Remote SSH does this automatically. `conf.py` sits at the repository root next to `requirements-docs.txt`, which is the standard layout, so any Sphinx extension you need can be added there.

Two gaps are worth naming before you invest time. First, licensing. The repository reports no license, and there is no LICENSE file in the tree. That is not an oversight you can infer your way around: absent an explicit grant, default copyright applies and you have no permission to redistribute, adapt or republish the text or the kernel code. For a book you plan to read privately or cite, that is a non-issue. For a course you teach, a translated edition, or a kernel you copy into a product, it is a blocker, and the answer is to ask the maintainers rather than to assume the Apache license their sibling MLC projects use.

Second, versioning. There are no releases and no tags, so there is no version number to cite and no changelog to check. The commit history is the only history, which is fine for a fast-moving book under continuous deployment and awkward for anyone who needs to say which version they consulted.

One navigational detail before you start. Alongside the numbered `chapter_*` directories the tree carries a separate `tirx_guide/`, which reads as a second route into the same material rather than as an appendix to it. If you already know TIRx and want the layout and scheduling reference specifically, start there and skip Part II. If you are coming from CUDA or Triton and have never met an explicit tensor layout, do not skip it, because Part II is what makes the rest legible and later chapters assume it.

Editorial conclusion

This is the rare GPU book that starts from the memory and compute engines rather than from a templated GEMM, and the Part IV Flash Attention chapter is a genuine payoff because it reuses the Part III techniques instead of introducing new ones. Read it even if you have no Blackwell card, because the hardware chapters transfer to any recent datacenter part. Three things to know first. There is no license file and repository metadata reports no license at all, so the default copyright applies and reuse terms are unclear, which matters if you plan to teach from it or paste its kernels. There are no tagged releases, so the only way to pin a version is to pin a commit, and the last push was on 2026-09-03. And the kernels need a Blackwell GPU, Apache TVM 0.26.0, a CUDA PyTorch build and a companion repo checked out to commit `5be39749e7dfd2c4bdae9b4d396f8ec35af07126`, which means the reference results are tied to one tested combination rather than to a version range.

Frequently asked questions

What is GPU programming?

Writing code that runs on a GPU's parallel processors rather than a CPU's, which means reasoning about warps, threads, shared memory and memory bandwidth instead of cores and caches. At the level this book works at, it also means writing explicit scheduling and layout decisions in an intermediate representation and letting a compiler handle instruction selection and register allocation.

How difficult is CUDA programming?

Hard, largely because performance is never automatic. Correctness is achievable fairly quickly, but reaching good bandwidth or tensor-core utilisation takes control over memory layout, launch configuration and instruction scheduling. Modern architectures add more levers, which is why this book introduces an IR-level DSL instead of asking readers to absorb raw CUDA.

What is the modern GPU architecture?

For datacenter work in 2026 that means the Blackwell generation, which this book targets at compute capability sm_100a. The features that changed kernel design are Tensor Memory for staging operands, TMA for bulk asynchronous copies, tensor cores exposed through single MMA instructions, warp groups, and cluster-level features such as two-CTA cooperation.

Which GPU is best for ML?

For large model training and inference the current generation datacenter parts lead, and this book assumes one of them, naming the B200 as an example and compiling for sm_100a. For smaller models, or for work that fits in one GPU's memory, the answer shifts toward cost and availability rather than peak throughput. The book does not benchmark or rank GPUs, so it does not answer this question on its own.

Official sources

  1. Issues
  2. mlc-ai/modern-gpu-programming-for-mlsys on GitHub
  3. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/mlc-ai-modern-gpu-programming-for-mlsys.svg)](https://hysenlabs.com/projects/mlc-ai-modern-gpu-programming-for-mlsys)