Model or dataset
cfregly/ai-performance-engineering avatar
cfregly/ai-performance-engineering

cfregly/ai-performance-engineering: a book companion repository with a commercial profiler in its setup

Code, labs, and resources for O'Reilly AI Systems Performance Engineering: GPU optimization, distributed training, inference scaling, and full-stack tuning.

2,031 stars279 forksPythonApache-2.0

At a glance

What is it?
This repository holds the code, labs and a two hundred item performance checklist for a November 2025 book on AI systems performance. The most instructive file in it is not a lab but a sample environment file, which points at a licensed commercial GPU profiler and its runtime injection library, and the root carries agent configuration and an audit remediation plan alongside the content.
Who is it for?
Read this repository if you want the checklist and the profiling methodology from a systems performance book without buying anything, since the checklist is a plain appendix document and the code is licensed permissively.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 5 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

A companion repository, and what that means for what is actually here

The project is the code, tooling and resources for a published book, one that appeared in November 2025 and covers GPU optimisation, distributed training, inference scaling and full stack tuning for modern AI workloads. That framing matters for what you should expect to find, because a companion repository for a book is not a library with a public interface. It is the worked material that would not fit in the book itself: the code listings that need to be run rather than read, the labs, and the reference tables that are too long to typeset.

The directory listing supports that reading. There is a code directory, a documentation directory, a resources directory, an image directory used by the book cover and the readme, and a meetups directory, which suggests the repository doubles as the archive for a recurring community event under the same name. There is also a gitmodules file, so some of the content does not live in this repository at all and is pulled in as a submodule.

The primary language is recorded as Python, which is consistent with the book's stated emphasis on PyTorch and on writing kernels through a compiler rather than in C++, though the material also includes CUDA C++ examples described as running into the thousands of lines. That is a sensible split for the subject: Python for the framework level where the profiler output has to be interpreted, and a lower-level language for the kernel work itself.

There are no GitHub releases, so there is no version to pin and no changelog to read, which for a repository whose contents track a published edition is a reasonable arrangement. The practical consequence is that the code is whatever is on the default branch, and if the book goes to a second edition, the branch will change underneath you without a version marker telling you it happened.

The checklist is the artefact most likely to outlive the book

One document in the repository is called out separately enough that it is worth treating as the main deliverable. It is a performance checklist of more than two hundred items, living in the documentation directory as an appendix, and it is organised into ten areas that cover the whole lifecycle of a system rather than the inside of a GPU.

The areas, in the order given, are a tuning mindset and cost optimisation section, reproducibility and documentation practices, system architecture and hardware planning, operating system and driver optimisations, GPU programming and CUDA tuning, distributed training and network optimisation, efficient inference and serving, power and thermal management, profiling tools and techniques, and architecture specific optimisations.

Two things about that ordering are worth noticing. Power and thermal management appears as its own area, sitting between serving and profiling, which is unusual in a performance curriculum and reflects an industry that has started counting watts as a first class constraint. And reproducibility and documentation appears second, before architecture, which tells you the author treats the ability to re-measure a result as a prerequisite for the rest of the work rather than as a final step.

The framing given for the checklist is that it captures field tested optimisations and can be applied immediately, and the book's own summary puts it as the last of its points, aimed at reproducing wins and preventing regressions across teams. That last half is the more interesting half. A list of things to try is easy to write and hard to use. A list that is meant to be applied again after every change, so that a fix does not silently regress three months later, is a different kind of document, and it is the one that would still be useful to someone who has never read the book.

The checklist is a plain markdown appendix, which means you can read it without installing anything, and it is the part of this repository with the clearest standalone value.

The chapter order is a debugging search sequence, not a syllabus

The table of contents is the most useful thing in the readme for someone deciding whether this material is relevant, and reading it in order reveals a method.

The first chapter sets out the role, benchmarking and profiling, scaling, resource management, cross team collaboration, and transparency and reproducibility. The second is hardware, moving from the CPU and GPU pairing through tensor cores and the transformer engine, then streaming multiprocessors, threads and warps, then networking, then the interconnect hardware, then multi GPU programming. The third drops straight into operating systems, containers and orchestration, with sections on operating system configuration, the driver and software stack, memory locality and CPU pinning, container runtime optimisations, topology aware scheduling, and memory isolation.

Then the order changes character. The fourth chapter is about communication: overlapping communication with computation, the collective communication library, topology awareness inside that library, data parallel strategies, an inference focused transfer library, and in-network aggregation. The fifth is storage: fast storage and data locality, direct storage access, distributed parallel file systems, a data loading library, and building training datasets. Only in the sixth chapter does the material reach GPU programming itself, with architecture, the thread and block hierarchy, the memory hierarchy, occupancy, and the roofline model. The seventh is profiling and tuning of memory access patterns.

Read as a whole, that sequence is a search order. The author places operating system configuration and container tuning before kernel programming, and networking before storage before the GPU, which is the order you would follow if you believed most reported GPU bottlenecks are actually host or fabric problems. Whether or not you accept that claim, it is a statement about where the author thinks the wins are, and it is the clearest signal in the repository about the book's thesis.

The tools named throughout are specific and current: system and compute profilers from the same vendor as the hardware, the PyTorch profiler, a kernel DSL compiler, and on the serving side a specific list of inference engines and a disaggregated prefill and decode architecture with a paged key value cache. The organising metric is named too, and it is not utilisation: it is goodput, defined in the text as the thing to profile for instead.

The sample environment file names a commercial profiler

There is one file in this repository that tells you more about the constraints on the material than any amount of readme prose, and it is the sample environment file. It is two lines long, and it is commented as a template to copy and fill in with local secrets:

code
# Copy this file to .env and fill in local secrets.
ZYMTRACE_LICENSE_KEY=
CUDA_INJECTION64_PATH=/var/lib/zymtrace/profiler/libzymtracecudaprofiler.so

The first line is a licence key for a commercial profiling product. The second is a path to a shared library loaded into every CUDA context at process start, which is the mechanism CUDA provides for attaching a profiler to a process that was not built to accommodate one.

So at least part of this material assumes you have a licensed commercial GPU profiler installed on your machine, and the setup is done by pointing an environment variable at its injection library. The repository is permissively licensed and free to read, which is entirely compatible with that, since licence terms govern the code in the repository rather than the tools you use to run it. But the practical consequence is real and worth stating plainly: a reader who follows the labs without that product will find that the profiling steps either do not work or produce nothing, and the failure will look like a configuration mistake rather than a missing dependency.

The second line also illustrates something about how GPU profilers work in general. Because the library is injected at load time, profiling a PyTorch script does not require modifying the script. That is a genuinely useful property, and it is also why profiler licensing and profiler activation tend to be environment configuration rather than a line in your training code. Anyone adapting these labs to their own environment should expect the same shape: an injection path, a licence, and a way to confirm the profiler actually attached before trusting an empty profile.

The file is committed with an empty key and a path into a root-owned location, which is the correct way to ship a template and also tells you that the labs are expected to be run on a machine set up for GPU profiling work rather than on a laptop.

Submodules, agent configuration and an audit plan in the repository root

The remaining root entries describe how the project is maintained, which for a companion repository is arguably the more interesting half.

There is a gitmodules file, so at least one piece of the content lives in another repository and is pulled in. For a reader that matters practically: a fresh clone without submodule initialisation will be missing content, and nothing in the directory names will tell you which part. Check the file before assuming a clone is complete.

Then there are four entries that are not what you would expect in a book's companion code. There is a file describing a handoff, a file describing an audit remediation plan, a contributing guide, and a configuration file at the root whose name is the conventional one for instructions aimed at coding agents rather than people. There is also a hidden directory for agent configuration and another hidden directory for a specific assistant tool's configuration.

The presence of an audit remediation plan is the most telling. It is a document type you associate with a security review, and its presence at the root of a teaching repository suggests either that the code was audited at some point and the findings tracked, or that it is a working document kept for ongoing work. Either way, it is the kind of file that tells you the project is being maintained by someone who reviews it, rather than published once and left.

Read together with the book's own stated adoption of AI assisted optimisation, which is one of its listed principles, the agent configuration is not surprising. A repository whose subject is kernel tuning and whose method is to generate and measure candidate changes is an obvious candidate for handing parts of the loop to a tool. The thing to check before trusting generated changes in a profiling context is the one the book itself insists on, which is that you measure the result rather than assuming the edit helped.

Editorial conclusion

Read this repository if you want the checklist and the profiling methodology from a systems performance book without buying anything, since the checklist is a plain appendix document and the code is licensed permissively. Use the code as a starting point rather than a dependency, because there are no releases, the repository pulls in content through submodules, and at least one part of the lab environment expects a commercial profiler whose licence key and injection library are not in the repository. Before running any lab, read the sample environment file and decide whether that profiler is something you have, because a missing licence key is the difference between a lab running and a lab silently collecting nothing. If you are deciding whether to read the book itself, the chapter ordering is the better argument, since it moves from operating system and network tuning to the GPU and then to profiling, which tells you the author treats bottlenecks as a search order rather than a topic list.

Frequently asked questions

What is in the cfregly/ai-performance-engineering repository?

It holds the code, tooling and resources for a published book on AI systems performance, including a code directory, documentation, a resources directory and material for a recurring meetup. The repository has no GitHub releases, so there is no tagged version to pin.

What is the performance checklist that ships with this book?

It is an appendix document of more than two hundred items organised into ten areas covering the lifecycle of a system, from tuning mindset and cost, through reproducibility, architecture, operating systems, GPU programming, distributed training, serving, power and thermal management, profiling and architecture specific work. It is intended to be reapplied after changes to prevent regressions.

Why does the repository ship a sample environment file with a licence key?

The sample environment file is a template for local secrets and it names a commercial GPU profiler, with a licence key variable and a path to the profiler's CUDA injection library. Labs that depend on that profiler will not produce results without it, so a reader should set it up before running the profiling steps.

How is the book organised, and what does the order suggest?

It starts with the engineer's role, benchmarking and reproducibility, then hardware, then operating system, container and orchestration tuning, then networking, then storage, then GPU architecture and occupancy, and finally profiling of memory access patterns. The sequence places host and fabric tuning before kernel programming, which reads as a search order for bottlenecks rather than a topic list.

Does cloning the repository give you all the material?

Not necessarily. A gitmodules file is present, so some content lives in another repository and is pulled in as a submodule, and a fresh clone without submodule initialisation will be missing it. Check that file before assuming you have everything.

Official sources

  1. cfregly/ai-performance-engineering on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/cfregly-ai-performance-engineering.svg)](https://hysenlabs.com/projects/cfregly-ai-performance-engineering)