Model or dataset
stas00/ml-engineering avatar
stas00/ml-engineering

stas00/ml-engineering: an operator's book for training and serving large models

Machine Learning Engineering Open Book

19,062 stars1,253 forksPythonCC-BY-SA-4.0

At a glance

What is it?
The ml-engineering repository is a CC-BY-SA-4.0 open book of scripts, benchmarks and copy-paste commands for people who actually run LLM and VLM training jobs. It is strongest on hardware, network and debugging, and it is not a course, a framework or a job guide.
Who is it for?
Adopt it if you already run multi-GPU or multi-node jobs and need concrete answers about accelerator throughput, inter-node bandwidth or a hung PyTorch process. Skip it if you want a structured course, a salary roadmap or a framework to install instead of reading.
Can I use it commercially?
Yes, with credit. CC-BY-SA-4.0 allows commercial use as long as you credit the authors and indicate what you changed. It is written for creative content, so check how it applies to any code.
Is it still maintained?
Yes. The repository last received commits 6 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

What ml-engineering actually is, and who it is written for

This is a book, not a library. The README describes it as "an open collection of methodologies, tools and step by step instructions" for training, fine-tuning and running inference on large language and multi-modal models, and it states plainly that the material is aimed at LLM and VLM training engineers and operators. The distinction matters. Nothing here is imported into your project. You read it, then paste commands into a terminal.

The content is a first-person record. The author says the know-how came from training BLOOM-176B in 2022, IDEFICS-80B in 2023 and RAG models at Contextual.AI in 2024, and describes the repository as an ongoing brain dump compiled mostly for personal reference. That origin explains both the strengths and the gaps. The parts covering cluster behaviour, NCCL throughput and PyTorch hangs read like notes taken during an incident. The parts covering topics the author has not personally operated are thinner or absent, and the README does not claim otherwise.

If you are looking for an ml engineering roadmap, an interview preparation list or a salary survey, this is the wrong artefact. The README never promises any of that. It promises scripts and copy-paste commands, and that is what the repository layout contains.

How the material is organised into parts and chapters

The table of contents splits into seven parts. Insights covers cloud provider selection and a framework for deciding whether a GPU generation upgrade is worth its cost, worked through an H200 to B200 comparison. Hardware splits into compute, storage and network. Orchestration covers containers and resources, with SLURM as its own chapter. Training, Inference and Development follow, and Development holds the debugging and testing material. Resources closes with an LLM and VLM chronicles collection.

The repository layout mirrors the table of contents almost exactly: compute/, storage/, network/, orchestration/, training/, inference/, debug/, testing/, resources/, plus courses/, insights/, model-parallelism/ and a build/ directory. That one-to-one mapping is the useful part. The README's shortcut list points at three benchmark scripts by path, all_reduce_bench.py under network/benchmarks, torch-distributed-gpu-test.py under debug, and mamf-finder.py under compute/accelerator/benchmarks, so you can go straight to a tool without reading a chapter first.

Two things sit outside the book proper. A SKILL.md file at the repository root is described as something you can use to teach an AI agent to train and operate large-scale models, with companion skills in the author's other repositories. And a courses/ directory contains a lesson-based reading path over the same insights. Neither replaces the chapters; they are alternate entry points into the same body of notes.

Installing nothing: reading the book and running the first script

There is no package to install. The README offers ebook builds hosted on the Hugging Face hub, a PDF and an EPUB, and says they are rebuilt once every few weeks, with build instructions in build/ for anyone who wants the current version. The canonical reading path is the rendered site linked from the repository homepage.

The first practical step is to clone the repository so the scripts are on the machine you are diagnosing.

bash
git clone https://github.com/stas00/ml-engineering.git
cd ml-engineering

From there the README's shortcut list gives the entry points. To check inter-node connectivity, the README points at debug/torch-distributed-gpu-test.py, described as a tool to quickly test your inter-node connectivity. To measure what your accelerator actually delivers rather than what its datasheet claims, it points at compute/accelerator/benchmarks/mamf-finder.py.

bash
python debug/torch-distributed-gpu-test.py

The README does not document the flags for either script, so read the file header before running it on a shared cluster. The network chapter also offers all_reduce_bench.py, which the README presents as an easier way to benchmark network throughput than nccl-tests. If you are on a SLURM cluster, orchestration/slurm/users.md is described as a cheatsheet, and it is the file to open before writing a submission script by hand.

The debugging chapter is the part with the highest hit rate

The PyTorch debugging guide at debug/pytorch.md is described in the README as quick copy-paste solutions for hanging or breaking PyTorch applications. This is the section most likely to pay for itself, because distributed training failures tend to be opaque: a job that stalls at the same step across every rank, or a process that dies with a stack trace pointing inside a collective. The guide is organised around symptoms rather than around API surface, which is the right shape for the problem.

The same chapter carries a technique the README lists separately: making tiny models, tokenizers and datasets so that a full training loop can be exercised in seconds instead of hours. That is a debugging method, not a performance trick, and it is the kind of thing that is obvious in hindsight and rarely written down.

The limitation is that this material is a collection of known cases. If your failure is not one the author has hit, the guide gives you a method and a set of tools rather than an answer. It also assumes you can reproduce the failure, which is not always true on a cluster where the same job succeeds on the next allocation.

Where the book stops being the right tool

The repository is explicit that the know-how came from a small number of very large training runs. That is a narrow base. If you are fine-tuning a 7B model on a single node, the network and storage chapters describe problems you will not have, and the orchestration chapter assumes a scheduler you may not run. The material is not wrong for smaller setups, it is simply aimed elsewhere, and the README does not pretend to cover every scale.

There is also a currency problem that any hardware reference faces. The README links comparison tables for theoretical accelerator TFLOPS, accelerator memory size and speed, and inter-node and intra-node networking. Those tables are only as good as the last time someone checked them against a spec sheet. The last push to the repository was on 2026-09-08, so the tables were current as of that date, but a table of theoretical numbers cannot tell you what a specific driver, CUDA version and interconnect topology will deliver on your cluster. That is exactly why the repository ships measurement scripts alongside the tables.

Finally, the book is one practitioner's experience, stated as such. Where a chapter gives a recommendation, it reflects what worked in those runs. Treat the scripts as reproducible and the opinions as opinions.

How it compares with a structured ML engineering course

The obvious alternative for someone starting out is a structured course, and the repository contains one of its own: courses/lesson-learned is described as a very different way of reading the open books, going over the terse learned insights and letting you dive deeper when needed. That is a reading path over the same notes, not a curriculum with exercises and grading.

The difference in approach is between reference and instruction. A course sequences material and assumes you know nothing at the start. This repository assumes you are already operating a cluster and need an answer now. It is organised for lookup: a shortcut list at the top of the README, symptom-based debugging, benchmark scripts you run against your own hardware. A course will not tell you why your all-reduce is slow on a particular fabric. This will, or it will give you the script that measures it.

If you want a course, the honest comparison is with courses outside this repository. The README does not position the book against any of them, and it does not claim to teach ML engineering from zero.

Licence, maintenance and the cost of keeping up

The repository is licensed CC-BY-SA-4.0, with the licence file at the root as LICENSE-CC-BY-SA. That is a content licence, not a software licence, which fits a book that happens to contain scripts. Share-alike means derivative works carry the same terms, and attribution is required. Whether that fits a company's internal wiki or a paid course is a question for your own legal review; nothing in the README addresses reuse in a commercial setting.

Maintenance is visible in two ways. The last push was on 2026-09-08, nine days before this writing, and the repository is not archived. The README also states that significant updates are announced on the author's Twitter channel, and that the ebook builds are refreshed every few weeks with build instructions in build/ for anyone who wants the latest. The Makefile shows the actual upgrade path: make html regenerates the HTML from chapters-md.txt, make pdf and make epub build from that, and make spell runs codespell with --write-changes over the tree. So a fork that tracks upstream has a defined rebuild, and the cost is running those targets rather than merging code.

The real upgrade cost is not mechanical. It is re-reading the hardware chapters when your cluster changes, because the numbers in the comparison tables describe specific generations and will not transfer to the next one.

Editorial conclusion

Adopt it if you already run multi-GPU or multi-node jobs and need concrete answers about accelerator throughput, inter-node bandwidth or a hung PyTorch process. Skip it if you want a structured course, a salary roadmap or a framework to install instead of reading. Before relying on any figure, open the comparison tables under compute/accelerator, network and storage and check the hardware generation they were measured on, because those tables age faster than the prose around them.

Frequently asked questions

What is stas00/ml-engineering?

It is an open book of methodologies, tools and step by step instructions for training, fine-tuning and running inference on large language and multi-modal models. The README describes it as technical material for LLM and VLM training engineers and operators, containing scripts and copy-paste commands.

Is ml-engineering hard?

The book assumes you already operate training jobs and need specific answers, which is a different starting point from a course. It does not teach ML engineering from zero, and the material is drawn from a small number of very large training runs rather than from a general curriculum.

Is stas00/ml-engineering a course?

Not in the usual sense. The repository contains courses/lesson-learned, described in the README as a different way of reading the open books by going over terse learned insights, but it is a reading path over the same notes rather than a graded curriculum.

How does ml-engineering differ from MLOps material?

The repository's table of contents is organised around hardware, networking, storage, orchestration and debugging rather than around deployment pipelines. Orchestration appears as a chapter covering containers, resources and SLURM, not as a platform to install.

What is the difference between ml engineering and data science?

The book does not define that boundary. Its chapters assume you are the person running the training and inference jobs, covering accelerators, storage, networking, orchestration and debugging rather than analysis or modelling workflows.

Official sources

  1. Issues
  2. License: CC-BY-SA-4.0
  3. Project website
  4. README
  5. stas00/ml-engineering on GitHub
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/stas00-ml-engineering.svg)](https://hysenlabs.com/projects/stas00-ml-engineering)
Community notes

Community notes