Open-source project
gpu-mode/lectures avatar
gpu-mode/lectures

gpu-mode/lectures: a GPU kernel course you clone, not sign up for

Material for gpu-mode lectures

6,674 stars671 forksJupyter NotebookApache-2.0

At a glance

What is it?
The gpu-mode lecture repository collects notebooks, slides and code from the GPU MODE YouTube series, covering CUDA, Triton, CUTLASS, NCCL, scan algorithms and Metal kernels. It is a study archive rather than a library, and the README is a lecture index, not a manual.
Who is it for?
Adopt this repository if you already write Python or PyTorch and want kernel-level material with runnable notebooks, and treat it as a reading list rather than a course with a defined order. Skip it if you need a supported library, versioned releases or a guided curriculum: the README is an index of lectures and the material spans many speakers and years.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 22 days ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What the gpu-mode lecture repository actually contains

This is supplementary material for a lecture series, not a software project. The README opens by calling itself exactly that, and the top level of the repository is a set of numbered folders, lecture_001 through lecture_114, plus a single utils.py and a LICENSE file. Each folder holds whatever that session produced: a notebook, a slide deck, a PDF, or nothing at all.

The README is an index of sessions. Lecture 1 is Mark Saroufim on profiling and integrating CUDA kernels in PyTorch, with the notebook and slides in lecture_001. Lecture 3 is Jeremy Howard on getting started with CUDA, with a Colab link. Lecture 29 is Kapil Sharma on Triton internals, lecture 31 is Nikita Shulga on Metal kernels, lecture 17 is Dan Johnson on NCCL, lecture 25 is Haocong Wang on AMD's Composable Kernel. The speakers are named, and each entry points at a folder or an external link.

The audience follows from the topics. Someone who writes PyTorch models and wants to know what happens below the framework, or an engineer who has read the PMPP book and wants worked examples. The README links to Programming Massively Parallel Processors: A Hands-on Approach, and several lectures recap or extend chapters from it. If you have never written a kernel and do not know what a block or a warp is, lecture 3 is the entry point the README points to.

How the lecture folders are organised, and where that breaks down

There is no build system, no package metadata, no test suite and no dependency file at the repository root. The mechanism is plain: one directory per lecture, named lecture_NNN, holding the artefacts for that session. When a speaker produced a notebook, it sits in the folder. When they produced slides, the README links to a Google Slides, Dropbox or Drive URL instead, and the folder may not exist locally at all.

That inconsistency is visible in the top-level listing. Folders are present for lecture_001 through lecture_005, then lecture_008, lecture_009, lecture_011, lecture_012, lecture_013, lecture_014, lecture_017, lecture_018, lecture_025, lecture_029 through lecture_031, and then a long run of higher numbers up to lecture_114. The gaps are sessions whose material lives only behind an external link, or sessions with no published artefacts. Reading the README is therefore mandatory before you go looking for a folder, because the folder you expect may not be there.

The numbering also does not imply a curriculum. Lectures were recorded over time with different speakers, and the README does not claim any prerequisite chain between them. Lecture 28, on Liger Kernel, ships five Colab notebooks on its own, covering RMSNorm correctness and performance, fused linear cross entropy, convergence comparison, contiguity and int32 address overflow. Those are self-contained exercises, not steps in a sequence. Treat the repository as a shelf of material, and pick what matches the problem in front of you.

Cloning the repository and running the first notebook

There is no install step for the repository itself. The README gives no pip command, no environment file and no setup script, because the material is notebooks and slides. What you need is a working Python environment with PyTorch and, for the CUDA lectures, a machine with a GPU. The Colab links in the README exist precisely so that you do not have to provision one.

The README lists the repository under the gpu-mode organization, so the clone target is the project's GitHub path. Once you have it locally, look at what the first lecture folder contains.

bash
git clone https://github.com/gpu-mode/lectures.git
cd lectures
ls lecture_001

The listing shows the notebook and slides referenced by the README for lecture 1, which covers profiling and integrating CUDA kernels in PyTorch. Which files appear depends on the session, so read the entry in the README alongside the directory.

If you would rather not set up a GPU locally, the README points at a Colab notebook for lecture 3, Getting Started With CUDA. Open that link and run the cells in order. For lectures where the README only links slides, there is nothing to execute, and the session is a talk.

When a notebook is local, open it from the repository root so relative paths inside the notebook resolve. The README does not prescribe a particular tool for this, so use whatever notebook front end you already have.

The repository carries no requirements.txt, so the import cells in a given notebook are the only statement of what that session needs. Expect to install PyTorch yourself if it is missing, and expect to adjust for the CUDA version on your machine.

The material is uneven, and that is the main limitation

The README documents no versioning, no compatibility matrix and no maintenance process for the notebooks. The repository's last push was on 2026-09-09, so it is still receiving changes, but a recent push says nothing about whether lecture_012 still runs on a current PyTorch. Kernel APIs move, and a notebook written for one release may fail on another. The repository gives you no way to know which lecture is affected.

Some sessions are code, some are slides, and some are both. Lecture 2, a recap of chapters 1 to 3 of the PMPP book, is a PowerPoint file and a Google Slides link, with no notebook. Lecture 6, on optimizing PyTorch optimizers, is a slides link only. Lecture 20 and lecture 21, both on the scan algorithm, share a single slides URL in the README entries. If you learn by typing along, those sessions give you nothing to run.

This is also the wrong tool for production work. There is no library to import, no API to call and no support commitment. If you need a maintained kernel implementation, you want the project a lecture is about, not the lecture folder. The repository's value is explanatory, and its unit of delivery is a session, not a release.

Finally, the README itself is truncated in places and points outward for a large share of the content. Google Slides, Dropbox and Drive links can rot independently of the repository, and nothing in the repository can detect that.

How it differs from a written GPU textbook

The obvious comparison is the book the README links to, Programming Massively Parallel Processors: A Hands-on Approach. That is a structured text with a fixed order, one voice and a defined progression from thread organisation through memory hierarchy to parallel patterns. The gpu-mode repository is the opposite arrangement: many speakers, many topics, no enforced order, and a mix of notebooks and slide decks.

The difference in approach matters for how you use each. A textbook tells you what to read next. This repository tells you what was covered, and leaves the sequencing to you. The compensation is breadth and recency. The lecture list reaches into territory a single book would struggle to keep current: Triton internals, AMD's Composable Kernel, SYCL on Intel GPUs, Metal kernels, speculative decoding in vLLM, NCCL collectives, CUTLASS, tensor cores, quantized training. Each of those is a session by someone working on it.

A second comparison is the framework documentation for PyTorch or Triton. That documentation is versioned and tested against releases. This repository is not, and does not claim to be. If you need an answer that is guaranteed correct for your installed version, go to the docs. If you want to understand why a kernel is written the way it is, a lecture notebook is often the faster route.

Editorial conclusion

Adopt this repository if you already write Python or PyTorch and want kernel-level material with runnable notebooks, and treat it as a reading list rather than a course with a defined order. Skip it if you need a supported library, versioned releases or a guided curriculum: the README is an index of lectures and the material spans many speakers and years. Before committing time, open lecture_001 and lecture_003 and check that the notebooks still run against your CUDA and PyTorch versions, because the repository carries no release tags and no compatibility statement.

Frequently asked questions

Do I need to install gpu-mode/lectures to use it?

No. The README contains no install command, package or setup script, because the repository holds notebooks and slides rather than a library. You clone it and open the notebooks, or follow the Colab links the README provides for sessions such as lecture 3.

Where do I start in gpu-mode/lectures if I am new to CUDA?

The README points to lecture 3, Getting Started With CUDA by Jeremy Howard, which has a notebook in the lecture_003 folder and a linked Colab version. Lecture 1 on profiling and integrating CUDA kernels in PyTorch is the other early session with a local notebook.

Why are some lecture numbers missing from the gpu-mode/lectures repository?

Not every session produced local artefacts. The README entries for sessions such as lecture 6, lecture 20 and lecture 21 link to external slide decks rather than a folder, so no corresponding directory exists in the top-level listing. The README is the authoritative index of what each session published.

Is gpu-mode/lectures a library I can use in production?

No. It is supplementary material for a lecture series, with no importable package, no release tags and no support commitment. For production kernel work you would use the project a given lecture covers, such as Triton, CUTLASS or NCCL, rather than the lecture folder.

Official sources

  1. gpu-mode/lectures on GitHub
  2. Issues
  3. License: Apache-2.0
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/gpu-mode-lectures.svg)](https://hysenlabs.com/projects/gpu-mode-lectures)