# LeetCUDA: A CUDA Kernel Notebook You Compile Yourself

> LeetCUDA is a GPL-3.0 collection of 200+ CUDA kernels, an HGEMM implementation, and a 400+ page open source book aimed at beginners who want to read real MMA PTX. Here is how the build works, what the benchmarks claim, and where it stops being the right tool.

**xlite-dev/LeetCUDA** — Open sources book with Modern CUDA Learn Notes for Beginners, includes FP16/BF16, FP8, HGEMM, FlashAttention, CuTe, etc.

- Repository: https://github.com/xlite-dev/LeetCUDA
- Website: https://github.com/xlite-dev/LeetCUDA
- Stars: 12,025 · Forks: 1,266
- Language: Cuda
- License: GPL-3.0
- Published: 2026-09-21 · Updated: 2026-09-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/xlite-dev-leetcuda

## What LeetCUDA Actually Is, and Who It Is Written For

LeetCUDA describes itself as an open source book with Modern CUDA Learn Notes for Beginners. The repository holds that book, a PDF build of it, and the kernel sources the book walks through. The README lists Tensor Cores, TF32, F16, BF16 and FP8, 200+ CUDA kernels with PyTorch, an HGEMM directory that the project says reaches 98% to 100% of cuBLAS TFLOPS, and a flash-attn directory implemented with pure MMA PTX.

The audience is narrow and specific. This is for someone who already writes C++ or Python, has a CUDA-capable GPU, and wants to move from calling torch.matmul to understanding what the hardware does underneath it. The kernels are organized as interview preparation material, which tells you the intended reader is preparing for GPU systems roles or trying to close a gap between framework-level and hardware-level work.

It is not a library you import. The README points at a separate project, ffpa-attn, described as a production-ready kernel library, for that purpose. LeetCUDA is the reading and the reference implementations; ffpa-attn is the thing you would ship. Keeping those two roles separate is the right call, and it is worth respecting when you decide whether to depend on this repository.

## How the Kernel Collection Is Organized and Built

The repository is a monorepo of directories rather than a package. Top-level entries include kernels/, docs/, slides/, others/ and third-party/, with the interview material under kernels/interview. The build is driven by a shell script, ./build.sh, which takes an architecture flag and produces a single binary named after the architecture, for example notes_v2_sm120a.bin.

The architecture flag matters more than it does in most projects. The README's own example targets sm_120a, which it associates with Blackwell parts such as the RTX 5090 and the PRO 5000/6000, and it states that this target needs CUDA Toolkit >= 13.2. Because the kernels use MMA PTX and TMA paths, the generated code is tied to the instruction set of the architecture you name. A binary built for one sm_ target is not a portable artifact you hand to a colleague with a different card.

The benchmark path is separate from the build. The same binary takes --bench along with shape arguments, and prints a table of kernels with a maximum error column and a TFLOPS column expressed relative to cuBLAS or cuDNN. That layout is useful for learning: it puts your kernel next to the vendor implementation on the same shapes, so a wrong result shows up as a large Max Err rather than as a silent slowdown.

## Installing LeetCUDA and Running a First Benchmark

The README's Quick Start begins with a clone and a recursive submodule update, because third-party material is pulled in as submodules. If you skip the submodule step the build will fail on missing headers rather than on anything you wrote.

The same block removes any previously installed CUDNN 9 packages for CUDA 13 and installs the versions the benchmarks expect, plus ccache for faster rebuilds. The removal is deliberate: mixing an older libcudnn with the benchmark binary produces misleading comparison numbers. Note that these are apt commands, so this path assumes a Debian or Ubuntu host.

```bash
git clone https://github.com/xlite-dev/LeetCUDA.git && cd LeetCUDA
git submodule update --init --recursive --force && cd kernels/interview
apt remove -y libcudnn9-cuda-13 libcudnn9-dev-cuda-13 libcudnn9-headers-cuda-13
apt install -y cudnn9-cuda-13 ccache
```

Then build for your target architecture. The README gives sm_120a as the example and states CUDA Toolkit >= 13.2 for it. Running the script with --help prints the available build options, which is the fastest way to find the flag for your own card instead of guessing.

```bash
./build.sh --arch sm_120a
./build.sh --help
```

For a first real run, the README's example uses bench mode with an MNK shape for the GEMM kernels and a BHND shape for the attention kernels. On the RTX 5090 example the README records cuBLAS v13.3.0.5-1 at 290T as the baseline, and reports HGEMM with pipe, shared memory and block swizzle at 1.07x over cuBLAS with F16 accumulation. What you should see is a table, not a single number: each row is a kernel variant, with its own maximum error and its own TFLOPS ratio.

```bash
./notes_v2_sm120a.bin --bench --mnk 4096,4096,4096 --bhnd 1,32,16384,128
```

## Reading the Benchmark Table Without Fooling Yourself

The published table is the most interesting artifact in the repository, because it does not flatter the project uniformly. On the RTX 5090 example, FA2 MMA Stages with F32 accumulation comes in at 0.75x and 0.81x of the cuDNN SDPA baseline, while the F16 accumulation variants of the same kernel reach 0.99x and 1.14x. The gap is not a bug; it is the cost of accumulating in F32, and the table makes that visible instead of hiding it behind a single headline figure.

The Max Err column is what makes the comparison honest. F16 accumulation rows show errors around 1.831e-04, F32 accumulation rows around 1.526e-05. A reader who only looks at TFLOPS will pick the F16 path everywhere and then wonder why numerical results drift. The table is effectively arguing that you should read both columns together.

The best numbers in the README come from the FA3-style path with two consumer warpgroups and TMA, reported at 1.37x over cuDNN with F16 accumulation. The Split-D variant for large head dimensions is the other standout: at D=320 with Sk=2, Sv=2, the README reports 182.6 TFLOPS against 83.1 for cuDNN SDPA, a 2.20x ratio, with F32 accumulation and the same 1.526e-05 error as the other F32 rows. That is the strongest claim in the document. Treat it as the project's own measurement on its own hardware, not as a number you will reproduce on a different part.

## Where LeetCUDA Is the Wrong Tool

The first limitation is portability. Everything here is built per architecture and depends on MMA PTX and TMA, so the artifacts do not travel. If you need one binary that runs across a fleet of mixed GPUs, this is the wrong starting point.

The second is that it is a teaching repository with a book attached, not a supported library. The release history shows v4.0.0 through v4.0.2 within a few days of each other, which is the cadence of active writing, not the cadence of an API contract. There is no statement in the README about backward compatibility for the kernel interfaces, and the README does not document rollback or a supported-version policy. If your code imports these kernels directly, an upstream change to a kernel signature is your problem to absorb.

The third is the licence. GPL-3.0 is a copyleft licence. For a personal study repository that is irrelevant. For a product that links this code, it is a decision your legal team has to make, and this article cannot make it for you. The README does not offer an alternative licence for the kernels.

Finally, the environment assumptions are specific: apt for the CUDNN packages, a recent CUDA Toolkit, and a GPU whose architecture the build script recognizes. On a machine that fails any of those, the Quick Start will not get you to a running binary.

## LeetCUDA Versus the Other CUDA Learning Routes

The obvious alternative is the vendor path: NVIDIA's own CUDA samples and the CUDA C++ Programming Guide. Those are broader in coverage and track the toolkit release cycle, but they are written as reference documentation. They will tell you what a TMA descriptor is; they will not hand you a FlashAttention implementation with a benchmark table next to it. LeetCUDA's difference is that every concept in the book has a runnable kernel and a measured number attached.

Another route is GPU Puzzles, which the related searches surface alongside this project. Puzzle-style material gives you small, self-contained exercises with immediate feedback and no build system to fight. LeetCUDA asks more of you up front: submodules, an architecture flag, a CUDNN install. In exchange you get to see an HGEMM that the project says reaches 98% to 100% of cuBLAS, and attention kernels written in MMA PTX rather than in a framework's operator layer. If you want short exercises, the puzzle route is faster. If you want to read a full kernel and then measure it against the vendor, this repository is the more direct path.

A third option is to stay at the PyTorch level and profile. That is cheaper, but it cannot answer the question the book is built around, which is what the SASS and the tensor core instructions are actually doing.

## The Agent Skill, the PDF, and the Maintenance Picture

Two distribution choices are worth noting. First, the book ships as a PDF, linked from the README as LeetCUDA.pdf with 400+ pages, so you can read it without cloning anything. Second, the repository includes a SKILL directory, kernels/interview/book/skills/leetcuda-cpp-kernel/, which the README describes as reusing the knowledge and examples from the book and repository with coding agents such as GitHub Copilot, Claude Code and Open Code. That is a genuine design decision: the same material serves a human reader and an agent that can pull kernel examples into a coding session.

On maintenance, the repository is not archived and the last push was on 2026-09-21, the same day as the v4.0.2 release. That is a current codebase. It is still worth being precise about what that means: an active push history tells you someone is working on it, not that any particular kernel interface will survive the next release.

The upgrade cost is concentrated in the toolchain. Each new CUDA Toolkit or CUDNN major version can invalidate the benchmark baselines printed in the README, because those baselines are pinned to specific versions such as cuBLAS v13.3.0.5-1, cuDNN v9.25.0.15 and PyTorch v2.11. If you are using the numbers as a reference point, record the versions you measured against, because the README's table will be regenerated against newer ones.

On licence: GPL-3.0 governs the repository contents. The README does not state a separate licence for the book text or the PDF, and it does not describe a commercial exception. If you intend to redistribute anything from here, read the LICENSE file at the repository root.

## Conclusion

Adopt LeetCUDA if you are learning CUDA seriously and want HGEMM, FlashAttention and FP8 material you can compile and read, and if you accept that the build is architecture specific: pick your sm_ target first, then run ./build.sh --arch and the resulting notes_v2 binary. Skip it if you need a supported library with a stable API and a release cadence you can plan around, or if you cannot use GPL-3.0 code. Before committing, verify that your CUDA Toolkit version meets the minimum the README states for your target, for example CUDA Toolkit >= 13.2 for sm_120a, and check the LICENSE file in the repository rather than this article.

## FAQ

### What is LeetCUDA and who is it for?

LeetCUDA is an open source book with Modern CUDA Learn Notes for Beginners, packaged with 200+ CUDA kernels, HGEMM and flash-attn implementations that use Tensor Cores and pure MMA PTX. It targets readers who already program and want to work at the kernel level rather than through a framework.

### How do I install and build LeetCUDA?

The README's Quick Start clones the repository, runs git submodule update --init --recursive --force, installs the CUDNN 9 packages for CUDA 13 plus ccache, then builds with ./build.sh --arch followed by an architecture such as sm_120a. The README states CUDA Toolkit >= 13.2 for that target.

### Does LeetCUDA work on any NVIDIA GPU?

No. The build takes an explicit architecture flag and the kernels rely on MMA PTX and TMA, so the resulting binary is tied to the sm_ target you name. The README's example targets sm_120a for Blackwell parts such as the RTX 5090 and the PRO 5000/6000.

### What licence does LeetCUDA use?

The repository is licensed GPL-3.0, and the README does not state a separate licence for the book text or the PDF, nor a commercial exception. Anyone planning to redistribute the code should read the LICENSE file at the repository root.

### Is there a production-ready library from the same project?

The README points to a separate project, ffpa-attn, described as a production-ready kernel library for fast and memory-efficient exact attention in BF16, FP16, FP8 and FP4 for large head dimensions. LeetCUDA itself is the book and reference kernel collection.

## Sources

- [License: GPL-3.0](https://github.com/xlite-dev/LeetCUDA/blob/main/LICENSE)
- [Project website](https://github.com/xlite-dev/LeetCUDA)
- [README](https://github.com/xlite-dev/LeetCUDA/blob/main/README.md)
- [Releases](https://github.com/xlite-dev/LeetCUDA/releases)
- [xlite-dev/LeetCUDA on GitHub](https://github.com/xlite-dev/LeetCUDA)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/xlite-dev-leetcuda
