# ThunderKittens: Tile Primitives for Writing Fast CUDA Kernels on Hopper and Blackwell GPUs

> ThunderKittens is a header-only CUDA framework from the Hazy Research Lab at Stanford that exposes tile-level primitives for writing high-performance deep learning kernels. It targets Hopper and Blackwell NVIDIA GPUs, requires CUDA 12.8 and C++20, and is designed to produce kernels that reach near-theoretical peak throughput without requiring deep GPU architecture expertise.

**HazyResearch/ThunderKittens** — Tile primitives for speedy kernels

- Repository: https://github.com/HazyResearch/ThunderKittens
- Stars: 3,734 · Forks: 331
- Language: Cuda
- License: MIT
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/hazyresearch-thunderkittens

## What ThunderKittens Solves and Who It Is For

Writing fast CUDA kernels for modern GPUs requires coordinating tensor core calls, shared memory layout to avoid bank conflicts, asynchronous memory copies to hide latency, and distributed shared memory for multi-GPU operations. Each of these is error-prone to implement correctly, and a small mistake in any one of them can cut throughput by a large factor.

ThunderKittens is a framework that provides abstractions for each of these concerns at the level of tiles. The README describes its design around a key observation: a GPU is not a large matrix multiply machine but a manycore processor where each core efficiently handles matrix multiplications of roughly 16x16 values. ThunderKittens is built around manipulating tiles no smaller than 16x16 values, which aligns with the hardware's actual structure.

The intended users are ML engineers and researchers at companies or labs that need custom kernels for training or inference at production scale. The README lists Together AI, Jump Trading, and Cursor as examples of organizations using ThunderKittens in production. A separate blog post from Hamza Elshafie provides a deep-dive on the codebase as a starting point for new users.

## Three Design Principles: Simplicity, Extensibility, and Speed

The README opens with three design principles that explain the trade-offs ThunderKittens makes.

Simplicity means the framework exposes a small number of well-chosen primitives. Kernels written in ThunderKittens are concise relative to the functionality they implement. The FlashAttention-3 implementation in the repository is cited as evidence that the framework can reach peak hardware utilization with readable code.

Extensibility means ThunderKittens is embedded natively in CUDA. If a specific operation is not covered by the library, the developer can drop to raw CUDA without the library getting in the way. There is no required class hierarchy or mandatory registration mechanism.

Speed means the library does the right thing at the hardware level. The six specific capabilities it addresses are: calling tensor core functions including asynchronous WGMMA calls on H100 GPUs and TCGEN05 calls on B200 GPUs, managing shared memory without bank conflicts, hiding memory latency through asynchronous TMA copies, using Distributed Shared Memory instead of L2, overlapping work and I/O through a Load-Store-Compute-Finish template, and transferring data over NVLink with NVSwitch acceleration.

## Installation and Build Requirements

ThunderKittens is a header-only library. Installation means cloning the repository and including the main header:

```bash
git clone https://github.com/HazyResearch/ThunderKittens.git
```

Then in your CUDA source file:

```Cuda
#include "kittens.cuh"
#include "prototype.cuh"
```

Set up the CUDA environment variables before compiling:

```bash
export CUDA_HOME=/usr/local/cuda-<YOUR-CUDA-VERSION>
export PATH=${CUDA_HOME}/bin:${PATH}
export LD_LIBRARY_PATH=${CUDA_HOME}/lib64:$LD_LIBRARY_PATH
```

The minimum CUDA version is 12.8. C++20 is required; if compilation fails with unexpected errors, the README instructs updating the compiler:

```bash
sudo apt update
sudo apt install gcc-11 g++-11
sudo update-alternatives --install /usr/bin/gcc gcc /usr/bin/gcc-11 100 --slave /usr/bin/g++ g++ /usr/bin/g++-11
```

For a libc10.so error that sometimes appears with PyTorch, the README provides:

```bash
python -c "import torch; print(torch.file)"
export LD_LIBRARY_PATH=<PRINTED_PATH>/lib:$LD_LIBRARY_PATH
```

As of ThunderKittens 2.0, individual kernels under the /kernels directory must be compiled separately. The top-level setup.py no longer exists. Each kernel directory contains its own Makefile, tests, and benchmarks.

## Data Types and a Matrix Multiplication Example

ThunderKittens provides four data type categories: register tiles, shared memory tiles, register vectors, and shared memory vectors. All are parameterized by layout, element type (bf16, fp16, fp32, or the new MXFP8 and NVFP4 types added in 2.0), and size.

A simple matrix multiplication kernel for H100 demonstrates the structure:

```Cuda
#include "kittens.cuh"
#include "prototype.cuh"

using namespace kittens;
using namespace kittens::prototype;
using namespace kittens::prototype::lcf;

template<int M_BLOCK, int N_BLOCK>
struct matmul_layout {
    using  base_tile      = st_bf<64, 64>;
    using  global_layout  = gl<bf16, 1, 1, -1, -1, base_tile>;
    struct globals        { global_layout A, B, C; };
    struct input_block    { base_tile a[M_BLOCK], b[N_BLOCK]; };
    struct finish_block   { base_tile c[M_BLOCK][N_BLOCK]; };
    struct common_state   { int2 coord; };
```

This is less than 100 lines of code in total, and the README states it achieves approximately 855 TFLOPs on an H100, which is 86% of the theoretical maximum. The educational kernel series at kernels/gemm/educational_h100 walks through this step by step.

## GPU Support: Hopper, Blackwell, Vera Rubin, and the End of Ampere

ThunderKittens 2.0, released January 2026, added full support for Blackwell GPUs along with MXFP8 and NVFP4 precision. Vera Rubin support was added in September 2026, bringing new PTX instructions for Rubin GPUs and additional GEMM kernels.

Ampere GPU support was actively maintained before 2.0. The 2.0 release documentation states that Ampere is no longer actively supported. ThunderKittens may still compile and run on Ampere hardware, but the developers do not plan further Ampere-specific work. Contributions from the community for Ampere support are welcomed by the project.

For AMD GPUs, a sister project called HipKittens is linked from the README, maintained separately.

The project is MIT licensed. Development is led by graduate students at the Hazy Research Lab, though the README notes that many AI companies contribute to production use and the 2.0 release includes major industry contributions.

## NVLink, Distributed Shared Memory, and Multi-GPU Operations

ThunderKittens 2.0 added GPU networking support. The library lets you transfer data over NVLink and use NVSwitch acceleration for fast multi-GPU operations. This capability is listed in the README as one of the six core primitives the framework addresses, alongside tensor cores, shared memory, loads and stores, distributed shared memory, and worker overlapping.

Distributed Shared Memory (DSMEM) allows one warp to directly access shared memory on another SM without going through L2. The README notes that L2 is "so last year" in describing this capability. DSMEM reduces latency for communication-heavy patterns that would otherwise require a global memory round-trip.

The Load-Store-Compute-Finish (LCF) template provides a structured way to overlap memory operations with computation. Instead of waiting for data to arrive before starting the next compute phase, the LCF template schedules loads and stores to interleave with compute passes. This is the mechanism that allows the matmul example to approach 86% of theoretical peak throughput on H100 despite the memory bandwidth requirements.

For teams comparing ThunderKittens with Triton, the Python-based GPU kernel language from OpenAI, the key difference is interface level: Triton compiles Python-like code to GPU machine code and abstracts away hardware layout concerns, while ThunderKittens provides lower-level CUDA C++ control closer to the hardware for finer-grained tuning. Teams comfortable with CUDA who need maximum throughput and direct access to Hopper and Blackwell features will find ThunderKittens' approach more direct; teams that prefer Python and portability will find Triton more accessible. ThunderKittens is not a Python library and has no pip install.

Engineers who need to write kernels for the attention mechanism specifically can look at the FlashAttention-3 implementation in the repository as a reference implementation. The README cites it as evidence that the library can reach near-peak hardware utilization. Kernels in the /kernels directory cover GEMM, attention, and other standard operations, each with its own Makefile for independent compilation. The tests/ and benchmarks reside alongside each kernel's source files, making it straightforward to run a specific kernel's validation independently of the rest of the codebase.

## Conclusion

ThunderKittens is a practical choice for ML engineers who need to write or adapt CUDA kernels for Hopper or Blackwell GPUs and want a framework that handles tensor core calls, shared memory bank conflicts, TMA-based async copies, and NVLink transfers without requiring them to implement those details from scratch. The library does not actively support Ampere GPUs as of ThunderKittens 2.0, so teams on A100 hardware should verify whether the kernels they need still compile and run correctly. Check the CUDA 12.8 and C++20 requirements against your compute environment before committing to it.

## FAQ

### What are ThunderKittens?

ThunderKittens is a header-only CUDA framework for writing fast deep learning kernels on NVIDIA GPUs. It provides tile primitives for tensor core operations, shared memory management, async TMA copies, and NVLink transfers, targeting Hopper and Blackwell GPU architectures.

### Does ThunderKittens support Ampere GPUs?

Ampere GPUs are no longer actively supported as of ThunderKittens 2.0. The library may still work on Ampere hardware but the team does not plan further Ampere-specific development. Contributions for Ampere are welcomed from the community.

### How do I install and use ThunderKittens?

ThunderKittens is header-only: clone the repository and include kittens.cuh in your CUDA source files. No pip install or build step is required for the library itself. CUDA 12.8 and a C++20-capable compiler (gcc-11 or later) are required. Individual kernels under /kernels must be compiled with their own Makefiles.

## Sources

- [HazyResearch/ThunderKittens on GitHub](https://github.com/HazyResearch/ThunderKittens)
- [Issues](https://github.com/HazyResearch/ThunderKittens/issues)
- [License: MIT](https://github.com/HazyResearch/ThunderKittens/blob/main/LICENSE)
- [README](https://github.com/HazyResearch/ThunderKittens/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/hazyresearch-thunderkittens
