Model or dataset
gpu-mode/Triton-Puzzles avatar
gpu-mode/Triton-Puzzles

Triton Puzzles: Interactive GPU Kernel Programming Exercises Using the Triton Language

Puzzles for learning Triton

2,618 stars258 forksJupyter NotebookApache-2.0

At a glance

What is it?
Triton Puzzles is a Jupyter notebook with a structured series of increasingly difficult exercises that teach GPU kernel programming in the Triton language, from element-wise operations to Flash Attention. The puzzles run on CPU via a Triton interpreter, so no GPU is needed to learn, and the progression starts from trivial examples and builds to real algorithms.
Who is it for?
Triton Puzzles is well suited for machine learning engineers and researchers who want to understand GPU memory access patterns and kernel programming without immediately dealing with a physical GPU setup. The Triton interpreter handles execution for all puzzles, so a CPU-only machine or a free Google Colab session is sufficient to start.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Why Puzzles Instead of Documentation for Learning Triton

Triton's official documentation describes the language API and compilation model but offers limited interactive exercises. Understanding GPU memory loading and storage requires writing actual kernels and seeing what happens when memory access patterns are wrong or suboptimal. Triton Puzzles takes the approach of teaching through constraints: each puzzle specifies exactly what operations are allowed and asks the learner to implement a target function using only those primitives.

The README describes memory loading and storage as an area where learners have particular difficulty, and the puzzle set is organized to address this directly. Early puzzles use only simple operations to build intuition; later puzzles introduce tiled memory access patterns that are essential for performance on real kernels. The README notes this is the seventh in a series of puzzle repositories by the same author, following gpu-puzzles, tensor-puzzles, autodiff-puzzles, transformer-puzzles, GPTworld, and LLM-Training-Puzzles.

The Triton Language and How It Differs from CUDA

Triton is an open-source language and compiler for writing GPU kernels. The README describes it as an alternative to CUDA that allows coding at a higher level while still compiling to GPU accelerators. Where CUDA requires managing thread indices, warp synchronization, and shared memory explicitly in a C-style syntax, Triton operates on tiles of data and handles many of the lower-level synchronization details through its compiler.

The syntax and semantics of Triton are described in the README as similar to NumPy and PyTorch. A developer familiar with array operations in either library will find the Triton programming model more approachable than raw CUDA. The tradeoff is that Triton's abstraction is leaky in some respects: memory layout, alignment, and block size choices still affect performance significantly, and these effects are one of the things the puzzles are designed to make visible. Triton is a separate open-source project maintained at github.com/openai/triton, not bundled with these puzzles.

Running the Puzzles: Colab, Local Notebook, and the Triton Interpreter

The entire puzzle set is contained in a single Jupyter notebook, `Triton-Puzzles.ipynb`, at the root of the repository. The README provides a Google Colab link that opens the notebook directly in a cloud environment with GPU support. The Colab path is the fastest way to start without any local setup.

The critical design property is that the puzzles do not require a physical GPU. The README states explicitly: "These puzzles do not need to run on GPU since they use a Triton interpreter." The Triton interpreter simulates kernel execution on the CPU, which means the puzzles run correctly on any machine that can install the Triton package. This removes the barrier of GPU access for engineers who are learning Triton on a CPU-only development machine. The interpreter is built on Triton-Viz, a project by Tejas Ramesh and Keren Zhou, which is credited in the README.

The Puzzle Progression: From Memory Copy to Flash Attention

The puzzles start with trivial operations and progress to algorithms that require understanding tiled memory access, masking, and multi-pass computation. The README describes the goal as building from trivial examples to real algorithms like Flash Attention and quantized neural networks.

Flash Attention is a memory-efficient self-attention algorithm that avoids materializing the full attention matrix in GPU memory by computing attention in tiles. Implementing it in Triton requires correctly managing the tiling pattern and accumulation across blocks. Reaching that level from a standing start is the arc the puzzle set is designed to create. Quantized neural networks require understanding how to load and store data types smaller than float32 and how to handle the resulting arithmetic correctly. These are practical skills for anyone writing production inference kernels, not just academic exercises.

What Triton Puzzles Does Not Cover

Triton Puzzles is a learning exercise, not a production kernel library. The puzzles do not cover profiling with Triton or the NVIDIA profiler, kernel autotuning through Triton's `triton.autotune` decorator, multi-GPU coordination, or deployment. An engineer who completes all the puzzles will understand Triton's programming model but will still need to read Triton's own documentation and examples to write production kernels.

The repository contains only the notebook file and a README; there are no solution files, test suites, or CI workflows. Checking whether a puzzle solution is correct depends on the verification cells within the notebook itself. The most recent direct predecessor in the series, cuda-puzzles, uses CUDA's numba-based API instead of Triton, making it a different but related learning path for engineers who prefer to understand CUDA directly rather than through Triton's abstraction layer.

Maintenance and Context Within the gpu-mode Series

The last push to the Triton Puzzles repository was on 2026-04-01. The repository is not archived. There are no GitHub releases; the notebook is updated in place. The gpu-mode organization on GitHub is the current home of the repository, though the README credits the puzzle concept and original design to the author of the broader series.

The Apache 2.0 license permits unrestricted use, modification, and redistribution. The Triton-Viz visualization tool that powers the interpreter is a separate project under its own license; the README credits it but does not reproduce its license terms. Teams building educational materials based on these puzzles should verify the Triton-Viz license independently. The Discord server linked in the README at discord.gg/gpumode has a dedicated `#triton-puzzles` channel for discussion.

Editorial conclusion

Triton Puzzles is well suited for machine learning engineers and researchers who want to understand GPU memory access patterns and kernel programming without immediately dealing with a physical GPU setup. The Triton interpreter handles execution for all puzzles, so a CPU-only machine or a free Google Colab session is sufficient to start. Engineers who have finished cuda-puzzles and want to move to a higher-level language will find the progression natural. The last push to the repository was on 2026-04-01, and the Apache 2.0 license permits unrestricted use.

Frequently asked questions

Is Triton a Python library?

Triton is an open-source language and compiler for GPU programming, not a traditional Python library. The README describes it as an alternative to CUDA that allows higher-level coding and compiles to GPU accelerators. It has a Python-like syntax and integrates with PyTorch.

What is Triton vs CUDA?

The README describes CUDA as a proprietary low-level language for GPU programming and Triton as an open-source alternative that operates at a higher level of abstraction. Triton's syntax is similar to NumPy and PyTorch, while CUDA uses a C-style syntax with explicit thread and warp management.

Does completing Triton Puzzles require a GPU?

No. The README explicitly states that the puzzles do not need to run on a GPU because they use a Triton interpreter that simulates execution on the CPU. Google Colab with GPU support is offered as an option but is not required.

Official sources

  1. gpu-mode/Triton-Puzzles on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/gpu-mode-triton-puzzles.svg)](https://hysenlabs.com/projects/gpu-mode-triton-puzzles)