CLI tool
triton-lang/triton avatar
triton-lang/triton

Triton: A Python-Embedded Language and Compiler for GPU Kernels

Development repository for the Triton language and compiler

20,239 stars3,221 forksMLIRMIT

At a glance

What is it?
Triton is a development repository for a language and compiler that let you write custom deep learning primitives in Python instead of CUDA C. Here is how it installs, how its build and caching work, and where it stops being the right tool.
Who is it for?
Adopt Triton if you write GPU kernels that are shaped like tiled tensor operations and you want to stay in Python: pip install triton and the interpreter path let you try kernels without a GPU. Do not adopt it if your workload is a plain elementwise or reduction kernel that an existing library already covers, or if you need a stable LLVM API to build against, since the README states the build will not work at an arbitrary LLVM version.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 5 days ago.
What is it written in?
Mainly MLIR, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Triton is for, and who it is written for

Triton is a language and compiler for writing custom deep learning primitives. The repository describes the aim as an open source environment to write fast code at higher productivity than CUDA, but with more flexibility than other existing DSLs. That sentence is the whole pitch, and it also defines the audience. If you have ever written a fused kernel by hand because the library version was too slow, you are the target user. If you only ever call torch.matmul, you are not.

The design point sits between two extremes. CUDA gives you everything and costs you a lot of code. A fixed operator DSL gives you a short program and takes away control over scheduling. Triton asks you to describe computation over tiles, and the compiler handles the parts that are tedious to write by hand. The foundations are described in a MAPL2019 paper, Triton: An Intermediate Language and Compiler for Tiled Neural Network Computations, which the README asks you to cite if you use the project.

The repository is the development repository, not a curated product. The primary language listed is MLIR, and the top level contains include/, lib/, python/, third_party/, unittest/ and test/ alongside a Makefile and a CMakeLists.txt. That layout tells you what you are signing up for: a compiler project with a Python front end, not a Python library with a small native extension.

How the compiler and the runtime fit together

Triton uses LLVM to generate code for GPUs and CPUs. The build normally downloads a prebuilt LLVM, so a default install does not require you to compile LLVM yourself. That default is the reason pip install triton is a one line operation for most users.

The build is a CMake and Ninja project driven through setuptools. pyproject.toml declares the build requirements as setuptools, cmake between 3.20 and 4.0, ninja, and nanobind pinned at 2.10.2, with setuptools.build_meta as the backend. The Python package is bound to the native code through nanobind, which is why the pinned version matters: nanobind is not a build detail you can float freely.

At runtime there is a cache. The README states that TRITON_HOME changes the location of the .triton directory where Triton's cache is located and downloads are stored during the build, defaulting to the user's home directory, and that it can be changed at any time. That directory is where compiled artifacts accumulate. On a shared machine or a container with a small home volume, this is the first thing to move.

For debugging, the project exposes MLIR pass dumps through an environment variable. The README lists MLIR_ENABLE_DUMP=1 as dumping the IR before every MLIR pass, and points at python/triton/knobs.py for the full set of configuration knobs, which can be set in Python or through environment variables.

Installing Triton and running a first kernel

The README gives a pip install for the latest stable release. Binary wheels are available for CPython 3.10 through 3.14, so check your interpreter version before anything else, because a mismatch here sends you down the source build path whether you wanted it or not.

bash
pip install triton

After that command, the triton package should be importable from the same interpreter. If it is not, the wheel for your platform and CPython version does not exist and you need the source route.

Building from source is a clone followed by a requirements install and an editable install. The README shows the virtualenv variant, which is the one worth copying, since it keeps the build dependencies away from your system Python.

bash
python -m venv .venv --prompt triton
source .venv/bin/activate

pip install -r python/requirements.txt
pip install -e .

The first pip command installs build-time dependencies and the second builds Triton in place. The README notes that passing --no-build-isolation to pip install makes rebuilds faster, because without it every invocation uses a different symlink to cmake and forces ninja to rebuild most of the .a files. If the build runs out of memory, the README suggests setting MAX_JOBS to limit the number of jobs, and TRITON_BUILD_WITH_CCACHE=true enables ccache.

If you want to try the language before touching a GPU, the README points at third-party Triton puzzles which it says can all be run using the Triton interpreter, no GPU required. That is the cheapest first contact with the language. For running the project's own tests, the README gives a one-time make dev-install followed by make test with a GPU or make test-nogpu without one, and warns that the setup step reinstalls local Triton because torch overwrites it with the public version.

Building against a custom LLVM is where the friction lives

The default build downloads a prebuilt LLVM. Overriding that is supported, but the README is explicit that LLVM does not have a stable API, so the Triton build will not work at an arbitrary LLVM version. This is not a warning about a rare edge case. It is a statement about how the project tracks LLVM.

To find the right revision you check the llvm_hash field in cmake/llvm-info.json. The README gives the example value 49af6502c6dcb4a7f7520178bd14df396f78240c and explains that this means the version of Triton you have builds against LLVM 49af6502. You then check out LLVM at that revision and build it.

The manual path sets three environment variables when installing Triton: LLVM_INCLUDE_DIRS, LLVM_LIBRARY_DIR and LLVM_SYSPATH, each pointing into your LLVM build directory. The README also offers make dev-install-llvm as a convenience command that builds LLVM and installs Triton against it. One honest note in the README: after starting the LLVM build, it says to grab a snack, this will take a while. That is the accurate expectation for anyone patching LLVM and rebuilding.

The practical consequence is that the custom LLVM path is for people modifying LLVM itself or targeting a configuration the prebuilt does not cover. If you are not doing either, the prebuilt default is the supported path, and the LLVM version is not something you get to choose.

The test setup assumes you are contributing, not just consuming

The README states plainly that there currently isn't a turnkey way to run all the Triton tests. That is an unusual thing for a project of this size to admit, and it is worth taking at face value rather than reading past it.

The Makefile is described in its own header as not the build system, just a helper to run common development commands, and it asks you to initialize the build system with make dev-install first. The targets are split by concern rather than by convenience. test-lit runs ninja check-triton-lit-tests, test-cpp runs ninja check-triton-unit-tests, and test-unit runs the Python unit suite through python -m triton._test_runner suite unit with --num-gpus and --num-procs arguments. There are separate targets for plugins, gluon, gsan, regression and microbenchmark, and test-interpret runs from python/test/unit.

For anyone evaluating Triton rather than patching it, this matters in a narrow way: the test suite is not a quick sanity check you run after install to confirm the wheel is good. It is a contributor workflow with GPU requirements on several targets and a warmup phase. Budget accordingly, and do not treat a failing test target as evidence that your install is broken until you have read what that target expects.

Where Triton is the wrong choice

The clearest boundary is the one the README draws itself: Triton exists to write custom primitives that are faster or more flexible than what you already have. If your operation is already served by a library kernel at acceptable speed, writing it in Triton adds a compiler, a cache and a build toolchain to your dependency graph for no gain.

The second boundary is portability of the build. Because the build pins a specific LLVM revision and the README states it will not work at an arbitrary LLVM version, Triton does not compose cleanly with an environment that already has its own LLVM requirements. If you are building another MLIR based project in the same image and you need a different LLVM, you have a conflict the project does not resolve for you.

The third is interpreter and platform coverage. Wheels are listed for CPython 3.10 through 3.14. Outside that range you are on the source build, which pulls in cmake, ninja and nanobind, and downloads or builds LLVM. That is a real cost for a project you might only be evaluating.

Finally, the cache is a persistent resource. The .triton directory holds compiled artifacts and build downloads, and it defaults to the user's home directory. In containers with ephemeral or size limited home volumes, this is a failure mode that appears as a full disk rather than a compiler error.

How it differs from writing CUDA or using a fixed kernel library

The comparison the README itself makes is against CUDA, and the stated difference is productivity at a similar level of speed. The mechanism behind that claim is the tiled programming model: you describe computation over blocks, and the compiler decides how to map that onto the hardware. In CUDA the mapping is your job, which is why CUDA kernels are longer and why tuning them is a separate skill.

The second comparison is against existing DSLs, where the README claims higher flexibility. A fixed operator DSL constrains you to the operators it defines. Triton's position is that you can express operations those DSLs do not cover, without dropping to CUDA.

For teams already on a kernel library, the difference is who owns the kernel. With a library you own a version pin. With Triton you own the kernel source, the build toolchain and the LLVM revision. That is a larger surface, and it is only worth taking on when the kernel is the thing you need to control.

Editorial conclusion

Adopt Triton if you write GPU kernels that are shaped like tiled tensor operations and you want to stay in Python: pip install triton and the interpreter path let you try kernels without a GPU. Do not adopt it if your workload is a plain elementwise or reduction kernel that an existing library already covers, or if you need a stable LLVM API to build against, since the README states the build will not work at an arbitrary LLVM version. Before committing, verify three things: that a wheel exists for your CPython version and platform, that the LLVM revision in cmake/llvm-info.json matches what you plan to build against, and that the .triton cache directory has room for the compiled artifacts, which is what TRITON_HOME relocates.

Frequently asked questions

How do I install Triton?

The README gives pip install triton for the latest stable release, with binary wheels available for CPython 3.10 through 3.14. Building from source is a git clone followed by pip install -r python/requirements.txt and pip install -e .

What is Triton?

The repository describes it as a language and compiler for writing highly efficient custom deep learning primitives, intended to give an open source environment to write fast code at higher productivity than CUDA and with more flexibility than other existing DSLs.

How do I use Triton without a GPU?

The README points to third-party Triton puzzles that it says can all be run using the Triton interpreter with no GPU required. The Makefile also provides a test-nogpu target.

Where does Triton store its cache?

The README states that the .triton directory holds the cache and build downloads, defaults to the user's home directory, and can be relocated with the TRITON_HOME environment variable at any time.

Can I build Triton against my own LLVM?

Yes, but the README states that LLVM does not have a stable API and the Triton build will not work at an arbitrary LLVM version. You find the required revision in the llvm_hash field of cmake/llvm-info.json, or use make dev-install-llvm.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/triton-lang-triton.svg)](https://hysenlabs.com/projects/triton-lang-triton)