TileLang: a tile-level DSL for writing GPU kernels in Python
Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels
At a glance
- What is it?
- Built on TVM with Pythonic syntax, TileLang compiles tile-level kernel code for CUDA, ROCm and Metal, and ships hundreds of examples of real attention and GEMM work.
- Who is it for?
- TileLang is the most interesting option here for kernel authors who want tile-level control without writing raw CUDA, and the example directory is the reason to believe it: DeepSeek MLA, block-sparse attention, Blackwell block-scaled GEMM and variable-shape paths are all ported rather than aspirational. Two things to weigh first.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 17 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 20, 2026, and from our analysis. They are not legal advice.
Editorial analysis
A Python surface over TVM rather than another kernel library
Tile Language, short-named tile-lang, is a domain-specific language for writing high-performance kernels: GEMM, dequantization GEMM, FlashAttention, LinearAttention. The syntax is Pythonic and the compiler infrastructure underneath is TVM. That pairing is the whole pitch. You express a kernel at tile granularity in Python, and the TVM stack handles lowering, so you get Python's iteration speed without giving up the low-level placement control that a kernel actually needs.
The distinction from writing raw CUDA or Triton is worth being precise about. Raw CUDA means owning every register and shared memory decision yourself. Triton means owning block-level choices inside a programming model that decides the rest. TileLang sits at tile granularity as the name implies, which is the level where most of the remaining performance in a modern GPU kernel is won or lost, and the README's example image is a tiled matrix multiplication rather than a toy elementwise loop.
The Python requirement is `>=3.10` and the project classifies itself as beta. The dependency list is short and mostly plumbing: `apache-tvm-ffi` for the runtime interface, `numpy`, `z3-solver` for layout constraints, `ml-dtypes` for the low-precision dtypes, and `torch` plus `torch-c-dlpack-ext` so tensors move between PyTorch and compiled kernels without a copy.
Install shape and what the package actually pulls in
TileLang ships to PyPI as `tilelang`, and installation is documented on tilelang.com rather than in the README, which links an Installation page as the next step. The repository also carries both `pyproject.toml` and a `requirements.txt` that agree on the runtime set.
pip install -r requirements.txtdependencies = [
"apache-tvm-ffi>=0.1.11,<0.1.13",
"torch-c-dlpack-ext; python_version < '3.14'",
"ml-dtypes",
"z3-solver>=4.13.0,<4.15.5",
]The version pins are narrow on purpose. `apache-tvm-ffi` is capped below 0.1.13 and `z3-solver` below 4.15.5, which tells you the FFI boundary and the solver are areas where breaking changes have bitten. `torch-c-dlpack-ext` is skipped on Python 3.14, and the packaging comment explains why it exists: it provides prebuilt torch extensions so that TVM FFI does not have to JIT compile on first import. On macOS the dependency set also requires `torch>=2.4` and `setuptools`.
One inconsistency is worth naming. The repository metadata reports the license as unasserted, while `pyproject.toml` declares `license = "MIT"` and a `LICENSE` file is present in the tree. Those two facts do not agree in form, so read the `LICENSE` file directly if licensing matters to you.
Three backends under one language dialect
The v0.1.13 release replaced a runtime-activated language facade with static per-backend re-exports, which is a refactor with real consequences rather than housekeeping. Before it, choosing CUDA over ROCm over Metal was a runtime decision. After it, each backend has a statically declared dialect, so mistakes surface at import time instead of mid-execution.
The three backends are CUDA, ROCm and Metal. CUDA is the mature path, and the release notes show hardware support moving quickly: native SM75 tensor-core GEMM for FP16, INT8 and INT4 on Turing, an arbitrary-layout TMA lowering path for swizzled shared memory, `T.copy_cluster` for TMA multicast and SM-to-SM cluster transfers, and `tile::gather4` and `tile::scatter4`. Blackwell work includes an SM120 NVF4 block-scaled MMA path for `T.mma_gemm_blockscaled`, and that same release cites roughly 1527 TFLOPS on an 8192 cubed problem measured on SM120 hardware.
ROCm support covers CDNA4 with FP4 E2M1 matrix cores on gfx950, and the v0.1.10 release broadened AMD work along with Blackwell. Metal is the newest and the most interesting outside CUDA: a first `T.gemm` path through `simdgroup_matrix` in May 2026, then Metal 4 cooperative-tensor GEMM for Apple M5 in July, with the simdgroup fallback retained for shapes and systems that cannot use cooperative tensors. There is also an LLVM backend for CPU lowering and execution, added in June 2026.
What the examples directory reveals about real workloads
The `examples/` directory is the most informative thing in the repository, because these are ports of named production workloads rather than tutorial snippets. The listing includes `deepseek_mla`, `deepseek_v3_2`, `deepseek_v4`, `deepseek_deepgemm`, `deepseek_nsa` and `deepseek_mhc`, plus `blocksparse_attention`, `blocksparse_gemm`, `block_causal_attention`, `attention_sink` and `dynamic_shape`.
That list tells you what the language is actually tuned for: attention variants, especially sparse and block-causal attention for diffusion language models, and the DeepSeek MLA family specifically. The README news log records individual optimizations against these kernels rather than generic claims. A top-k selector for sparse attention got a better memory access pattern and roughly 1.9x higher performance in the reported benchmark. The DeepSeek V3.2 sparse MLA backward pass picks its launch width from the head-block size. Block-causal attention came in fixed-length and variable-length forms for diffusion language models.
There is also `bitnet-1.58b`, `convolution`, `elementwise`, `cast` and `dequantize_gemm`, which are the connective tissue rather than the headline. A directory with `deepseek_v4` and `convolution` side by side is a good sign: it suggests the tile abstractions generalise past the attention-shaped problems that drove the early design.
Debugging the compiler rather than debugging your kernel
TileLang has invested noticeably in tools for looking at what the compiler did, which is the right investment for a DSL. The pass visualizer, added in July 2026, is a structure-tree browser for inspecting compiler transformations. The IR lower trace, added the same month, lets you inspect IR changes across every compiler pass and the final code generation step. Pass Diff, from June, compares compiler-pass IR so you can see exactly which pass changed something. Pass timing, from mid-July, profiles passes with a configurable reporting threshold. IKET profiler integration adds CUDA timeline instrumentation.
Compiler diagnostics also carry Python source locations through TIRX into errors, so a traceback points at the line you wrote rather than at generated C. That single feature changes the debugging experience more than any of the visualisers.
The language moved its IR usage to TVM's TIRX representation in May 2026, and v0.1.12 added the LLVM backend, the tile scheduler, a backend CodeGen registry and expanded Blackwell support. If you are following compiler internals, the `benchmark/`, `testing/`, `src/` and `tilelang/` directories plus `format.sh` and a `VERSION` file are where the mechanics live.
Autotuning, cache behavior and the cost of a search
TileLang's stated philosophy is search over heuristics. The parallel autotuning work landed in May 2026, and v0.1.14 reworked `T.alloc_reducer` into deferred reduction epochs whose physical lowering is planned by layout inference, with automatic vectorization of contiguous reducer updates. The same release added an IO-aware cost model for free-mode layout selection and restored register count as the default, with an environment override to switch models.
That override is worth knowing about. It means the default layout cost model is not the only option, and you can change which one runs without editing code. If a kernel's layout choice looks wrong, that switch is the first thing to try.
Compilation speed has been a repeated theme, and the direction of travel is good. The v0.1.14 release reports up to roughly 4x faster cold parallel and AOT compilation, achieved by materializing Z3 solvers lazily and isolating analyzer contexts per kernel compilation. A cross-host CUDA binary cache, added in June, lets compiled binaries be reused across compatible hosts. For a DSL where every edit triggers a recompile, that difference is the difference between an edit loop you keep open and one you avoid.
The v0.1.14 release also added warp specialization schedules and unified TMA copy lowering on CuTe algebra. A separate TileLang LSP was open sourced in August 2026, with inlay hints for buffer shapes, dtypes, scopes and inferred layouts.
Editorial conclusion
TileLang is the most interesting option here for kernel authors who want tile-level control without writing raw CUDA, and the example directory is the reason to believe it: DeepSeek MLA, block-sparse attention, Blackwell block-scaled GEMM and variable-shape paths are all ported rather than aspirational. Two things to weigh first. The `0.1.x` line moves fast enough that v0.1.13 removed several legacy APIs, so read the compatibility notes before pinning a version. And the repository metadata reports the license as unasserted while `pyproject.toml` declares MIT and a `LICENSE` file sits at the root, so check the file yourself before shipping anything derived from it. Start by running one example from `examples/` and then reading the pass visualizer, since the compiler's behaviour is far easier to debug with the tooling than without.
Frequently asked questions
Is TileLang a replacement for Triton?
Not exactly. Both let you write kernels without raw CUDA, but they sit at different granularities. Triton works at block level and lets the runtime decide the rest, while TileLang works at tile level on top of TVM and lowers through the TVM compiler stack. The practical test is whether your kernel needs tile-level placement decisions, which is where TileLang claims its advantage.
Which hardware backends does TileLang support?
CUDA, ROCm and Metal are the three backends, with an LLVM backend for CPU lowering. CUDA support extends from SM75 tensor cores through Blackwell, including SM120 NVF4 block-scaled MMA. ROCm covers CDNA4 with FP4 matrix cores on gfx950. Metal gained a simdgroup GEMM path in May 2026 and Metal 4 cooperative-tensor GEMM for Apple M5 in July, with a simdgroup fallback.
Is TileLang stable enough to pin a version?
Pin it, but read the notes first. The project is on a 0.1.x line and v0.1.13 explicitly removed several legacy APIs and packages, so an upgrade can break working code. Pinning an exact version is the safer habit here, and the repository also publishes a TileLang LSP separately for editor support.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/tile-ai-tilelang)