cuTile Python: writing tile-based GPU kernels from a decorator instead of inline PTX
cuTile is a programming model for writing parallel kernels for NVIDIA GPUs
At a glance
- What is it?
- NVIDIA's Apache-2.0 tile language for Python, where a kernel is a decorated function and the compiler emits Tile IR. It needs driver r580 or newer and skips Hopper entirely, and the README, the repository description and the package metadata each call it something slightly different.
- Who is it for?
- cuTile Python is worth a look if you are writing or tuning GPU kernels in Python and want to stay out of inline PTX and raw CUDA C++, and it is not usable today if your fleet includes Hopper GPUs or your driver is older than r580.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 15 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
A kernel is a decorated Python function
The worked example in the README is a vector add, which is the right way to show a language's structure rather than its performance. A kernel is an ordinary Python function with a decorator, and tiles are loaded, combined and stored with explicit shapes.
import cuda.tile as ct
import cupy
import numpy as np
TILE_SIZE = 16The decorated body reads tiles by index and shape rather than by pointer arithmetic:
@ct.kernel
def vector_add_kernel(a, b, result):
block_id = ct.bid(0)
a_tile = ct.load(a, index=(block_id,), shape=(TILE_SIZE,))
b_tile = ct.load(b, index=(block_id,), shape=(TILE_SIZE,))
result_tile = a_tile + b_tile
ct.store(result, index=(block_id,), tile=result_tile)Launch is explicit, with a grid computed from the array size and the current stream from CuPy rather than from a framework.
grid = (ct.cdiv(a.shape[0], TILE_SIZE), 1, 1)
ct.launch(cupy.cuda.get_current_stream(), grid, vector_add_kernel, (a, b, result))The arithmetic between load and store is written as a normal Python operator, and the compiler is what turns that into instructions on the device. That is the whole idea: you describe a tile computation, not a thread layout. The example depends on CuPy for host arrays and stream handling, and the README notes it assumes a CUDA toolkit of 13.1 or newer is installed.
The sample directory is where the real range shows. It holds MatMul, BatchMatMul, LayerNorm, Transpose, FFT, a mixture-of-experts kernel, fused attention, an all-gather matmul, and both block-scaled and NVFP4-scaled matmul variants, plus a quickstart folder, a templates folder and a shared utilities module. NVIDIA also points at TileGym, a separate repository, as a place for more exercises.
Blackwell and Ampere and Ada, but not Hopper
The hardware support list is short and specific, and this is the single most important fact in the README for anyone planning a deployment.
Kernels are generated as Tile IR, and running that IR requires an NVIDIA driver at r580 or later. The compiler that consumes it, called tileiras at version 13.2, supports Blackwell GPUs and Ampere or Ada generation GPUs. Hopper support is described as coming in later versions. The README links a prerequisites page for the full requirement list.
So the exclusion is specific rather than general: this is not a project that lags the newest architecture, it is a project that has not reached one generation back from the newest. If you have Hopper hardware, the answer today is that the toolchain will not target it.
Two version facts sit in tension and both are worth checking against your own machine. The example's comment asks for a CUDA toolkit of 13.1 or newer, while the compiler named in the same README is version 13.2, and the package metadata requires the tileiras extra at version 13.2 or newer and below 13.5. Three statements, two floors. The Debian package names the README recommends also carry the 13.2 version, which suggests 13.2 is what the supported path actually is.
One project, three one-line descriptions
It is worth putting the project's own summaries side by side, because they describe different products and you will meet all three.
The repository description calls cuTile a programming model for writing parallel kernels for NVIDIA GPUs. The README opens by calling it a programming language for NVIDIA GPUs. The package metadata in pyproject.toml calls the project CUDA Tile Compiler. The documentation site it links is hosted under the CUDA path on the NVIDIA developer documentation site.
A language, a programming model and a compiler are different claims. A language implies you are writing source that is compiled. A programming model is looser and describes an abstraction over hardware. A compiler describes the artifact that comes out, which matches the existence of the tileiras compiler that the tileiras extra installs. None of the three is wrong, and the practical reading is that the language surface sits on top of a compiler that emits Tile IR, which is what the README says it does in the system requirements section.
The same ambiguity shows up in the product's scope. The topics list tile-based programming alongside GPU, kernel and parallel kernels, which is consistent, and the package name on PyPI is cuda-tile, not cutile-python. The repository name and the import name differ from the distribution name, so if you are searching a requirements file for this, look for cuda-tile.
Installing from PyPI, or building a C++ extension in editable mode
The PyPI route is one command, and the extra matters because it decides whether the compiler comes with the package.
pip install cuda-tile[tileiras]The README states that this optional dependency installs the tileiras compiler directly into the Python environment. If you would rather not have it inside the environment, install the bare package and put CUDA Toolkit 13.1 or newer on the system yourself. On Debian there is a middle path worth knowing: instead of the full toolkit, two narrower packages install just the pieces you need.
Building from source is the more interesting path, because this is mostly Python with one C++ extension underneath. The README lists a C++17-capable compiler, CMake 3.18 or newer, GNU Make or msbuild, Python 3.10 or newer with development headers, and the CUDA toolkit. On Ubuntu the first four come from apt.
sudo apt-get update && sudo apt-get install build-essential cmake python3-dev python3-venvThe build is driven from a setup.py that wraps CMake, and the repository also has CMakeLists.txt, a cmake directory, a cext directory and a print_env.sh diagnostic script. The editable install is the documented default for development.
python3 -m venv env
source env/bin/activate
pip install -e .One detail in setup.py pays off immediately. The CMake build is invoked once, and after that a plain make in the build directory recompiles the extension, which is much faster than re-running the install. The build command also exposes switches for disabling the internal extension, building a minimal one, enabling development features, and pointing at a custom NVVM, libdevice or ptxas, which is what you would need if you wanted to test against a different CUDA component build.
The CMake step also downloads DLPack from GitHub during the build. If that is not acceptable in a locked-down environment, the README documents setting an environment variable to a local DLPack source tree to supply your own copy.
A preview namespace, an Apache-2.0 claim, and a licence field that reads NOASSERTION
cuTile ships an explicitly unstable surface and labels it as such, which is unusual and useful. A separate preview package holds APIs that are under active development, they are not part of the stable cuda.tile namespace, and they may change. You install it from a checkout or from a Git subdirectory, and then import from a separate preview namespace.
from cuda.tile_preview import foreign_callSet against that, the package metadata declares itself Development Status 5, Production or Stable. Both statements describe the same project: the stable namespace is meant to be stable, and the preview namespace exists precisely so that unstable work does not have to live inside it. The preview package is a separate install, which is what makes the split enforceable rather than aspirational.
The licensing metadata tells a more confused story. Every source file opens with an SPDX header identifying the copyright holder and naming Apache-2.0, the pyproject metadata sets the licence to Apache-2.0 with the LICENSES folder as the licence files, and the README states that cuTile Python is licensed under the Apache 2.0 licence. The repository's own licence field, however, reads as unrecognised rather than Apache-2.0. So the automated identifier and the files disagree, and the files are the authority: the LICENSES directory at the tree root holds the actual text, with a LICENSE.md file alongside it.
There is also a console script in the metadata named cutile-cache, bound to a cache CLI, which suggests compiled kernels are cached somewhere on disk. That is the kind of detail worth knowing when a rebuild seems not to happen.
Editorial conclusion
cuTile Python is worth a look if you are writing or tuning GPU kernels in Python and want to stay out of inline PTX and raw CUDA C++, and it is not usable today if your fleet includes Hopper GPUs or your driver is older than r580. Start by installing with the tileiras extra so the compiler arrives with the package, run one of the samples such as MatMul.py or AttentionFMHA.py against hardware you actually have, and read the samples README before the API reference, because the sample set is where the intended shape of a kernel is clearest. Two details are worth resolving in your own environment first: the version requirements are stated three different ways in the README, and the repository's licence field reads NOASSERTION even though every file header and the package metadata say Apache-2.0, so take the text from the LICENSES folder rather than from the badge.
Frequently asked questions
What are the differences between cuTile and the CuTe DSL?
cuTile Python is a tile-oriented language where a kernel is a decorated Python function and Tile IR is generated from it, distributed as the cuda-tile package. The CuTe DSL is C++ template metaprogramming, so the practical difference is the host language and where the programming effort sits: Python plus a compiler in this case.
Which NVIDIA GPUs does cuTile support?
The tileiras compiler at version 13.2 supports Blackwell and Ampere or Ada generation GPUs, and the README says Hopper support will come in later versions. Generated Tile IR also requires an NVIDIA driver at r580 or newer.
How do I install cuTile Python?
Use pip install cuda-tile with the tileiras extra to pull the compiler into your environment, or install the bare package and provide CUDA Toolkit separately. On Debian, two narrower packages install the compiler and its dependencies without the full toolkit. Building from source needs a C++17 compiler, CMake 3.18 or newer, and Python 3.10 or newer.
What is the difference between cuda-tile and cutile-python?
cuda-tile is the name on PyPI and the project name in the packaging metadata, while cutile-python is the repository name. The import namespace is cuda.tile, so it is the distribution name you will find in a requirements file.
What licence is cuTile Python under?
Apache 2.0. The source headers carry an SPDX identifier naming Apache-2.0, the packaging metadata declares the same licence with a LICENSES folder, and the README states it directly. The repository's automated licence field reads as unrecognised, so take the text from the LICENSES folder.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/nvidia-cutile-python)