TileKernels, TileLang kernels for MoE routing and low-precision quantization
A kernel library written in tilelang
At a glance
- What is it?
- TileKernels collects GPU kernels for the operations that dominate mixture-of-experts inference and low-precision training, written in TileLang and wrapped in PyTorch autograd functions. It is a reference implementation set rather than a finished library, and the README says so, and the hardware floor of SM90 or SM100 with CUDA 13.1 decides whether you can run it at all.
- Who is it for?
- Use TileKernels as a reference when you are writing your own kernels for mixture-of-experts routing or low-precision quantization on Hopper or Blackwell hardware and want working code and PyTorch baselines to check against. Do not adopt it as a production dependency on a 0.x line with no tagged releases and no changelog, and do not expect documented speed figures, because the README publishes none.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 2 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What TileKernels is: a kernel collection, not a framework
The operations this project covers are not arbitrary. They are the ones that sit on the critical path of a mixture-of-experts model and of low-precision inference: deciding which experts a token goes to, moving tokens to those experts and back, and casting activations into the narrow formats the hardware prefers. Writing those in raw CUDA is slow to iterate on and easy to get subtly wrong, so the project takes the DSL route.
TileLang is the substrate. The README describes it as a domain-specific language for expressing high-performance GPU kernels in Python, with easy migration, agile development and automatic optimization, and TileKernels is built on it. The value proposition is iteration speed: a kernel is Python you can edit and run, and the compiler does the tuning work that would otherwise be a week of manual autotuning.
What the repository is not is a training framework. There is no model, no trainer and no inference server. It is a set of kernels grouped by purpose, plus PyTorch reference implementations to check them against, plus a test harness. That makes it the right shape for a research group publishing what it has already used internally, and the wrong shape for a team that wants to install something and move on, which is the tension the README itself acknowledges in the next section.
The citation block at the bottom of the README names seven authors and a 2026 year, which tells you this is a recent, small, lab-internal release rather than a long-lived community project.
The hardware floor: SM90, SM100, CUDA 13.1 and PyTorch 2.10
The requirements list is short and unforgiving, and it is the first thing to check.
You need Python 3.10 or higher, PyTorch 2.10 or higher, TileLang 0.1.9 or higher, an NVIDIA GPU on the SM90 or SM100 architecture, and the CUDA Toolkit at 13.1 or higher. In the manifest those become `requires-python` of `>=3.10` and dependencies on `torch>=2.10` and `tilelang>=0.1.9`.
The architecture requirement is the real filter. SM90 and SM100 correspond to the most recent generation of data centre parts, so Ampere and anything older are out, and so is most of the older GPU fleet that inference code is still deployed on. If your cluster is A100 or L4, none of this runs and no amount of configuration changes that.
The TileLang floor deserves a second look. `tilelang>=0.1.9` is a pre-1.0 dependency, which means TileLang does not promise compatibility across minor versions, and a permissive lower bound in this project does not protect you from a TileLang release that changes behaviour. Pin TileLang alongside TileKernels if you care about reproducibility.
CUDA 13.1 is the third constraint, and it is the one that most often ends an evaluation. A CUDA toolkit that recent is not present on many shared cluster images, and mixing a system CUDA with a pip-installed PyTorch build is a known source of hours lost. The project declares the requirement and leaves the reconciliation to you.
Also worth noting from the manifest: the classifier is `Development Status :: 3 - Alpha`. That is the author's own statement about stability, and it matches the README's language.
Installing a tagged release or a working copy, and a name that changes
There are two install paths and the README distinguishes them clearly.
For a released version from the package index:
pip install tile-kernelsFor a local development version, which is what you want if you intend to read or modify the source:
pip install -e ".[dev]"The extras matter for the second one. The `dev` extra in `pyproject.toml` adds `setuptools`, `wheel`, `setuptools-scm>=8`, `pytest`, `pytest-xdist` and `pytest-repeat`, so an editable install brings the test tooling with it. Without the extra you have kernels and no way to run the test suite that validates them.
There is a naming wrinkle that will trip up a scripted install. The package on the index is `tile-kernels` with a hyphen, the project name in the manifest is `tile_kernels` with an underscore, and the importable package directory is `tile_kernels/`. So the pip requirement and the import name differ, which is normal but worth writing down before you start grepping for a `tile-kernels` directory that does not exist.
The version itself is derived, not written. The build requires `setuptools-scm>=8`, the version is declared dynamic, and the resolved value is written to `tile_kernels/_version.py`. In other words, the version comes from git tags at build time. Since the repository has no GitHub releases listed and no `CHANGELOG.md` at the top level, whatever `pip install tile-kernels` gives you is an opaque tag with no published notes. Pin an exact version and record it somewhere, because there is nothing in the project that will tell you later what changed.
tile_kernels/torch is the correctness argument, and pytest is how you use it
A kernel library's central problem is not performance, it is being wrong in a way that produces plausible numbers. This project addresses that structurally rather than aspirationally: there is a `torch/` directory in the package holding PyTorch reference implementations, so every kernel has a plain PyTorch counterpart to be compared against.
That reference is what makes the test commands in the README worth running. The default run checks correctness only:
pytest tests/transpose/test_transpose.py -n 4 # Correctness only with 4 workers
pytest tests/transpose/test_transpose.py --run-benchmark # Correctness + Benchmarking
TK_FULL_TEST=1 pytest -n 4 --count 2Three things are encoded there. The `-n 4` is `pytest-xdist` parallelising across workers, which matters because these are large-tensor tests on a single GPU. The `--run-benchmark` flag switches the same test from verifying against PyTorch to also timing it, so correctness and performance are measured in one pass rather than in two divergent setups. And the third line, with `TK_FULL_TEST=1` and `--count 2` from `pytest-repeat`, is the pressure test: run everything twice rather than the fast subset.
The honest limitation is that the README publishes no results from any of this. There is no benchmark table, no hardware description for a published figure, no shapes, no sequence lengths, no batch sizes, and no baseline comparison. The performance claim in the README is qualitative, that most kernels approach the limit of hardware performance in terms of compute intensity and memory bandwidth, and that some have been used in internal training and inference scenarios. You can reproduce the measurement on your own hardware with `--run-benchmark`; you cannot check the authors' numbers, because they are not written down.
For a user planning to evaluate a kernel, that ordering matters: get the correctness suite green first, then run the benchmark yourself, and treat the README's characterisation as a starting hypothesis.
Seven kernel families, from MoE routing to Sinkhorn normalisation
The feature list is organised by the part of the model being served, and it is worth reading as a map of where the current bottleneck discussions are.
Gating covers top-k expert selection and scoring for mixture-of-experts routing, which is the decision that determines which experts a token visits. MoE Routing is the movement around that decision: token-to-expert mapping, fused expansion and reduction, and weight normalization, which are the permute-style operations that dominate MoE inference time. Transpose is batched transpose, unglamorous and frequently the limiting factor when layout conversions sit between two kernels.
Quantization covers per-token, per-block and per-channel casting to FP8, FP4 and E5M6, including fused SwiGLU plus quantization operations so the activation function and the narrowing happen in one pass instead of writing an intermediate tensor to memory. The README does not define the E5M6 layout, so if that format matters to your hardware, the hosted TileLang documentation is where to look.
The last three are more specialised. Engram provides gating kernels with fused RMSNorm, forward and backward passes and weight gradient reduction, which is training-side rather than inference-side. Manifold HyperConnection covers hyper-connection kernels including Sinkhorn normalisation and mix splitting and application, a mechanism for routing information between layers that is not part of a conventional residual stack. Modeling sits on top of all of it as high-level `torch.autograd.Function` wrappers that compose the low-level kernels into trainable layers, specifically the engram gate and the mHC pipeline.
That last layer is the reason to look at the project even if you only intend to read the code. A kernel you can read is a kernel you can copy, and the autograd wrapper tells you how the author intended the pieces to be composed.
What the README admits, and what the tool configuration does not cover
It is worth quoting the README's own assessment, because it is unusual and it tells you how to read everything else on the page. The project says that most kernels approach the limit of hardware performance with respect to compute intensity and memory bandwidth, that some have already been used in internal training and inference scenarios, and then that they do not represent best practices and that work is under way on code quality and documentation.
That last clause is the important one. This is a set of working kernels published for reuse, not a maintained library with a compatibility promise, and the manifest's Alpha classifier agrees. Combined with the absence of GitHub releases, of a changelog and of any published benchmark numbers, the correct reading is: study it, run its tests, and vendor what you need rather than depend on it.
The tooling configuration supports that reading. The `ruff` section in `pyproject.toml` sets a line length of 150 and selects exactly one rule, `Q000`, which is the flake8-quotes check for single-quoted strings, and the `inline-quotes` setting is set to single. There is no type checker configured, no formatter, no static analyser beyond that single rule, and no coverage measurement. For a research kernel library that may be a deliberate choice, since a strict style gate on experimental GPU code is its own kind of friction, but it does mean the code quality is exactly what the README says it is.
The repository layout is small and legible by comparison: `.editorconfig`, `.gitignore`, `LICENSE`, `README.md`, `pyproject.toml`, `tests/` and `tile_kernels/`. There is no CI configuration directory in the top-level entries, no examples directory and no documentation site of its own, so the README is the whole of the user-facing text and the linked TileLang repository is the actual documentation.
TileKernels against hand-written CUDA and against PyTorch alone
Two alternatives, with genuinely different costs.
The first is writing the kernel in CUDA yourself. You get total control and no dependency on a DSL that is itself at version 0.1.9, and you lose everything this project is selling: the ability to write the kernel as Python, change it in minutes, and let a compiler schedule the memory movement. If your kernel is unusual enough that a DSL cannot express it, or if you need a guarantee of behaviour that only your own code can give, the hand-written route is correct. The cost is measured in engineer-weeks for one kernel, and in the correctness risk that comes with it.
The second alternative is the one already in the package: `tile_kernels/torch` holds PyTorch reference implementations. For many workloads the PyTorch version is fast enough, and it is what you would fall back to if a TileLang kernel fails to compile on your setup or if a future TileLang release changes behaviour underneath you. The README's own framing supports this, since it says the reference implementations exist in the project and the tests compare against them.
For completeness there is a middle path through other Python-level kernel compilers, where the same trade applies with a different tool. The choice between them is not about the DSLs in the abstract; it is about whether your shapes and your compiler versions are supported today. TileKernels narrows that question for you by pinning TileLang at 0.1.9 or higher and by naming the GPU architectures it targets, which is more information than most projects in this space give you.
The honest summary: this repository is most valuable as a worked example of how a lab turned working internal kernels into a published reference set, and as a source of correct starting points. Anyone expecting a supported dependency with documented speedups will not find either in the README.
Editorial conclusion
Use TileKernels as a reference when you are writing your own kernels for mixture-of-experts routing or low-precision quantization on Hopper or Blackwell hardware and want working code and PyTorch baselines to check against. Do not adopt it as a production dependency on a 0.x line with no tagged releases and no changelog, and do not expect documented speed figures, because the README publishes none. Verify first by running the transpose test with pytest -n 4 to confirm the toolchain works on your GPU, comparing a kernel against the matching implementation in tile_kernels/torch, and pinning an exact version because pip install tile-kernels resolves to whatever was tagged last.
Frequently asked questions
What hardware and software does TileKernels require?
The README lists Python 3.10 or higher, PyTorch 2.10 or higher, TileLang 0.1.9 or higher, an NVIDIA GPU on the SM90 or SM100 architecture, and the CUDA Toolkit at 13.1 or higher. The manifest states requires-python of >=3.10 with dependencies on torch>=2.10 and tilelang>=0.1.9.
How do I install TileKernels?
Use pip install tile-kernels for a released version from the package index, and pip install -e ".[dev]" for a local development version, whose dev extra adds pytest, pytest-xdist and pytest-repeat. Note that the index package is tile-kernels while the importable directory is tile_kernels.
How do I run the TileKernels tests and benchmarks?
The README shows pytest tests/transpose/test_transpose.py -n 4 for correctness only with four workers, the same command with --run-benchmark for correctness plus benchmarking, and TK_FULL_TEST=1 pytest -n 4 --count 2 as the pressure test that runs everything twice.
What is the torch directory in TileKernels used for?
The package includes a torch/ directory holding PyTorch reference implementations alongside the TileLang kernels. The tests compare against these, which is how the project argues that a kernel is correct rather than merely fast.
Is TileKernels ready for production use?
The README says most kernels approach the limit of hardware performance for compute intensity and memory bandwidth and that some have been used in internal training and inference, but also that they do not represent best practices and that code quality and documentation are being improved. The manifest classifier is Development Status :: 3 - Alpha.
What licence is TileKernels released under?
The code is released under the MIT License, with the text in the LICENSE file at the repository root, and the manifest declares license as MIT. The README also provides a BibTeX citation listing seven authors.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/deepseek-ai-tilekernels)