Model or dataset
tile-ai/TileRT avatar
tile-ai/TileRT

TileRT: a tile-level runtime for ultra-low-latency LLM decoding

Tile-Based Runtime for Ultra-Low-Latency LLM Inference

1,831 stars128 forksPythonMIT

At a glance

What is it?
TileRT is a Python-served inference runtime from tile-ai that decomposes LLM operators into tile-level tasks and reschedules compute, I/O and communication across a single 8-GPU node. It is aimed at latency, not throughput, and its v0.1.5 wheel is pinned to one specific hardware and software stack.
Who is it for?
Adopt TileRT if your workload is single-request latency on an 8x B200 node and you can reproduce the pinned stack: Python 3.12, torch==2.11.0+cu130, transformers==4.46.3, tokenizers==0.20.3. Do not adopt it for throughput-oriented batch serving on mixed or older hardware, and do not expect the wheel to load on a host that drifts from those pins.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 8 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The latency problem TileRT is built around

Most inference servers are tuned for tokens per second across a batch. TileRT is tuned for the time between one output token and the next, on a single request. The README frames the goal as pushing latency limits "without compromising model size or quality", and names the scenarios it cares about: high-frequency trading, interactive AI, real-time decision-making, long-running agents and AI-assisted coding. In each of those, a user or a program is waiting on the next token, so a queue that keeps the GPU busy but delays any individual response is the wrong shape.

The project targets models with hundreds of billions of parameters and claims millisecond-level time per output token. The published milestones go further: v0.1.3 reports up to 500 tokens/s on GLM-5-FP8 and up to 600 tokens/s on DeepSeek-V3.2, and a June 2026 post describes pushing MiMo-V2.5-Pro-UltraSpeed past 1000 tokens/s on a 1-trillion-parameter model on a single 8-GPU node. Those are vendor figures from the project's own release notes, not independent measurements, and they describe best-case configurations.

Who this is for, concretely: teams that already own or rent an 8x B200 node and have a latency-sensitive product on top of it. If your serving fleet is A100 or H100, or if you batch requests to amortise cost, TileRT is not addressing your bottleneck.

How the tile-level engine overlaps compute, I/O and communication

The mechanism the README describes is a compiler-driven decomposition. LLM operators are broken into fine-grained tile-level tasks rather than being dispatched as whole kernels. The runtime then reschedules those tasks across multiple devices, overlapping computation, I/O and communication. The stated payoff is less idle time and better hardware utilisation.

This is a different lever from the usual ones. Tensor parallelism and pipeline parallelism decide how a model is split; they do not by themselves decide when each piece runs. TileRT's claim is about the schedule: if a tile of one operator is waiting on an all-reduce, the runtime can put a tile of another operator on the same device instead of stalling. On a single 8-GPU node with fast interconnect, that overlap is where the latency win is supposed to come from.

The README also states that the underlying compiler techniques will be shared with the community as they are integrated into TileLang and TileScale. That matters for evaluation: the runtime is the shipped artefact, while the compiler work behind it is described as arriving later through separate projects. If you want to understand or modify the scheduling decisions rather than consume them, the public surface is thinner than the runtime itself.

Multi-token prediction is a second lever. The v0.1.2-alpha.1 notes report decoding rates up to 590 tokens/s with mtp=3 under synthetic workloads, and the GLM-5.1 chart in the README compares TileRT without MTP, with MTP at an average acceptance length of 3.2, and a peak under best-case acceptance of 4.0. Synthetic workloads and best-case acceptance are exactly the conditions where speculative decoding looks strongest, so treat the peak bar as an upper bound rather than an expectation.

Installing TileRT: the pinned wheel and the Docker path

The README is explicit that v0.1.5 ships as a pre-built binary wheel linked against an exact ABI. Python, CUDA and PyTorch combinations outside the listed set are described as untested and not guaranteed to work, and the README asks you to reproduce the environment rather than treat the versions as lower bounds. The pinned set is Python 3.12, torch==2.11.0+cu130, transformers==4.46.3 and tokenizers==0.20.3, on Linux x86_64 with glibc 2.28 or newer, on a machine with 8x NVIDIA B200 and a driver that supports the CUDA 13.2 runtime.

The recommended route is the prebuilt image, which the README says avoids version drift on the host. It is mirrored to two registries, so pull from whichever is reachable:

bash
docker pull ghcr.io/tile-ai/tilert:cu132-latest

If that registry is blocked, the same image is published to Docker Hub:

bash
docker pull tileai/tilert:cu132-latest

If you install from PyPI instead, the order matters. torch has to come from PyTorch's cu130 index first, because the default PyPI torch is a different CUDA build and the repository notes say it will not load the cu130-linked tilert binary:

bash
pip install --index-url https://download.pytorch.org/whl/cu130 torch==2.11.0
pip install -r requirements.txt

The repository's requirements.txt carries the same warning in a comment, so the two-step install is the intended sequence rather than a workaround. After that, the README points to a generation example and an MTP generation example as the first real use; the README excerpt available here shows the navigation links to those sections but not their contents, so follow the README's own commands for the first generation run rather than improvising an entry point.

Where TileRT will not help you

The hard constraint is the hardware. The v0.1.5 wheel was built against 8x NVIDIA B200 with a CUDA 13.2-capable driver. That is not a recommendation in the README, it is the build target, and the README states that other combinations are untested. A team with H100s, A100s, or a B200 node with fewer than eight GPUs has no supported path in this documentation.

The second constraint is the software pin. transformers is held at 4.46.3, and the Dockerfile comments say the 5.x branch is not backward compatible with TileRT's tokenizer and model loading paths. If your stack already depends on a newer transformers for another service, you are looking at a separate environment, not a shared one. The same logic applies to torch: the wheel is linked to the cu130 build, so a container that resolves torch from the default index will not work.

The third is the workload itself. TileRT explicitly prioritises responsiveness over high-throughput batch processing. If your traffic is many concurrent users where aggregate tokens per second determines cost, a throughput-oriented server is the better fit and TileRT's scheduling work is largely wasted on you. Latency-optimised runtimes and throughput-optimised runtimes make opposite trade-offs, and picking the wrong one shows up as either a slow product or an expensive one.

Finally, the Dockerfile describes a release pipeline that rebuilds the wheel, starts a fresh container, installs it and runs pytest on B200 GPUs. That is the project's own validation loop. Nothing in the README suggests a CPU-only or single-GPU fallback for development, so plan on having the target hardware available before you write integration code.

TileRT compared with vLLM, and with PD disaggregation

vLLM is the natural comparison because the two projects now meet. TileRT v0.1.5 introduced prefill-decode disaggregation, described in the release notes as vLLM prefill plus TileRT decode, behind an OpenAI-compatible endpoint, supported on GLM-5/5.1 and DeepSeek-V3.2. The v0.1.5.post2 release is tagged as exactly that combination.

The difference in approach is the unit of scheduling. vLLM is a general-purpose serving engine built around paged attention and continuous batching, which is a throughput-first design that keeps many sequences in flight. TileRT decomposes operators into tile-level tasks and reschedules them across devices to shorten the critical path of a single request. The disaggregated setup uses each for what it is better at: prefill is a large, parallel, throughput-friendly phase, so it runs on vLLM, while decode is the latency-critical token-by-token loop, so it runs on TileRT.

That is a pragmatic arrangement rather than a winner-takes-all claim, and it is also a useful signal about where the project sees its own boundary. If TileRT were the better prefill engine too, there would be no reason to route that phase elsewhere. The trade-off for you is operational: two engines, two sets of version pins, and an OpenAI-compatible endpoint joining them. The README does not document rollback behaviour for that split, so the failure modes of the disaggregated path are not something you can plan from the documentation alone.

Licence, maintenance and the cost of staying pinned

TileRT is MIT licensed, both in the repository metadata and in pyproject.toml. That is permissive: you can use it commercially, modify it and redistribute it, provided the copyright notice and licence text travel with it. The runtime dependencies you install alongside it carry their own terms, and the Dockerfile pulls from a PyTorch manylinux builder image plus EPEL packages such as glog and zstd, so your container's licence surface is wider than the tilert package alone. This is not legal advice; check the terms of the components you actually ship.

The maintenance picture is current. The last push to the default branch was on 2026-08-13, and the release cadence in the notes runs from v0.1.0-alpha.1 in November 2025 through v0.1.5.post2 in August 2026. The repository is not archived.

The upgrade cost is the part to budget for. Because the wheel is ABI-linked, a version bump is not a pip upgrade in practice. The Dockerfile comments instruct that no dependency be bumped without re-running the full release pipeline, and the pinned lock set in the image was resolved against a specific torch and transformers pair. When TileRT moves to a new torch or a new transformers, your environment moves with it, including any other service sharing that container. The safest posture is to treat the TileRT container as an appliance with its own image tag, and to keep the rest of your stack outside it.

Editorial conclusion

Adopt TileRT if your workload is single-request latency on an 8x B200 node and you can reproduce the pinned stack: Python 3.12, torch==2.11.0+cu130, transformers==4.46.3, tokenizers==0.20.3. Do not adopt it for throughput-oriented batch serving on mixed or older hardware, and do not expect the wheel to load on a host that drifts from those pins. Before you commit, verify three things on your own node: that the prebuilt image starts with all eight GPUs visible, that the model you intend to serve appears in the supported list for the release you install, and that the PD disaggregation path you want is the one shipped in v0.1.5.post2 rather than the earlier v0.1.5 tag.

Frequently asked questions

What is TileRT?

TileRT is a tile-based runtime for ultra-low-latency LLM inference. It decomposes LLM operators into fine-grained tile-level tasks and reschedules computation, I/O and communication across devices to reduce the time per output token.

How does TileRT compare with vLLM?

vLLM is a throughput-oriented serving engine, while TileRT prioritises responsiveness for individual requests. The two are not mutually exclusive: TileRT v0.1.5 introduced PD disaggregation that runs vLLM prefill with TileRT decode behind an OpenAI-compatible endpoint, supported on GLM-5/5.1 and DeepSeek-V3.2.

Which GPUs and software versions does TileRT require?

The v0.1.5 wheel was built against 8x NVIDIA B200 with a CUDA 13.2-capable driver, on Linux x86_64 with glibc 2.28 or newer, using Python 3.12, torch==2.11.0+cu130, transformers==4.46.3 and tokenizers==0.20.3. The README describes other combinations as untested and not guaranteed to work.

How do I install TileRT?

The README recommends the prebuilt Docker image, available as ghcr.io/tile-ai/tilert:cu132-latest or tileai/tilert:cu132-latest, because it avoids version drift on the host. Installing from PyPI requires torch to be installed first from PyTorch's cu130 index, because the default PyPI torch is a different CUDA build.

Which models does TileRT support?

The release notes and README reference DeepSeek-V3.2, GLM-5 and GLM-5.1, with PD disaggregation supported on GLM-5/5.1 and DeepSeek-V3.2. The 1000 tokens/s milestone was reported on MiMo-V2.5-Pro-UltraSpeed in collaboration with Xiaomi MiMo.

Official sources

  1. License: MIT
  2. Project website
  3. README
  4. Releases
  5. tile-ai/TileRT on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/tile-ai-tilert.svg)](https://hysenlabs.com/projects/tile-ai-tilert)