AITemplate: compiling PyTorch models into standalone CUDA/HIP C++ for fp16 inference
AITemplate is a Python framework which renders neural network into high performance CUDA/HIP C++ code. Specialized for FP16 TensorCore (NVIDIA GPU) and MatrixCore (AMD GPU) inference.
At a glance
- What is it?
- AITemplate is a Python framework that renders neural networks into self-contained CUDA or HIP binaries, with horizontal, vertical and memory fusion as its main selling points. It is aimed at teams running fp16 inference on Ampere-class NVIDIA or CDNA2 AMD GPUs who are willing to trade portability for kernel-level control.
- Who is it for?
- AITemplate fits teams with fixed fp16 inference graphs on SM80+ NVIDIA or CDNA2 AMD hardware, a working CUDA 11.6 or ROCm 5.2.3 toolchain, and a willingness to compile per model. It does not fit anyone on T4, V100 or CDNA1 hardware, teams that need dynamic shapes today, or projects that cannot absorb a codegen learning curve for every unsupported operator.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem AITemplate targets: inference graphs that outgrow general-purpose kernel libraries
Most fp16 GPU inference stacks assemble a graph from library calls: cuBLAS or rocBLAS for matrix multiplies, cuDNN or MIOpen for convolutions, and a runtime such as TensorRT or MIGraphX to schedule them. AITemplate takes the opposite route. It is a Python framework that renders a neural network into CUDA (NVIDIA) or HIP (AMD) C++ code, then compiles that code into a self-contained portable binary. The README states plainly that it does not depend on cuBLAS, cuDNN, rocBLAS, MIOpen, TensorRT or MIGraphX. The audience is therefore narrow and specific: engineers serving fp16 models on NVIDIA SM80+ GPUs (Ampere and later) or AMD CDNA2 GPUs (MI-210 and MI-250), who want the compiler to own the kernel schedule rather than delegate it.
The stated goal is performance close to roofline fp16 TensorCore (NVIDIA) or MatrixCore (AMD) on models such as ResNet, MaskRCNN, BERT, VisionTransformer and Stable Diffusion. That word, roofline, is the honest framing of the trade-off: AITemplate is chasing the hardware limit, and the cost of that chase is a build step and a hardware floor. If your workload is a small model on a T4, or a graph that changes shape on every request, the framework's own README tells you it is not the right starting point.
How the codegen pipeline works: graph nodes in, fused kernels out
The architecture is a two-part codegen contract. The README describes extension work as adding two Python files: one for a graph node definition and another for the backend codegen. A CUDA or HIP kernel written in a text header file can be used directly by the codegen, so a hand-written kernel does not need to be ported into an internal DSL before it can participate in fusion.
Fusion is the mechanism the project leans on hardest, and it comes in three named forms. Vertical fusion collapses elementwise operations, reductions and layout permutations into TensorCore or MatrixCore operations, and supports back-to-back TensorCore/MatrixCore operation fusion. Horizontal fusion takes parallel GEMM, LayerNorm and other operators with different input shapes and merges them into a single GPU kernel. Memory fusion merges GEMM or LayerNorm with concatenation, split and slice, again into one operator. These are not three names for the same pass; they attack different boundaries, one across layers, one across independent branches, one across tensor reshaping.
The runtime side has a second design decision worth noting. The generated Python runtime can accept PyTorch tensors as inputs and outputs without an extra copy, but the README also states that for environments without PyTorch the Python and C++ runtime is self-contained. So the compiled artifact does not carry a PyTorch dependency into production unless you want the zero-copy interop.
Installing AITemplate with Docker or from source
The README recommends Docker first, specifically to avoid using a mismatched NVCC or HIPCC. Clone with submodules, since the build depends on them:
git clone --recursive https://github.com/facebookincubator/AITemplateThen build the image for your platform. The README gives two commands, one per vendor, and both produce an image tagged ait:latest:
./docker/build.sh cudaDOCKER_BUILDKIT=1 ./docker/build.sh rocmIf you build from source instead, the README pins CUDA 11.6, and for ROCm it says the project was tested on ROCm 5.2.3 with a customized HIPCC build whose command lives in docker/Dockerfile.rocm. The README warns that an incorrect compiler will lead to performance regression, and asks you to confirm all submodules are cloned before continuing. The wheel build is two commands:
cd python
python setup.py bdist_wheel
pip install dist/*.whl --force-reinstallFor a first real use, the README points to the documentation site for API reference and to three onboarding tutorials: how to inference a PyTorch model with AIT, how to add an op to AIT codegen, and how to visualize AIT's optimization. Model templates live under examples/, covering ResNet-50 with TIMM, MaskRCNN-FPN with Detectron2, BERT with Hugging Face Transformers, Vision Transformer with TIMM and Stable Diffusion with Hugging Face Diffusers. Start there rather than with your own graph.
The hardware floor and the dynamic shape gap
Two constraints are stated outright in the README and should decide adoption before anything else. First, NVIDIA support is tested only on SM80 and above. The README says not all kernels work with old SM75 or SM70 GPUs, naming T4 and V100. Second, AMD support is tested only on CDNA2, meaning MI-210 and MI-250, with possible compiler issues on CDNA1 MI-100. If your fleet is V100 or T4, the framework is simply not addressed to you.
The more interesting limitation is dynamic shape support. It appears in the README's mid-term plan, not as a shipped feature: better dynamic shape support, focused on the dynamic sequence in Transformers, with symbolic shape support listed as something to add. The long-term plan lists automatic ONNX and Open-XLA conversion, quantization at fp8/int8/int4, sparsity pruning for GEMM, and PT2 integration via Aten2AIT, which the README describes as under active development. So a model with variable sequence length is a case where AITemplate is currently the wrong tool, and the project says so by placing it on the roadmap.
Operator coverage is the second failure mode. AITemplate does not support all PyTorch operators, and the FX2AIT tool exists precisely because of that gap.
FX2AIT and partial lowering for unsupported operators
FX2AIT is a Python tool bundled in the repository that converts PyTorch models into an AITemplate engine. Its output is an AITModule for inference serving, and the README says conversion needs only a PyTorch model and an input. The part that matters is the AITLowerer: when a model contains operators AITemplate cannot handle, the lowerer performs partial AIT acceleration rather than failing the whole conversion. The README points to fx2ait/fx2ait/example/03_lowering_split for the worked case.
This is a pragmatic answer to the coverage problem, but it is a split, not a solution. Part of the graph runs as generated AIT kernels and the rest stays in PyTorch, which means the boundary between the two becomes something you have to reason about. The README does not document the cost of that boundary. Anyone evaluating FX2AIT should read the lowering example before assuming the split is free.
How AITemplate differs from TensorRT and MIGraphX
The README names TensorRT and MIGraphX directly, as dependencies AITemplate does not use. That is the clearest statement of the difference in approach. TensorRT and MIGraphX are runtimes that take a graph and select from a set of prebuilt, vendor-tuned kernels; AITemplate generates the C++ kernels themselves and compiles a binary per model. The README's claim about the result is that each model becomes a self-contained portable binary usable on any software environment with the same hardware.
That portability claim has a specific boundary: same hardware. A binary built for SM80 is not a binary for CDNA2, and the build is per model, not per family of models. The payoff the project claims is a wider range of fusions than existing solutions on both GPU platforms, in three flavors (horizontal, vertical, memory). The cost is that every model is a compilation target. With TensorRT you ship an engine file and a runtime; with AITemplate you ship a compiled artifact and own the codegen path when an operator is missing. Teams that value a stable, vendor-supported operator library over fusion headroom should stay with the runtime approach.
Licence, maintenance and the cost of upgrading
AITemplate is Apache-2.0, so the licence permits commercial use and modification under its terms; the repository also carries a licenses/ directory and a CITATION.cff. Nothing here suggests a copyleft obligation, but the licence text is the authority, not this summary.
The repository is not archived, and the last push was on 2026-08-07, which is recent. Release tags are a different story: the two listed releases are 0.1 from 2022-10-11 and 0.2 from 2023-01-31. The README explains the policy directly, stating that releases are not on a set schedule and will only be tagged for significant feature releases, and that all current development updates can be seen in the repository. The practical consequence for upgrades is that you should track main and the submodules rather than wait for a tag, and that a rebuild means recompiling every model against a possibly different CUDA or HIPCC version. The README's own warning about compiler mismatch causing performance regression is the upgrade risk in one sentence. Pin your toolchain, and treat a toolchain bump as a re-validation event for each compiled model.
Editorial conclusion
AITemplate fits teams with fixed fp16 inference graphs on SM80+ NVIDIA or CDNA2 AMD hardware, a working CUDA 11.6 or ROCm 5.2.3 toolchain, and a willingness to compile per model. It does not fit anyone on T4, V100 or CDNA1 hardware, teams that need dynamic shapes today, or projects that cannot absorb a codegen learning curve for every unsupported operator. Before committing, clone with --recursive, build the Docker image, and confirm that the operator coverage of your model matches what examples/ and fx2ait/fx2ait/example/03_lowering_split actually demonstrate.
Frequently asked questions
Which GPUs does AITemplate support?
The README states that NVIDIA support is tested only on SM80 and above, and that not all kernels work on older SM75 or SM70 GPUs such as T4 and V100. For AMD, it is tested only on CDNA2 GPUs (MI-210 and MI-250), with possible compiler issues on CDNA1 MI-100.
Does AITemplate depend on TensorRT, cuBLAS or cuDNN?
No. The README states that AITemplate does not depend on third-party libraries or runtimes such as cuBLAS, cuDNN, rocBLAS, MIOpen, TensorRT or MIGraphX, and that each model is compiled into a self-contained portable binary for the same hardware.
How do I install AITemplate?
The README recommends Docker, using ./docker/build.sh cuda for NVIDIA or DOCKER_BUILDKIT=1 ./docker/build.sh rocm for AMD, which produces an image tagged ait:latest. Building from source requires CUDA 11.6, or ROCm 5.2.3 with a customized HIPCC build, and then python setup.py bdist_wheel followed by pip install dist/*.whl --force-reinstall.
What does FX2AIT do when a model uses operators AITemplate does not support?
FX2AIT's AITLowerer performs partial AIT conversion, so models with unsupported operators can still get partial acceleration instead of failing outright. The README points to fx2ait/fx2ait/example/03_lowering_split for the worked example.
Does AITemplate support dynamic shapes?
The README lists better dynamic shape support, focused on the dynamic sequence in Transformers with symbolic shape support, as a mid-term plan rather than a shipped capability. Models with variable input shapes are therefore not the framework's current strength.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/facebookincubator-aitemplate)