# NVIDIA TransformerEngine: FP8 and FP4 building blocks for Transformer training

> Transformer Engine is NVIDIA's library of fused Transformer modules that keep FP8, MXFP8 and NVFP4 scaling factors inside the layer. It is aimed at teams already training on Hopper, Ada or Blackwell hardware, and it is not a framework replacement.

**NVIDIA/TransformerEngine** — A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hopper, Ada and Blackwell GPUs, to provide better performance with lower memory utilization in both training and inference.

- Repository: https://github.com/NVIDIA/TransformerEngine
- Website: https://docs.nvidia.com/deeplearning/transformer-engine/user-guide/index.html
- Stars: 3,537 · Forks: 827
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/nvidia-transformerengine

## The problem TransformerEngine targets: low precision without hand-written scaling

Training a large Transformer in FP8 means every weight and activation tensor needs a scale factor, and those scales have to be updated as the distribution of values shifts during training. Doing that by hand inside a model definition is error-prone and easy to get subtly wrong. TransformerEngine packages the scaling logic into the modules themselves. The README states that modules provided by TE internally maintain scaling factors and other values needed for FP8 training, which is the whole point: the model code stays close to ordinary PyTorch or Flax code while the numerics underneath run at reduced precision.

The audience is narrow and specific. This is for engineers who already have NVIDIA GPUs of the right generation and who are hitting memory or throughput walls in training or inference. The README lists support for FP8 on Hopper, Ada and Blackwell, MXFP8 and NVFP4 on Blackwell, and optimizations across FP16 and BF16 on Ampere and later. If your hardware is older than Ampere, or not NVIDIA at all, the library has nothing to offer you.

## How the autocast recipe and internal scaling work together

The mechanism has two halves. The first is a recipe object that describes the format and the scaling strategy. The second is an autocast context that tells the TE modules to run in low precision for the duration of the block. In the PyTorch example in the README, a DelayedScaling recipe is constructed with margin=0 and fp8_format=recipe.Format.E4M3, then passed to te.autocast(enabled=True, recipe=fp8_recipe). Inside that context, a te.Linear layer consumes a CUDA tensor and produces an output; outside it, the same layer would behave like a normal linear layer.

The data flow is what makes this different from a global precision flag. The scale factors live with the module, so a forward pass reads the current scales, quantizes inputs, runs the fused kernel, and updates the scales for the next iteration. DelayedScaling, as the name implies, does not recompute the scale on every step; it uses a margin to decide when to update. That is a deliberate trade-off between accuracy and the cost of the scale computation, and it is the reason the recipe exposes a margin parameter at all.

On the JAX side the shape of the API is the same. The README's Flax example builds a DenseGeneral layer inside te.autocast with a DelayedScaling recipe, then applies it with a params dict. The framework changes; the recipe and autocast pattern does not.

## Installing TransformerEngine and running a first FP8 layer

The build system in pyproject.toml requires setuptools>=61.0, cmake>=3.21, wheel, pybind11[global], ninja, pip, torch>=2.1, jax>=0.5.0, flax>=0.7.1 and nvidia-cudnn-frontend>=1.28.0. That list is a fair warning about what a source build pulls in: a CUDA toolchain, CMake, Ninja, and both major frameworks unless you trim them. setup.py reads the installed frameworks through get_frameworks() and picks a different build extension depending on whether PyTorch or JAX is present, so the wheel you get is tied to the framework set at build time.

A typical source install from a checkout looks like this. The build compiles CUDA kernels, so expect it to take a while and to need nvcc on PATH.

```bash
pip install .
```

If you want the PyTorch extension built with parallel compilation, setup.py exposes NVTE_UB_WITH_MPI for the userbuffers path, but the README does not document that variable, so treat it as a build-system detail rather than a supported knob.

Once installed, the minimal FP8 example is the one from the README. Create the layer, create the recipe, and wrap the forward pass.

```python
import torch
import transformer_engine.pytorch as te
from transformer_engine.common import recipe

model = te.Linear(768, 3072, bias=True)
inp = torch.randn(2048, 768, device="cuda")
fp8_recipe = recipe.DelayedScaling(margin=0, fp8_format=recipe.Format.E4M3)

with te.autocast(enabled=True, recipe=fp8_recipe):
    out = model(inp)

loss = out.sum()
loss.backward()
```

What you should see is a forward and backward pass that completes without a precision-related error, with the layer holding its own scales. If the GPU does not support the requested format, the failure surfaces at this point rather than at import.

## Where TransformerEngine is the wrong tool

The clearest limitation is hardware. FP8 requires Hopper, Ada or Blackwell; MXFP8 and NVFP4 require Blackwell. On an older NVIDIA card you can still get FP16 and BF16 fused kernels, but the low-precision formats that justify the dependency are unavailable, and you have taken on a heavy build for a smaller return. On non-NVIDIA accelerators the library does not apply at all, and the search phrase about ROCm support is not answered anywhere in the README.

The second limitation is build and version coupling. pyproject.toml pins torch>=2.1, jax>=0.5.0, flax>=0.7.1 and nvidia-cudnn-frontend>=1.28.0, and setup.py selects its build extension based on which frameworks are installed. A wheel built without JAX will not serve a JAX training job. That means the artifact you install is a function of the machine that built it, which is awkward if your team mixes PyTorch and JAX workloads.

Third, this is not a convergence fix. The README points to a Convergence section and to external posts about NVFP4 training, but the library itself gives you modules and kernels. If your loss is diverging in BF16, switching to FP8 will not diagnose it.

## TransformerEngine against Flash Attention

The two are often compared, and they solve different layers of the stack. Flash Attention is an attention kernel: it changes how the attention computation is tiled and how memory is moved, but it does not change the numeric format of the weights or the surrounding linear layers. TransformerEngine is a module library with a precision policy. It ships fused kernels, and the README lists fused operations and MoE support among its optimizations, but its distinguishing feature is that its modules carry FP8, MXFP8 or NVFP4 scaling state.

In practice they are complementary rather than substitutes. A model can use a Flash Attention style kernel for attention and TE modules for the linear and normalization layers. The reason to pick TE specifically is that you want the low-precision formats and you are willing to accept the CUDA and framework version constraints that come with them. If you only want faster attention at BF16 and your stack is already stable, the heavier dependency is hard to justify.

## Maintenance cadence, licence and upgrade cost

The repository is not archived, and the last push was on 2026-09-10. Releases are frequent: v2.18 on 2026-08-14, v2.17.1 on 2026-08-07 and v2.17 on 2026-07-28. That cadence is a real cost as well as a benefit. A library that ships a minor release roughly every few weeks will occasionally move a recipe default or a build requirement, and because the build is compiled, an upgrade is not a pure Python dependency bump. Budget for rebuilding the extension and re-running your convergence checks after each version change, not just for editing a version pin.

The licence is Apache-2.0, which is permissive and generally straightforward for commercial use. The README defers to LICENSE for the full terms, and the source headers carry an NVIDIA copyright notice. Apache-2.0 includes a patent grant and requires that you preserve notices; if your legal team has specific questions about the patent or notice clauses, that is a question for them rather than something the README resolves.

## Conclusion

Adopt TransformerEngine if you train or serve Transformer models on Hopper, Ada or Blackwell GPUs and you can pin CUDA, cuDNN and framework versions for the whole cluster; the payoff is lower memory use and fused kernels, not a new training API. Do not adopt it for CPU-only work or for non-NVIDIA accelerators, and do not expect it to fix an unstable training run. Before committing, verify the CUDA architecture list your build targets, confirm which framework extra the wheel was built for, and check that your parallelism setup is covered by the examples directory.

## FAQ

### What is TransformerEngine?

It is a library for accelerating Transformer models on NVIDIA GPUs, using FP8 on Hopper, Ada and Blackwell and MXFP8 or NVFP4 on Blackwell. It provides optimized building blocks plus a framework-agnostic C++ API for integrating FP8 support into other deep learning libraries.

### How do I install TransformerEngine?

The build system requires CMake, Ninja, pybind11, a CUDA toolchain and either PyTorch or JAX, and setup.py selects its build extension based on which frameworks are installed. A source install from a checkout is done with pip install . and compiles the CUDA kernels.

### How is TransformerEngine different from Flash Attention?

Flash Attention is an attention kernel, while TransformerEngine is a set of Transformer modules that maintain their own FP8, MXFP8 or NVFP4 scaling factors. The README describes TE as providing fused kernels and an autocast-style API rather than a single attention implementation.

## Sources

- [License: Apache-2.0](https://github.com/NVIDIA/TransformerEngine/blob/main/LICENSE)
- [NVIDIA/TransformerEngine on GitHub](https://github.com/NVIDIA/TransformerEngine)
- [Project website](https://docs.nvidia.com/deeplearning/transformer-engine/user-guide/index.html)
- [README](https://github.com/NVIDIA/TransformerEngine/blob/main/README.md)
- [Releases](https://github.com/NVIDIA/TransformerEngine/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/nvidia-transformerengine
