NVIDIA/apex: PyTorch Extensions for Mixed Precision and Distributed Training
A PyTorch Extension: Tools for easy mixed precision and distributed training in Pytorch
At a glance
- What is it?
- NVIDIA/apex is a PyTorch extension repository that provides fused CUDA kernels for mixed precision training, distributed data parallelism, and normalization layers. It is maintained by NVIDIA as an upstream source for PyTorch utilities, with builds requiring a matching CUDA environment.
- Who is it for?
- NVIDIA/apex is the right choice for researchers and engineers running large-scale training on NVIDIA hardware who need fused kernels not yet in mainstream PyTorch. It is the wrong choice if your PyTorch version does not meet the torch>=2.6.0 requirement or if build infrastructure for custom CUDA extensions is unavailable.
- Can I use it commercially?
- Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What NVIDIA/apex Provides and Who Needs It
NVIDIA/apex is a collection of PyTorch utilities for two primary training scenarios: mixed precision (combining 16-bit and 32-bit floating point to reduce memory and speed up computation) and distributed training across multiple GPUs. The repository is maintained by NVIDIA and described in its README as the upstream source for several utilities that eventually move into PyTorch itself.
The intended audience is machine learning engineers and researchers running large-scale training jobs on NVIDIA GPU hardware. The core value of apex over vanilla PyTorch is fused CUDA kernels. Fused kernels combine multiple operations into a single GPU kernel launch, reducing memory bandwidth usage and kernel launch overhead. The practical difference is faster and more numerically stable training for large models.
The README states: the intent of Apex is to make up-to-date utilities available to users as quickly as possible. Some of the code here will be included in upstream PyTorch eventually. This means apex sometimes provides features that are not yet in a stable PyTorch release.
Fused Kernels: What They Are and Why They Matter
The performance-critical utilities in apex depend on custom CUDA extensions that must be compiled during installation. The key ones are: FusedAdam in apex.optimizers (an Adam optimizer implemented as a fused CUDA kernel), FusedLayerNorm and FusedRMSNorm in apex.normalization (layer normalization with a fused kernel), and FusedLayerNorm improvements for apex.parallel.SyncBatchNorm.
A fused kernel for the Adam optimizer merges the parameter update steps that would otherwise be separate GPU operations into one kernel. This reduces the number of times data must move between GPU memory and compute units. For models with hundreds of millions of parameters, this translates to measurable training time reduction.
The README is explicit about what a Python-only build loses. A Python-only build omits fused kernels required to use FusedAdam, FusedLayerNorm, and FusedRMSNorm. DistributedDataParallel, amp, and SyncBatchNorm remain usable in a Python-only build, but may be slower.
A Python-only build is functionally a compatibility fallback: you can run the code without the performance gains.
Installing Apex from Source
Apex is not available as a pre-built wheel from PyPI for all configurations. The recommended install method is from source, cloning the repository and running pip with build flags. The README recommends installing Ninja first to speed up compilation.
For full functionality with C++ and CUDA extensions:
git clone https://github.com/NVIDIA/apex
cd apex
APEX_CPP_EXT=1 APEX_CUDA_EXT=1 pip install -v --no-build-isolation .To also build all contrib extensions at once:
APEX_CPP_EXT=1 APEX_CUDA_EXT=1 APEX_ALL_CONTRIB_EXT=1 pip install -v --no-build-isolation .For faster builds on systems with many CPU cores, parallel building can be enabled:
NVCC_APPEND_FLAGS="--threads 4" APEX_PARALLEL_BUILD=8 APEX_CPP_EXT=1 APEX_CUDA_EXT=1 pip install -v --no-build-isolation .The Python-only build (no CUDA extensions) uses:
pip install -v --disable-pip-version-check --no-build-isolation --no-cache-dir ./NVIDIA also provides pre-built containers through the NGC catalog at catalog.ngc.nvidia.com/orgs/nvidia/containers/pytorch, which include all apex extensions precompiled for the specific CUDA and driver versions in each container.
Key Modules: FusedAdam, FusedLayerNorm, and DistributedDataParallel
apex.optimizers.FusedAdam provides an Adam optimizer with fused CUDA kernels. It requires the CUDA extension build (APEX_CUDA_EXT=1). This is a drop-in replacement for torch.optim.Adam for training on NVIDIA hardware and requires no changes to the training loop beyond importing from the apex namespace.
apex.normalization provides FusedLayerNorm and FusedRMSNorm. FusedLayerNorm is numerically equivalent to PyTorch's LayerNorm but uses a fused kernel for better performance. This matters for Transformer architectures where layer normalization is applied at every attention and feed-forward block.
apex.parallel.DistributedDataParallel is apex's distributed training wrapper. The README notes it is usable in a Python-only build but may be slower. The CUDA extension build enables fused kernels that improve its performance.
The contrib modules in apex.contrib are a separate category. They require one or more install options beyond the core --cpp_ext and --cuda_ext flags. The README also notes that contrib modules do not necessarily support stable PyTorch releases and some may only be compatible with nightly builds.
The module table in the README maps each module name to its required environment variable and install option, making it straightforward to determine which APEX_ flags to set for a given module.
Build Complexity and Dependency Requirements
Building apex from source requires a working NVIDIA GPU environment: CUDA must be installed and accessible via the CUDA_HOME environment variable, and the CUDA version must be compatible with the installed PyTorch version. The requirements.txt file specifies torch>=2.6.0 as a dependency. Building against an older PyTorch version is not supported.
The contrib modules have additional requirements. For example, the cudnn_gbn module requires cuDNN 8.5 or higher. The nccl_p2p module requires NCCL 2.10 or higher. These requirements are listed in the module table in the README, but checking them before attempting a build avoids failed compilations.
Compilation time is significant. The README recommends Ninja to reduce it, and the parallel build flags (APEX_PARALLEL_BUILD and NVCC_APPEND_FLAGS) provide further speedup. On machines with limited CPU cores or memory, the README notes that the --parallel option is generally preferred over --threads.
Windows support is described as experimental. A Python-only build (pip install -v --no-cache-dir .) is described as more likely to work on Windows than the full CUDA build.
apex vs. PyTorch Native torch.amp
PyTorch introduced native automatic mixed precision through torch.cuda.amp in PyTorch 1.6 (2020). The torch.amp API provides autocast context managers and gradient scaling, which covers the basic mixed precision training use case without requiring apex.
The relationship between apex and torch.amp is one of precedence: apex's amp module predates the native PyTorch implementation, and several of its innovations influenced the design of torch.amp. For engineers starting a new training project today, torch.amp is the simpler starting point because it requires no separate installation.
apex remains relevant for the fused kernel implementations. torch.amp provides the mixed precision infrastructure; apex provides optimized kernel implementations for specific operations like FusedAdam and FusedLayerNorm. These two can coexist in the same training pipeline: torch.amp for autocast and gradient scaling, plus apex's FusedAdam for the optimizer step.
For engineers evaluating whether to add apex as a dependency, the relevant question is whether the specific fused kernels they need are already in their PyTorch version. If FusedLayerNorm is already available natively (PyTorch has been adding these), apex may add build complexity without adding new capabilities.
License, Maintenance, and Repository Status
NVIDIA/apex is licensed under the BSD-3-Clause license. This is a permissive license that allows use, modification, and distribution with attribution, a copyright notice, and without endorsement claims using NVIDIA's name. The license is in the LICENSE file at the top of the repository.
The repository is not archived. The last push was on 2026-09-23. It has no GitHub releases; updates are pushed directly to the master branch. The pyproject.toml and setup.py are present for building, with the build system set to setuptools. The requirements.txt specifies minimum versions for numpy, tqdm, PyYAML, pytest, packaging, and torch.
The repository includes an examples/ directory with subdirectories for imagenet, dcgan, and simple training examples. These examples demonstrate how to integrate apex modules into standard training loops.
Editorial conclusion
NVIDIA/apex is the right choice for researchers and engineers running large-scale training on NVIDIA hardware who need fused kernels not yet in mainstream PyTorch. It is the wrong choice if your PyTorch version does not meet the torch>=2.6.0 requirement or if build infrastructure for custom CUDA extensions is unavailable. Run pip show apex after installation to verify which extensions were built.
Frequently asked questions
How do I install NVIDIA apex?
Clone the repository with git clone https://github.com/NVIDIA/apex, then install with APEX_CPP_EXT=1 APEX_CUDA_EXT=1 pip install -v --no-build-isolation . from inside the apex directory. This requires a working CUDA environment and PyTorch 2.6.0 or higher. NVIDIA also provides NGC containers with apex precompiled.
What is NVIDIA apex?
NVIDIA/apex is a PyTorch extension repository providing fused CUDA kernels for mixed precision training (FusedAdam, FusedLayerNorm) and distributed training. NVIDIA maintains it as an upstream source for utilities that may eventually be added to PyTorch itself.
What is the difference between apex and PyTorch's built-in amp module?
PyTorch's torch.amp provides autocast and gradient scaling for mixed precision training without extra installation. apex provides additional fused kernel implementations such as FusedAdam and FusedLayerNorm that are separate from the torch.amp infrastructure. Both can be used together: torch.amp for autocast, apex's optimizers and normalization layers for fused performance.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/nvidia-apex)