# NVIDIA DALI: GPU Data Pipelines for PyTorch, TensorFlow and PaddlePaddle

> DALI moves decoding and augmentation off the CPU and onto the GPU, with a pipeline mode and a newer dynamic mode. It pays off when the input stage is the bottleneck and your hardware is NVIDIA-only.

**NVIDIA/DALI** — A GPU-accelerated library containing highly optimized building blocks and an execution engine for data processing to accelerate deep learning training and inference applications.

- Repository: https://github.com/NVIDIA/DALI
- Website: https://docs.nvidia.com/deeplearning/dali/user-guide/docs/index.html
- Stars: 5,769 · Forks: 680
- Language: C++
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/nvidia-dali

## The CPU preprocessing bottleneck DALI targets

The README states the problem plainly: deep learning pipelines involve loading, decoding, cropping, resizing and other augmentations, and those stages currently execute on the CPU, which limits training and inference throughput. DALI's answer is to offload that work to the GPU and to run it through its own execution engine, with prefetching, parallel execution and batch processing handled without the user writing that logic.

The audience follows from that. This is for teams whose GPU utilization dips because the input stage cannot feed it fast enough, and who are already committed to NVIDIA hardware. The library is portable across TensorFlow, PyTorch, PaddlePaddle and JAX, so a pipeline written once can be retargeted. If your data loading is already fast enough, or your accelerator is not an NVIDIA GPU, the premise of the project does not apply to you.

## Pipeline mode versus dynamic mode: two execution models

The README shows two APIs side by side. Pipeline mode is the older, declarative style: you decorate a function with @pipeline_def, describe the graph with fn.readers.file, fn.decoders.image_random_crop, fn.resize and fn.crop_mirror_normalize, and hand the resulting pipeline to a framework iterator such as DALIGenericIterator from nvidia.dali.plugin.pytorch. The graph is defined once, then executed repeatedly.

Dynamic mode is the newer experimental API, imported as nvidia.dali.experimental.dynamic and aliased ndd. There is no decorator and no graph definition. You create a reader with ndd.readers.File, call reader.next_epoch(batch_size=16) in a loop, and apply operators such as ndd.decoders.image_random_crop and ndd.resize directly to the returned arrays. Tensors convert with torch.as_tensor. The README also points to a dali-dynamic-mode skill for AI agents that documents the Dynamic Mode API and best practices.

The trade-off is real. Pipeline mode requires you to describe the whole graph up front, which is more ceremony but lets the engine plan prefetching and parallelism. Dynamic mode reads like ordinary imperative code, which is easier to debug and to interleave with training logic, but the README labels it experimental, so its API is the one more likely to move between releases.

## Installing NVIDIA DALI and running a first pipeline

The README does not include installation instructions; it links to the documentation at docs.nvidia.com, and related search traffic shows people looking for a pip install path. The repository layout includes a conda/ directory and a docker/ directory alongside the Python package under dali/, so packaged and container routes both exist. Check the documentation page for the exact command matching your CUDA and framework versions rather than copying a version string from anywhere else.

The README examples depend on an environment variable pointing at a sample dataset. It references https://github.com/NVIDIA/DALI_extra and reads os.environ['DALI_EXTRA_PATH'], joining it with db/single/jpeg. Without that variable set, the example fails before it reaches any DALI operator.

Once installed, a minimal pipeline-mode script looks like this. The decorator takes num_threads and device_id, the reader shuffles files from a directory, and the decoder runs on device="mixed" so decoding happens on the GPU while the reader stays on the CPU.

```python
from nvidia.dali.pipeline import pipeline_def
import nvidia.dali.types as types
import nvidia.dali.fn as fn

@pipeline_def(num_threads=4, device_id=0)
def get_dali_pipeline():
    images, labels = fn.readers.file(
        file_root=images_dir, random_shuffle=True, name="Reader")
    images = fn.decoders.image_random_crop(
        images, device="mixed", output_type=types.RGB)
    images = fn.resize(images, resize_x=256, resize_y=256)
    return images, labels
```

The same operations in dynamic mode drop the decorator and loop over epochs instead.

```python
import nvidia.dali.experimental.dynamic as ndd

reader = ndd.readers.File(file_root=images_dir, random_shuffle=True)
for images, labels in reader.next_epoch(batch_size=16):
    images = ndd.decoders.image_random_crop(images, device="gpu", output_type=types.RGB)
    images = ndd.resize(images, resize_x=256, resize_y=256)
```

Note the device argument differs between the two: pipeline mode uses device="mixed" for the decoder, dynamic mode uses device="gpu". That is not a cosmetic difference, and copying one example into the other API will not work.

## Where DALI is the wrong tool

The largest constraint is hardware. DALI is built around GPU offload, and the README's own framing is that it addresses the CPU bottleneck by moving work to the GPU. On a CPU-only machine, or on accelerators from other vendors, the central mechanism is unavailable. The README does state that CPU and GPU execution are both supported, so CPU execution exists, but the performance argument for adopting the library rests on the GPU path.

A second limitation is the operator surface. DALI ships its own readers and augmentations, and a custom pipeline that does not map onto those operators requires writing a custom operator, which the README lists as an extensibility feature rather than a small task. Teams with unusual preprocessing, or with augmentation logic that changes weekly, will find a plain framework dataloader easier to modify.

Third, the two APIs carry different stability expectations. Dynamic mode is imported from nvidia.dali.experimental.dynamic, and the README describes the accompanying skill as guidance on the Dynamic Mode API, which is consistent with an interface still settling. Production training code that must survive upgrades has a reason to prefer pipeline mode today. Finally, DALI does not remove I/O limits: if your storage cannot deliver bytes fast enough, moving decode to the GPU does not create data.

## NVIDIA DALI versus the PyTorch DataLoader

The natural comparison is torch.utils.data.DataLoader with a Dataset. That combination runs augmentation in Python worker processes on the CPU, and its performance comes from increasing num_workers, which consumes CPU cores and host memory. DALI replaces that model with a graph executed by its own engine, where decoding and augmentation run on the GPU and prefetching is handled internally. The README describes DALI as usable as a portable drop-in replacement for built-in data loaders and iterators in popular frameworks, but the examples show a DALI-specific pipeline definition and a DALIGenericIterator rather than a Dataset subclass, so the migration is a rewrite of the input stage, not a one-line swap.

The practical difference is where the work happens and what you give up. PyTorch's DataLoader is pure Python, trivially debuggable, and works on any device. DALI trades that flexibility for throughput on NVIDIA hardware, and adds framework portability: the same pipeline can be retargeted to TensorFlow, PyTorch, PaddlePaddle or JAX. If you are not CPU-bound in the input stage, that trade buys you nothing.

## Maintenance, releases and the Apache-2.0 licence

The repository is not archived, and the last push was on 2026-09-10, so development is ongoing. Release cadence is visible in the tags: v2.3.0 on 2026-08-28, v2.2.0 on 2026-06-29 and v2.1.1 on 2026-06-15. That pattern suggests minor releases every couple of months plus patch releases, which is a normal rhythm for a library that must track CUDA, framework and driver versions.

That tracking is the real upgrade cost. DALI sits between your framework and your GPU stack, and the repository carries version pins such as DALI_DEPS_VERSION, DALI_EXTRA_VERSION and a VERSION file, plus a conda/ and docker/ directory. Upgrading DALI can therefore mean coordinating a framework version, a CUDA version and a container image. The README does not document a rollback procedure, so pinning the DALI version in your environment is the practical safeguard.

The licence is Apache-2.0, which permits commercial and closed-source use and requires preservation of notices. This is a description of the licence identifier, not legal advice; check the LICENSE file and your organisation's policy for the terms that apply to you.

## Conclusion

Adopt DALI when an NVIDIA GPU sits idle while your CPU dataloader decodes and augments images, video or audio, and when you can accept a custom pipeline definition instead of a torch.utils.data.Dataset. Skip it if you train on non-NVIDIA accelerators, on CPU-only machines, or if your bottleneck is disk or network throughput rather than preprocessing. Before committing, verify that the operators you need exist for your data format, that DALI_EXTRA_PATH is set if you intend to run the README examples unchanged, and that your framework plugin version matches the DALI release you install.

## FAQ

### What is NVIDIA DALI?

It is a GPU-accelerated library for data loading and preprocessing, providing optimized building blocks for image, video and audio data plus its own execution engine. It is meant to replace built-in data loaders in frameworks such as TensorFlow, PyTorch, PaddlePaddle and JAX.

### How to install NVIDIA DALI?

The README does not contain install steps; it links to the documentation at docs.nvidia.com, and the repository includes conda/ and docker/ directories alongside the Python package. Follow the documentation page for the command matching your CUDA and framework versions.

### How does NVIDIA DALI compare with the PyTorch DataLoader?

The PyTorch DataLoader runs augmentation in Python worker processes on the CPU, while DALI executes a graph through its own engine with decoding and augmentation on the GPU. DALI's README calls it a portable drop-in replacement, but the examples define a DALI pipeline and use DALIGenericIterator rather than a Dataset subclass.

### What is an alternative to NVIDIA DALI?

The framework-native loaders are the direct alternative: torch.utils.data.DataLoader with a Dataset for PyTorch, and the equivalent built-in iterators in TensorFlow or PaddlePaddle. They run on the CPU and work on any device, which is the difference that matters when you are not on NVIDIA hardware.

## Sources

- [License: Apache-2.0](https://github.com/NVIDIA/DALI/blob/main/LICENSE)
- [NVIDIA/DALI on GitHub](https://github.com/NVIDIA/DALI)
- [Project website](https://docs.nvidia.com/deeplearning/dali/user-guide/docs/index.html)
- [README](https://github.com/NVIDIA/DALI/blob/main/README.md)
- [Releases](https://github.com/NVIDIA/DALI/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/nvidia-dali
