Library / SDK
NVIDIA/cudnn-frontend avatar
NVIDIA/cudnn-frontend

cuDNN Frontend: NVIDIA's Open-Source C++ and Python Graph API for Deep Learning Kernels

cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.

953 stars299 forksPythonApache-2.0

At a glance

What is it?
cuDNN Frontend is NVIDIA's modern entry point to the cuDNN library, exposing a header-only C++ API and a Python interface with native PyTorch integration, alongside a growing set of open-source high-performance kernels for attention, MoE GEMM, and normalization on Hopper and Blackwell GPUs.
Who is it for?
cuDNN Frontend is appropriate for deep learning framework developers and ML infrastructure engineers who need to compose fused GPU kernels for attention, GEMM, and normalization with precise control over precision and layout, without writing raw CUDA. The Python path via pip install nvidia-cudnn-frontend is the lowest-friction entry point.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What cuDNN Frontend Is and Who It Is For

cuDNN is NVIDIA's library for deep learning primitives: convolutions, attention, normalization, and related operations. Accessing it directly through the cuDNN C API requires building execution plans, selecting algorithms, and managing workspace memory at a low level. The cuDNN Graph API is a higher-level abstraction that expresses operations as a graph and lets cuDNN select and compile efficient execution plans automatically.

cuDNN Frontend is the open-source layer that makes the Graph API usable from both C++ and Python. The README describes it as NVIDIA's modern, open-source entry point to the cuDNN library plus a growing collection of high-performance open-source kernels. The keyword here is open-source: the kernels can be inspected, modified, and contributed to, which is a change from earlier cuDNN components that were closed binary releases.

The primary audience is deep learning framework developers (implementing or optimizing attention, GEMM, or normalization layers in PyTorch, JAX, or custom frameworks) and ML infrastructure engineers building efficient model serving or training kernels for NVIDIA hardware. End users of existing frameworks do not interact with cuDNN Frontend directly; they benefit through framework integrations.

C++ Header-Only API and Python Interface

The C++ API is header-only: including the cudnn_frontend directory in a C++ project's include path is sufficient to use the Graph API without linking against a cuDNN Frontend library. The C++ API targets the cuDNN Graph API version, which abstracts individual operation nodes, their data descriptors, and their connections into a graph that cuDNN compiles into a kernel plan.

The Python interface, available as the nvidia-cudnn-frontend package on PyPI, wraps the same Graph API through pybind11 bindings. The README describes native PyTorch integration as part of the Python interface, meaning PyTorch tensors can be passed directly to the graph execution without explicit data conversion. The Python package requires Python 3.10 or higher.

For Python users, the entry point is the `cudnn.pygraph` API. The FROST GEMM engine and other open-source kernels are reachable through the same `cudnn.pygraph` interface when enabled, so switching between the backend's own plans and the open-source kernels does not require changing the calling code.

Installing cuDNN Frontend via PyPI

The Python package is available as nvidia-cudnn-frontend on PyPI. The current dependency list in pyproject.toml requires nvidia-cutlass-dsl[cu13]>=4.6.2 and apache-tvm-ffi>=0.1.11. These are heavier dependencies than a typical Python package; the cutlass-dsl package in particular pulls in NVIDIA CUTLASS DSL infrastructure.

The requirements.txt for development adds pybind11, pytest, and cuda-python. For C++ builds, the CMakeLists.txt at the repository root drives the build, and the setup.py wraps it for the Python package build.

The project is on a rolling development cadence, with release tags like v1.31.0.dev in the September 2026 timeframe. The last push was on 2026-09-27, one day before this writing, confirming active development. For production use, pinning to a specific release tag is advisable; development builds may carry breaking changes.

Open-Source Kernels: Flash Attention, MoE GEMM, and Normalization

The most widely used kernel in cuDNN Frontend is scaled dot-product attention (SDPA), also referred to as Flash Attention in the README. The open-source implementation covers Hopper (SM90) and Blackwell (SM100/SM103) architectures, including the backward pass for D=256 on SM100.

For MoE (Mixture-of-Experts) workloads, the repository includes a range of grouped GEMM fusions. The README lists separate implementations for grouped GEMM with SwiGLU, with GLU, with sReLU, with dsReLU, with backward passes for each, and with quantization outputs. Dense and discrete MoE weight layouts are both supported.

The BSA (block-sparse attention) implementation supports block-level routing metadata for forward and backward passes using CuTe DSL. The NSA (Native Sparse Attention) implementation is described in a research paper on hardware-aligned trainable sparse attention. Flex Attention supports reusable interval-mask plans with PyTorch autograd across SM90, SM100, and SM103.

For normalization, the repository includes a fused RMSNorm + SiLU kernel and DSA/CSA kernels for DeepSeek model variants. The HSTU Attention implementation targets packed variable-length sequences specifically on Blackwell SM100 and SM103.

The FROST GEMM Engine for Blackwell

The README highlights the FROST GEMM engine as a new open-source JIT-compiled GEMM engine for Blackwell GPUs. FROST compiles matmul kernels at runtime for the specific graph configuration in use, covering standard matmul, grouped (MoE) matmul, block-scaled FP4/FP8, and chained pointwise epilogues fused into a single kernel.

FROST is opt-in: it requires setting the environment variable CUDNN_FRONTEND_ENABLE_FROST_ENGINES=1 before running. Once enabled, it becomes a candidate for every matmul graph that can be served, ranked against the cuDNN backend's existing plans. The README states the graph built with the ordinary `cudnn.pygraph` API is unchanged when FROST is enabled; the engine selection happens inside cuDNN's planning phase.

The Blackwell targeting is explicit: SM100 and SM103 (the B200/GB200/GB300 family). Using FROST on Hopper hardware is not described in the README. This makes FROST a forward-looking addition for teams deploying on the Blackwell generation.

Precision Targets and Platform Limitations

The full precision coverage listed in the README is FP16, BF16, FP8, and MXFP8 (microscaling FP8). FP8 and MXFP8 operations require Hopper or Blackwell GPU hardware: the H100, H200, B200, GB200, and GB300 product lines. FP16 and BF16 operations work on a wider range of NVIDIA hardware, but the cuDNN Frontend library itself requires a compatible cuDNN version and a modern CUDA toolkit.

The README notes the supported GPU architectures as Hopper (H100/H200) and Blackwell (B200/GB200/GB300), with Rubin mentioned in the PyPI package description as a future target. Consumer GPUs from earlier generations (Ampere, Ada Lovelace) can use the library but will not benefit from FP8 paths or Blackwell-specific kernels.

A practical alternative for teams that want Flash Attention specifically without the full cuDNN dependency is the flash-attention package, which also runs on Hopper and Ampere. Flash Attention handles the attention kernel only; cuDNN Frontend covers attention plus MoE GEMM, normalization, and the full breadth of cuDNN's operation graph, making it the broader tool for framework-level integration.

Editorial conclusion

cuDNN Frontend is appropriate for deep learning framework developers and ML infrastructure engineers who need to compose fused GPU kernels for attention, GEMM, and normalization with precise control over precision and layout, without writing raw CUDA. The Python path via pip install nvidia-cudnn-frontend is the lowest-friction entry point. The Hopper and Blackwell targeting means consumer and data center GPUs before the H100 generation will not benefit from most of the advanced precision formats. The FROST engine requires an opt-in environment variable and targets Blackwell specifically; it is not a drop-in for Hopper environments. Anyone integrating this into a framework should review the dual Apache-2.0 and MIT licensing before redistribution.

Frequently asked questions

What is cuDNN Frontend?

cuDNN Frontend is NVIDIA's open-source C++ and Python interface to the cuDNN Graph API, plus a growing collection of open-source high-performance kernels for attention, grouped GEMM, and normalization, targeting Hopper and Blackwell GPUs.

What is Nvidia cuDNN?

cuDNN is NVIDIA's library of deep learning primitives for GPU-accelerated operations such as convolutions, attention, and normalization. cuDNN Frontend is the open-source layer that makes the cuDNN Graph API accessible from C++ and Python.

What is the difference between cuDNN and CUDA?

CUDA is NVIDIA's parallel computing platform and programming model for writing general-purpose GPU code. cuDNN is a higher-level library of deep learning operations built on top of CUDA, providing optimized implementations of convolutions, attention, and other neural network primitives.

Official sources

  1. License: Apache-2.0
  2. NVIDIA/cudnn-frontend on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/nvidia-cudnn-frontend.svg)](https://hysenlabs.com/projects/nvidia-cudnn-frontend)