NVIDIA CUTLASS: CUDA Templates and the CuTe Python DSL for GEMM Kernels
CUDA Templates and Python DSLs for High-Performance Linear Algebra
At a glance
- What is it?
- CUTLASS is NVIDIA's library of CUDA C++ template abstractions for high-performance GEMM and related linear algebra, plus a Python DSL family called CuTe DSL. It is a tool for kernel authors, not a drop-in BLAS replacement.
- Who is it for?
- Adopt CUTLASS if you are writing or tuning CUDA kernels for GEMM, convolutions, or attention and you need control over tiling, data types, and the thread and data hierarchy. Do not adopt it if you want a drop-in BLAS call, if your target is not an NVIDIA GPU, or if you cannot pin a CUDA toolkit and driver combination.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 6 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
What CUTLASS solves, and who is expected to use it
CUTLASS exists for the gap between a naive CUDA kernel and a vendor-tuned library call. The README describes it as "a collection of abstractions for implementing high-performance matrix-matrix multiplication (GEMM) and related computations at all levels and scales within CUDA," built around hierarchical decomposition and data movement. That framing matters: the project is not selling a function you call to get a fast multiply. It is selling the moving parts, decomposed into reusable components that a kernel author assembles and tunes.
The intended audience is narrow and technical. Someone writing a fused kernel, a custom epilogue, a mixed-precision path, or an attention variant needs control over tiling sizes, data types, and algorithmic policy. CUTLASS exposes exactly those knobs. A team that just wants a fast FP16 GEMM on a supported GPU has no reason to be here; a library call will do, and will do it with less work.
The C++ template abstractions have been shipping since 2017, and the README states they cover FP64, FP32, TF32, FP16, BF16, FP32 emulation via tensor core instructions, 8-bit floating point types (e5m2 and e4m3), block scaled types (NVFP4 and OCP MXFP4, MXFP6, MXFP8), narrow integers at 4 and 8 bits, and 1-bit binary types where the architecture supports them. That list spans Volta, Turing, Ampere, Ada, Hopper, and Blackwell, with Rubin support introduced in the 4.8 release notes.
The CuTe DSL: Python kernels without the C++ metaprogramming
CUTLASS 4 added a second surface: Python-native DSLs built on core CUTLASS and CuTe concepts. The README claims these offer "no performance compromises" while giving a smoother learning curve, faster compile times, and integration with deep learning frameworks without glue code. The first of the family, shipped in 4.0, is CuTe DSL, described as a low-level programming model consistent with the CuTe C++ abstractions: layouts, tensors, hardware atoms, and control over the hardware thread and data hierarchy.
The README states that CuTe DSL is currently in public beta and is expected to graduate out of beta by end of summer 2026. Treat that as a real constraint, not a footnote. The 4.8 notes describe an opt-in preview of the CuTe DSL extensions compiler pipeline, enabled with an environment variable, which lets users mix extension APIs directly into @cute.jit and @cute.kernel code. The notes say the pipeline is expected to preserve program behavior but that generated PTX and SASS may differ, and that it is planned to become the default in a future release. If you depend on instruction-level output stability, that sentence is the one to read twice.
The Python package metadata in pyproject.toml lists dependencies on cuda-python, networkx, numpy, pydot, scipy, and treelib, with requires-python >=3.8. Note the version string in that file: 4.2.0.0, while the README is for 4.8.0. The packaging metadata lags the release notes.
How the abstraction hierarchy is organized
The architecture is a hierarchy of decomposition. The README's opening image is captioned "Complete CUDA GEMM decomposition," and the text says primitives at different levels of a conceptual parallelization hierarchy can be specialized and tuned via custom tiling sizes, data types, and other algorithmic policy. In practice that means a GEMM is split across threadblock, warp, and instruction levels, with explicit data movement between memory levels and async copy abstractions for the transfers.
The repository layout reflects that layering. The include/ directory holds the C++ headers; examples/ contains numbered programs that walk from a basic GEMM (examples/00_basic_gemm/) through tile iterators, batched and split-K GEMM, tensorop GEMM for Volta and Turing, planar complex, and fusion examples such as examples/12_gemm_bias_relu/ and examples/13_two_tensor_op_fusion/. There are also visualization and diagnostic examples: examples/02_dump_reg_shmem/ and examples/03_visualize_layout/. The numbering is a reading order, and it is more useful than the prose in the README for understanding how the pieces fit.
Alongside include/ and examples/ sit operators/ and python/, plus tools/, test/, cmake/, and a cutlass_compiler/ directory. The presence of cuBLAS.cmake and cuDNN.cmake at the top level is worth noting: the build system can locate those libraries, which is relevant if you intend to compare against or interoperate with them.
Installing CUTLASS and running a first kernel
The README does not give install commands. It points to two guides: the CUTLASS C++ Quick Start Guide and the CuTe DSL Quick Start Guide, both hosted under docs.nvidia.com/cutlass. Those are the authoritative sources for build steps, and the exact CMake invocations are not reproduced in the README itself. What the repository does show is a Python packaging path via pyproject.toml, whose project name is nvidia-cutlass and whose declared dependencies are listed above.
For the C++ side, the repository is a CMake project: CMakeLists.txt at the top level, with CUDA.cmake, cuBLAS.cmake, cuDNN.cmake, customConfigs.cmake, and bin2hex.cmake as supporting modules, plus a cmake/ directory. The examples are built through that CMake tree rather than through any per-example script the README documents.
For the DSL side, the 4.8 notes give one concrete command, for testing the opt-in extension compiler pipeline. The notes state you can try it with:
CUTE_DSL_USE_EXTENSION_COMPILER=1 python your_program.pyThat variable switches the compiler pipeline used for CuTe DSL kernels. The notes say the pipeline is expected to preserve program behavior, but that generated PTX and SASS may differ from the default path, and that it is planned to become the default in a future release. If you are evaluating it, run the same program with and without the variable and compare the emitted PTX or SASS rather than assuming equivalence.
The pyproject.toml also declares the package metadata a packaging tool would read:
[project]
name = "nvidia-cutlass"
version = "4.2.0.0"
requires-python = ">=3.8"
license = {text = "BSD-3-Clause"}That version field does not match the 4.8.0 README, so do not use it to decide which release you are installing.
Where CUTLASS is the wrong tool
The clearest failure mode is treating CUTLASS as a BLAS. It is not one. There is no documented single call that replaces a GEMM routine for an arbitrary shape. You pick a configuration, instantiate templates, and build. That is a compile-time cost and a maintenance cost, and for a straightforward dense multiply on a supported architecture the trade is usually bad.
Architecture coverage is another boundary. The README lists Volta, Turing, Ampere, Ada, Hopper, and Blackwell for the C++ abstractions, and the 4.8 notes add Rubin (SM107) support in CuTe DSL and primitives. But the notes carry a hard constraint on that Rubin path: executing Rubin kernels requires the R615 driver, which the notes say will be released with CUDA Toolkit 13.4 GA, and they explicitly state that R610 from the CUDA Toolkit 13.4 Developer Preview is not sufficient. If you are on a preview driver, the Rubin examples will not run, regardless of what the code compiles to.
The CuTe DSL beta status is a third boundary. The README says the DSLs are in public beta and expected to graduate by end of summer 2026. A team that needs a stable API surface for a long-lived internal kernel should weigh that. The extension compiler pipeline being opt-in today and planned as default later means the code path you validate now may not be the default path you get later.
Finally, the README does not document a rollback or downgrade procedure between CUTLASS versions, and it does not document a compatibility matrix tying a CUTLASS release to a specific CUDA toolkit beyond the driver note for Rubin. If you need that matrix, the README is silent and you will have to assemble it from the release notes and the docs site.
How it differs from cuBLAS and from hand-written CUDA
The obvious alternative is cuBLAS, which NVIDIA also ships. The difference is in the level of control. cuBLAS gives you a call and a set of parameters; it decides the tiling, the scheduling, and the instruction selection. CUTLASS gives you the decomposition itself, so you can fuse an epilogue, change the data layout, or insert a custom reduction. The repository even carries cuBLAS.cmake, which suggests the two are expected to coexist rather than replace one another: you use cuBLAS where a call suffices and CUTLASS where it does not.
The second alternative is writing the CUDA kernel by hand. CUTLASS's value there is that the reusable primitives already encode the async copy, multiply-accumulate, and data-movement patterns for each architecture. Writing those from scratch means rediscovering the same structure. The cost is that you now depend on CUTLASS's abstractions and their evolution, including the beta DSL surface.
A third comparison is within CUTLASS itself: C++ templates versus CuTe DSL. The README positions the DSL as giving faster compile times and a smoother learning curve, with the claim of no performance compromise. The 4.8 notes list a large set of new examples for the DSL across Rubin, Blackwell, and Ampere, including dense GEMM, blockscaled GEMM, grouped GEMM, attention (GQA decode), and top-K. That example inventory is the best available signal of which paths are actually exercised.
Maintenance, releases, and licence
The repository is not archived, and the last push was on 2026-09-08. Releases are frequent and versioned in three tracks at once: v4.8.0dev (2026-08-27), v4.7.1 (2026-08-26), and v4.6.3 (2026-08-26). A dev tag alongside two patch tags on adjacent days indicates parallel maintenance of a current line and at least one older line. For an adopter, that means upgrade cost is real: you should expect to re-validate kernels against a new release rather than assume drop-in compatibility, and the README does not offer a migration guide between minor versions.
Licensing needs care. The GitHub metadata reports the licence as NOASSERTION, meaning the platform could not classify it automatically. The repository contains both LICENSE.txt and EULA.txt, and pyproject.toml declares BSD-3-Clause for the Python package. Those three signals do not obviously agree, and the README does not reconcile them. Before shipping CUTLASS inside a product, read LICENSE.txt and EULA.txt directly and have someone qualified determine which terms apply to the components you use. The presence of an EULA alongside a BSD-style declaration is the specific thing to resolve.
The CHANGELOG.md and CITATION.cff files at the top level are the other maintenance artifacts worth reading; the README's "What's New" section covers only the current release and is truncated in the repository listing.
Editorial conclusion
Adopt CUTLASS if you are writing or tuning CUDA kernels for GEMM, convolutions, or attention and you need control over tiling, data types, and the thread and data hierarchy. Do not adopt it if you want a drop-in BLAS call, if your target is not an NVIDIA GPU, or if you cannot pin a CUDA toolkit and driver combination. Before committing, verify three things: that your GPU architecture appears in the supported list, that your required data types are covered at that architecture level, and, for CuTe DSL work, that you are willing to track a public beta whose extension compiler pipeline is opt-in today and planned to become the default in a future release.
Frequently asked questions
How do I install NVIDIA CUTLASS?
The README does not give install commands. It points to the CUTLASS C++ Quick Start Guide and the CuTe DSL Quick Start Guide on docs.nvidia.com/cutlass for getting started. The repository is a CMake project with a top-level CMakeLists.txt, and pyproject.toml declares a Python package named nvidia-cutlass.
How do I use NVIDIA CUTLASS?
You use it by assembling CUDA C++ template abstractions for GEMM and related computations, tuning tiling sizes and data types, and building through the CMake tree. For the Python path, the README describes CuTe DSL as a low-level programming model exposing layouts, tensors, and hardware atoms, and the 4.8 notes give CUTE_DSL_USE_EXTENSION_COMPILER=1 python your_program.py for testing the opt-in extension compiler pipeline.
What is NVIDIA CUTLASS?
The README describes it as a collection of abstractions for implementing high-performance matrix-matrix multiplication and related computations at all levels and scales within CUDA, built around hierarchical decomposition and data movement. Since CUTLASS 4 it also includes Python-native DSLs, of which CuTe DSL was the first.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/nvidia-cutlass)
Community notes