DeepGEMM: runtime-compiled FP8 and FP4 tensor core kernels for Hopper and Blackwell
DeepGEMM: clean and efficient BLAS kernel library on GPU
At a glance
- What is it?
- DeepGEMM is DeepSeek's CUDA kernel library for the matrix primitives behind large language models. It compiles at runtime through DeepJIT, so installation needs no CUDA build, but it only targets SM90 and SM100 GPUs.
- Who is it for?
- Adopt DeepGEMM if you run inference or training on SM90 or SM100 hardware and your bottleneck is the GEMM, grouped MoE or MQA scoring kernel itself. Do not adopt it on consumer GPUs, on SM120, or if you need the library to handle transposition and FP8 casting for you, because the README states those must be done by the user.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 6 days ago.
- What is it written in?
- Mainly Cuda, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What DeepGEMM is for, and who should reach for it
DeepGEMM is a CUDA kernel library, not a framework. It collects the tensor core primitives that show up repeatedly in large language model execution: FP8, FP4 and BF16 GEMMs, fused MoE with overlapped communication (Mega MoE), MQA scoring for the lightning indexer, and HyperConnection. The README describes it as a unified, high-performance tensor core kernel library built as one cohesive CUDA codebase.
The audience is narrow and specific. You need an NVIDIA SM90 or SM100 GPU, Python 3.8 or higher, C++20 support in your compilers and standard libraries, CUDA Toolkit 12.9 or higher, PyTorch 2.3 or higher, and CUTLASS 4.0 or higher. If your hardware is older, or if you are on a consumer card, the library has nothing to offer you. The README's own framing is telling: DeepGEMM borrows concepts from CUTLASS and CuTe but deliberately avoids heavy reliance on their templates, and it keeps the number of core kernel functions small so that it stays readable as a learning resource for GPU kernel optimization. That dual purpose matters when you evaluate it. You are adopting both a production kernel set and a codebase meant to be studied.
How DeepJIT and the GEMM interfaces actually work
The mechanism that separates DeepGEMM from a conventional CUDA extension is DeepJIT. Kernels are compiled at runtime, which is why the README states no CUDA compilation is required during installation. The trade-off is a first-call compile cost that the user absorbs at runtime rather than at install time. The release notes list faster JIT compilation as a recurring theme, including a low-CPU-overhead JIT CPP module in the 2025.07.20 refactor, which suggests the cost was noticeable enough to keep working on.
The naming convention is `D = C + A @ B`, with an NT layout by default: non-transposed A, transposed B. So `fp8_gemm_nt` computes `D = C + A @ B.T`. The SM90 implementation supports only the NT memory layout. The SM100 implementation supports all four: NT, TN, NN and TT. That is a real capability gap between the two architectures, not a documentation detail.
Scaling factors are where the two architectures diverge again. Both require the LHS scaling factor in a TMA-aligned, transposed layout, but SM90 wants FP32 while SM100 wants packed UE8M0 format, four UE8M0 values packed into a single `torch.int`. Getting this wrong is a silent correctness problem, not a crash. The README also states plainly that input transposition and FP8 casting must be handled by the user, and that the PyTorch utility functions the library ships may be slower than a fused implementation. The library optimizes the GEMM kernel and expects you to optimize everything around it.
Grouping works differently from what CUTLASS users expect. DeepGEMM groups only the M-axis; N and K must stay fixed. That fits MoE experts that share a shape. For the contiguous layout, tokens from different experts are concatenated into one tensor and each expert segment must be aligned to the GEMM M block size, which you query with `get_mk_alignment_for_contiguous_layout()`. For inference decoding under CUDA graphs, where the CPU does not know how many tokens each expert receives, the masked variant `m_grouped_fp8_gemm_nt_masked` takes a mask tensor and computes only the valid portions. The README names DeepEP's low-latency kernel output as one intended input to that path. A separate K-axis-grouped API, `k_grouped_fp8_gemm_tn_contiguous`, exists for MoE weight backward, with M and N fixed instead.
Installing DeepGEMM and running a first grouped FP8 GEMM
The repository ships two shell scripts. The README's development path clones the submodules, which is required, then links includes and builds the C++ extension:
git clone --recursive [email protected]:deepseek-ai/DeepGEMM.git
cd DeepGEMM
cat develop.sh
./develop.shFor a normal installation the README points at `install.sh`:
cat install.sh
./install.shAfter that, the README says to import `deep_gemm` in your Python project. Because kernels compile through DeepJIT at runtime, the first invocation of a given kernel is where compilation happens, and you should expect that latency rather than treat it as a hang.
The setup script exposes environment variables that change the build. `DG_SKIP_CUDA_BUILD=1` skips the CUDA build, `DG_FORCE_BUILD=1` forces it, and `DG_USE_LOCAL_VERSION` defaults to 1. The build pulls includes from `deep_gemm/include`, `third-party/deep_jit/include` and `third-party/cutlass/include`, compiles `csrc/python_api.cpp` with `-std=c++20`, and links `cudart`. If you set `DG_SKIP_CUDA_BUILD`, the wheel URL pattern in `setup.py` points at GitHub releases under the `DeepSeek-AI/DeepGEMM` repository, which is how a prebuilt path avoids local compilation.
A first real use is a grouped FP8 GEMM in contiguous layout, which is the MoE training-forward and inference-prefill case. The README does not print a full call signature, so the honest starting point is the documented function name and the alignment helper. Before calling it, query the required M block alignment and pad your expert segments to it:
import deep_gemm
alignment = deep_gemm.get_mk_alignment_for_contiguous_layout()
# concatenate expert token segments, each padded to `alignment`
# deep_gemm.m_grouped_fp8_gemm_nt_contiguous(...)The README directs you to the function documentation for the full argument list; it is not reproduced in the README itself. Plan to read `deep_gemm/` and `docs/` in the repository rather than working from the README alone.
The constraints that decide whether DeepGEMM fits
The hard boundary is architecture. SM90 or SM100 only. The related searches include DeepGEMM SM120 and DeepGEMM H20, and the README offers nothing for either; there is no fallback path documented for other compute capabilities. If your fleet is mixed, you are maintaining two code paths, and the SM90 path is the weaker one: NT layout only, FP32 scaling factors, no TN, NN or TT.
The second constraint is that DeepGEMM is a kernel library with a deliberately thin surface. Transposition and FP8 casting are your problem. The README is explicit that the bundled PyTorch utilities may be slower and that the project's focus is the GEMM kernels. If you were hoping to hand it a float tensor and get an FP8 result, this is the wrong tool. The same applies to the grouped APIs: N and K fixed for M-grouped, M and N fixed for K-grouped. Shapes that vary along the wrong axis are outside the design.
Third, the masked grouped GEMM exists because CUDA graph decoding hides the per-expert token count from the CPU. That is a narrow, well-defined use case, and reaching for it outside a graph-captured decode loop adds a mask tensor you do not need.
Finally, the runtime compilation model means the first call in a fresh process pays a compile cost. The release notes list faster JIT compilation and a low-CPU-overhead JIT module as improvements, which tells you the cost has been real. If your serving path cannot tolerate a slow first request, you need a warmup step before you take traffic. The README does not document a warmup procedure.
DeepGEMM versus CUTLASS and Triton
CUTLASS is the closest comparison and the one DeepGEMM positions itself against directly. The README says DeepGEMM leverages concepts from CUTLASS and CuTe but avoids heavy reliance on their templates or algebras. That is the substantive difference: CUTLASS gives you a template metaprogramming framework for composing kernels, and DeepGEMM gives you a small set of concrete kernels with hand-written CUDA. The cost is flexibility. CUTLASS can be bent toward shapes and layouts DeepGEMM does not cover, and CUTLASS grouped GEMM does not restrict you to grouping the M-axis alone. The benefit is that DeepGEMM's kernels are tuned for the specific shapes MoE inference produces, and the codebase is small enough to read. CUTLASS is also a build-time dependency here, vendored as a submodule at version 4.0 or higher.
Triton is the other comparison the search data raises, and the difference is in kind rather than degree. Triton compiles Python-level tile programs to GPU code and lets you express new kernels without writing CUDA. DeepGEMM is a fixed library of CUDA kernels with a runtime JIT, not a language for writing your own. If your problem is a novel fused operator, Triton is the general tool and DeepGEMM has no answer. If your problem is a standard FP8 GEMM or a grouped MoE GEMM on Hopper or Blackwell and you want the last increment of performance, DeepGEMM is the specialist. The README claims performance that matches or exceeds expert-tuned libraries across various matrix shapes, and the release notes cite up to 1550 TFLOPS on H800 as of the 2025.04.18 update. Those are the project's own numbers from its own pull requests, not independent measurements.
Maintenance, licence and the cost of upgrading
The repository is not archived and the last push was on 2026-09-14, eight days before this writing, so it is under current development. The release tags are not semantic versions; they are named `nv_dev_<hash>`, with the most recent listed as `nv_dev_f8e8fb5` from 2026-07-20. The README's news entries track pull requests rather than releases, and the 2026.09.10 entry covers Sparse Indexer, Mega Gate, Mega mHC, DeepJIT and MoE and Indexer optimizations. The practical consequence is that pinning to a release tag gives you a commit identifier, not a version with a compatibility promise. You will be reading pull request numbers to understand what changed.
Upgrade cost is coupled to your toolchain. CUDA Toolkit 12.9 or higher is required, and a 2025.07.20 note states that NVCC 12.9 performs FFMA interleaving automatically, so all post optimizations are no longer supported. Moving to a newer toolkit can therefore change generated code in ways the library no longer tries to control. CUTLASS must be at 4.0 or higher, pulled as a submodule, and the include paths in `setup.py` reference `third-party/cutlass/include` and `third-party/deep_jit/include`, so a submodule update touches the build.
The licence is MIT. That is permissive and compatible with commercial use, but note that the vendored CUTLASS submodule carries its own licence terms, and the build links against CUDA components. This is not legal advice; check the licence files in the repository and in `third-party/` if redistribution matters to you.
Editorial conclusion
Adopt DeepGEMM if you run inference or training on SM90 or SM100 hardware and your bottleneck is the GEMM, grouped MoE or MQA scoring kernel itself. Do not adopt it on consumer GPUs, on SM120, or if you need the library to handle transposition and FP8 casting for you, because the README states those must be done by the user. Before committing, verify three things: that your GPU reports SM90 or SM100, that CUDA Toolkit 12.9 or higher and PyTorch 2.3 or higher are present, and that your scaling factors match the layout your architecture expects (FP32 on SM90, packed UE8M0 in a torch.int on SM100).
Frequently asked questions
How does DeepGEMM work?
It is a CUDA kernel library whose kernels compile at runtime through DeepJIT, so no CUDA compilation happens during installation. Kernels follow the convention D = C + A @ B with an NT layout by default, and the SM100 path additionally supports TN, NN and TT.
How do I install DeepGEMM?
Clone the repository with submodules using git clone --recursive, then run ./install.sh; the README also documents a development path through ./develop.sh. Requirements include CUDA Toolkit 12.9 or higher, PyTorch 2.3 or higher, Python 3.8 or higher and CUTLASS 4.0 or higher.
What is DeepGEMM?
It is a tensor core kernel library from DeepSeek that gathers the GEMM, fused MoE, MQA scoring and HyperConnection primitives used by large language models into one CUDA codebase. The README describes it as designed for simplicity, with a limited number of core kernel functions.
DeepGEMM vs Triton: what is the difference?
Triton is a way to write your own GPU kernels from Python-level tile programs, while DeepGEMM is a fixed set of hand-written CUDA kernels compiled at runtime. If you need a novel fused operator, DeepGEMM does not provide a language for expressing it.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/deepseek-ai-deepgemm)