Library / SDK
deepseek-ai/FlashMLA avatar
deepseek-ai/FlashMLA

FlashMLA: DeepSeek's Optimised Attention Kernels for GPU Inference

FlashMLA: Efficient Multi-head Latent Attention Kernels

12,960 stars1,168 forksC++MIT

At a glance

What is it?
FlashMLA is a CUDA kernel library from DeepSeek that provides optimised sparse and dense attention implementations for the H800 and B200 GPU architectures, powering the DeepSeek-V3 and later model families. It is a low-level component, not a standalone model or inference framework.
Who is it for?
FlashMLA is the right component for teams deploying DeepSeek-V3, V3.2, V4, or V4.1 at scale on H800 or B200 hardware who need to replace a slower attention implementation. It requires CUDA 12.8 or later, SM90 or SM100 architecture, and PyTorch 2.0 or later.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 15 days ago.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

What FlashMLA Is and Why It Exists

FlashMLA is DeepSeek's library of GPU attention kernels, written to accelerate inference on the DeepSeek model family. The README states that it powers DeepSeek-V3 and DeepSeek-V3.2-Exp. The library provides two categories of kernels: sparse attention kernels that implement DeepSeek Sparse Attention (DSA) introduced with DeepSeek-V3.2, and dense attention kernels for prefill and decoding stages.

The project is not a model, a framework, or a serving stack. It is a collection of CUDA kernels that a serving system calls during attention computation. Teams that want to use it must integrate it into an existing inference stack such as vLLM or a custom implementation. The README covers the Python API for the decoding kernel but does not provide a full serving stack example.

The repository documents a performance update from April 2025: a kernel release that achieved 5 to 15 percent improvement for compute-bound workloads on H800 SXM5, with an interface that remained fully compatible with the previous version. This means upgrading FlashMLA within a minor version does not require changes to calling code.

Sparse and Dense Kernel Types in the Support Matrix

The repository contains five kernel variants as of the most recent documented state. Dense decoding targets SM90 (H800 class) in MQA mode and supports DeepSeek V3 and V3.1. Sparse decoding targets both SM90 and SM100 (B200 class) and supports V3.2, V4, and V4.1, with an FP8 KV cache option. Dense prefill targets SM100 only, in MHA mode, for V3 through V3.2. Sparse prefill targets SM90 and SM100 for V3.2, V4, and V4.1. The fused norm-RoPE-attn-RoPE-cast kernel targets SM100 only and is for V4 and V4.1.

This matrix has a concrete implication: using sparse decoding for DeepSeek V4.1 requires SM100 hardware. The dense decoding kernel uses bf16 or fp16 KV caches, while the sparse decoding kernel uses an FP8 format with a scale factor described in the README.

The README notes that the MLA mode differs between model generations. V3 and V3.1 use head dimensions of 576 for keys and 512 for values. V4 and V4.1 use 512 for keys and 512 for values. The calling code must pass the correct dimensions.

Installing FlashMLA

Installation requires cloning the repository and building from source:

bash
git clone https://github.com/deepseek-ai/FlashMLA.git flash-mla
cd flash-mla
git submodule update --init --recursive
pip install -v .

The build requires CUDA 12.8 or later. SM100 kernels additionally require CUDA 12.9 or later. The setup.py reads the NVCC version at build time and enforces this constraint. If CUDA 12.9 is not present and `FLASH_MLA_DISABLE_SM100` is not set to 1, the build will abort with an error message.

The build can be customised through environment variables. `FLASH_MLA_DISABLE_SM100=1` skips SM100 compilation for environments that only have CUDA 12.8. `FLASH_MLA_DISABLE_SM90=1` skips SM90 targets. `FLASH_MLA_DISABLE_FP16=1` disables FP16 support by passing the corresponding compile flag.

The MLA Decoding API

The README provides a concrete usage example for the decoding path. Before the decoding loop, you call `get_mla_metadata` once to compute tile scheduler metadata:

python
from flash_mla import get_mla_metadata, flash_mla_with_kvcache

tile_scheduler_metadata, num_splits = get_mla_metadata(
    cache_seqlens,
    s_q * h_q // h_kv,
    h_kv,
    h_q,
    is_fp8,
    topk,
)

Then at each decoding step across layers:

python
o_i, lse_i = flash_mla_with_kvcache(
    q_i, kvcache_i, block_table, cache_seqlens, dv,
    tile_scheduler_metadata, num_splits,
    is_causal, is_fp8_kvcache, indices,
)

The `s_q` parameter is the number of query tokens per sequence, which should be 1 when multi-token prediction (speculative decoding) is disabled. The FP8 KV cache path requires the `indices` argument for sparse attention; the dense decoding kernel does not support FP8.

The fused norm-RoPE-attn-RoPE-cast kernel requires TileLang, TileKernels, and DeepGEMM as additional dependencies. These are not installed by the main pip install step.

Performance Numbers from the Repository

The README documents benchmark commands and their results for each kernel type. For the dense MLA decoding kernel on H800 SXM5 with CUDA 12.8, the README states: up to 3000 GB/s in memory-bound configuration and up to 660 TFLOPS in compute-bound configuration. The sparse decoding kernel (with FP8 KV cache) achieves 410 TFLOPS in compute-bound configuration on H800 SXM5 and up to 700 TFlops on B200.

For the sparse prefill kernel, the README documents up to 640 TFLOPS on H800 SXM5 with CUDA 12.8 and up to 1450 TFlops on B200 with CUDA 12.9. For the dense prefill kernel on B200, NVIDIA reported up to 1460 TFLOPS forward and 1000 TFLOPS backward.

The test scripts to reproduce these benchmarks are in the `tests/` directory:

bash
python tests/test_flash_mla_dense_decoding.py
python tests/test_flash_mla_sparse_decoding.py

These numbers are reported for specific hardware configurations. Performance on different hardware or with different batch sizes will vary.

Hardware Constraints and When FlashMLA Does Not Apply

FlashMLA is architecturally specific. The minimum GPU requirement is SM90, which corresponds to NVIDIA Hopper architecture (H100 and H800). Consumer GPUs based on Ada Lovelace (RTX 4000 series, SM89) or Ampere (RTX 3000 series, SM86/SM80) are not supported. SM100 (Blackwell, B200) adds support for the SM100-specific kernels but requires CUDA 12.9.

The kernels are implemented for the DeepSeek model architecture specifically. They use the MQA attention mode with the head dimensions of the DeepSeek models. The head dimensions are 576 for keys and 512 for values for V3, V3.1, and V3.2, and 512 for both keys and values for V4 and V4.1. The fused norm-RoPE-attn-RoPE-cast kernel additionally requires TileLang, TileKernels, and DeepGEMM dependencies that are separate from the main pip install step.

The kernels are not a general-purpose drop-in for arbitrary transformer architectures. Teams working with LLaMA, Mistral, Qwen, or other model families need different kernel implementations. The MQA mode used by DeepSeek, where the head_dim_k and head_dim_v follow the DeepSeek-specific layout, is not the standard multi-head attention layout used by most other open model families.

FlashInfer is a related project that provides a broader set of attention kernels for different architectures and attention patterns. The README does not compare the two, but FlashInfer covers a wider range of hardware and model configurations than FlashMLA.

Editorial conclusion

FlashMLA is the right component for teams deploying DeepSeek-V3, V3.2, V4, or V4.1 at scale on H800 or B200 hardware who need to replace a slower attention implementation. It requires CUDA 12.8 or later, SM90 or SM100 architecture, and PyTorch 2.0 or later. It is not useful for other model families, does not run on consumer GPUs without the SM90 constraint, and is not a drop-in solution for teams without deep knowledge of attention kernel integration. Before using FlashMLA in a production serving stack, run the provided test scripts against your hardware configuration to confirm that the performance characteristics match the numbers documented in the README.

Frequently asked questions

flashinfer vs flash mla

FlashMLA is DeepSeek's own kernel library specifically for its MLA attention variant on Hopper and Blackwell GPUs (SM90 and SM100). FlashInfer is a separate project that provides attention kernels across a broader set of architectures and model types. FlashMLA is not a general-purpose replacement for FlashInfer.

Does FlashMLA work on consumer NVIDIA GPUs?

No. FlashMLA requires SM90 (Hopper, such as H100 and H800) or SM100 (Blackwell, such as B200). Consumer GPUs based on Ada Lovelace (RTX 4090, SM89) or earlier architectures are not in the support matrix.

Does FlashMLA support FP8 inference?

The sparse decoding kernel supports an FP8 KV cache format. The README describes the on-disk layout for the FP8 format and notes that it dequantizes to bfloat16 before performing attention. FP8 KV cache is only supported together with the sparse (indexed) attention path, not the dense decoding kernel.

Official sources

  1. deepseek-ai/FlashMLA on GitHub
  2. Issues
  3. License: MIT
  4. README
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/deepseek-ai-flashmla.svg)](https://hysenlabs.com/projects/deepseek-ai-flashmla)
Community notes

Community notes