Open-source project
Dao-AILab/flash-attention avatar
Dao-AILab/flash-attention

FlashAttention: installing the CUDA, ROCm and CuTeDSL builds

Fast and memory-efficient exact attention. FlashAttention This repository provides the official implementation of FlashAttention and FlashAttention-2 from the following papers.

25,011 stars3,100 forksPythonBSD-3-Clause

At a glance

What is it?
FlashAttention is an IO-aware exact attention kernel for PyTorch, shipped as four separate builds (v2, v3, v4 and the ROCm Triton backend). This article covers what each one requires, how to install it, and when you should not use it.
Who is it for?
Adopt FlashAttention if you train or serve transformer models on Ampere, Ada, Hopper or Blackwell NVIDIA GPUs, or on MI200/MI300/RDNA3/4 AMD GPUs under ROCm 6.0 or newer, and you can pin a CUDA or ROCm version that matches your PyTorch build. Do not adopt it if you are on Turing hardware and need the full feature set, on a platform where you cannot compile CUDA extensions, or on a CPU-only deployment.
Can I use it commercially?
Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 4 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What FlashAttention actually replaces in a PyTorch model

Attention in a standard PyTorch transformer materializes the full score matrix, which for long sequences costs memory proportional to sequence length squared. FlashAttention computes the same exact attention result, not an approximation, by tiling the computation and keeping intermediate values in on-chip memory rather than writing the full matrix to HBM. The repository describes itself as the official implementation of FlashAttention and FlashAttention-2 from the associated papers, and the same kernel family now spans four distinct code paths.

The audience is narrow but deep: people training or serving transformer models who have already hit a memory ceiling or a throughput ceiling on attention, and who control their own GPU stack. If you are calling a hosted API, or if your sequences are short enough that attention is not the bottleneck, the installation cost of this project will not pay for itself. The README notes that FlashAttention and FlashAttention-2 are free to use and modify under the LICENSE file, and asks that you cite and credit the project if you use it.

Four builds in one repository, and why that matters for install

The repository is not one package. The top level holds the FlashAttention-2 implementation under flash_attn/ and csrc/, a separate hopper/ directory for the FlashAttention-3 beta, and a CuTeDSL implementation released as flash-attn-4. Each has its own GPU requirements and its own install command, and the README treats them as separate products rather than versions of one thing.

The v2 path targets Ampere, Ada and Hopper GPUs, supports fp16 and bf16, and handles head dimensions up to 256. The v3 beta is restricted to H100 and H800 class Hopper GPUs and requires CUDA 12.3 or newer, with CUDA 12.8 recommended for best performance; it currently ships FP16 and BF16 forward and backward plus FP8 forward. The v4 CuTeDSL implementation targets both Hopper and Blackwell, so H100 and B200 class hardware. The ROCm side is itself split: a composable_kernel backend that is the default, and a Triton backend installed through the aiter package, which is pulled in as a git submodule at third_party/aiter.

That structure is the single most important thing to understand before running any install command. A wheel name, a CUDA version and a GPU generation all have to line up, and picking the wrong subdirectory gives you a package that imports but does not cover your hardware.

Installing flash-attn and running a first forward pass

The v2 package installs from PyPI. The README recommends disabling build isolation, because the build needs to see your installed PyTorch and CUDA toolkit rather than a fresh environment.

bash
pip install flash-attn --no-build-isolation

If you prefer to compile from source, the repository's setup.py supports the standard setuptools path. The README gives this as the alternative:

bash
python setup.py install

Compilation is the part people underestimate. The README states that without ninja, compiling can take a very long time, around two hours, because the build does not use multiple CPU cores. With ninja installed and working, the same build takes 3 to 5 minutes on a 64-core machine using the CUDA toolkit. The README also warns that on a machine with less than 96GB of RAM and many CPU cores, ninja can launch enough parallel jobs to exhaust memory. The documented fix is to cap the job count:

bash
MAX_JOBS=4 pip install flash-attn --no-build-isolation

Before any of this, confirm the prerequisites the README lists: a CUDA or ROCm toolkit, PyTorch 2.2 or above, and the packaging, psutil and ninja Python packages. The README suggests installing the NVIDIA PyTorch container if you want a known-good toolchain, since it ships the required tools. Once the package is in place, the interface lives at src/flash_attention_interface.py. The v4 package uses a different import path, which the README shows as:

python
from flash_attn.cute import flash_attn_func

out = flash_attn_func(q, k, v, causal=True)

That single call is the whole surface for the common case. Everything else in the repository is about which hardware that call runs on.

The Hopper beta and the CuTeDSL build are separate installs

FlashAttention-3 does not install from the repository root. The README instructs you to change into the hopper directory and run setup.py there, then set PYTHONPATH to the working directory before running the test file with pytest. The import name is flash_attn_3, not flash_attn, which means code written against v2 will not silently pick up v3.

bash
cd hopper
python setup.py install

The README also documents a uv-based install, declaring flash-attn-3 as a dependency and pointing uv at the hopper subdirectory as a git source, with no-build-isolation set to true. That is the cleanest route if your project already manages dependencies through uv, because it pins the subdirectory rather than asking you to remember it.

FlashAttention-4 is simpler to install than v3, because it publishes to PyPI directly. The README gives the plain install and a CUDA 13 variant with an extra:

bash
pip install flash-attn-4
pip install "flash-attn-4[cu13]"

The README recommends the cu13 extra for best performance on CUDA 13. Both v3 and v4 are described in the README as beta lines, and the release tags in the repository carry beta numbering, so the install commands are stable but the interfaces behind them are still moving.

ROCm installs, and the Triton backend's real constraint

On AMD hardware the install path forks. The README tells you to get PyTorch for ROCm from the PyTorch site first, then install FlashAttention with an environment variable that selects the Triton backend:

bash
cd flash-attention
FLASH_ATTENTION_TRITON_AMD_ENABLE="TRUE" pip install --no-build-isolation .

ROCm 6.0 or above is the stated requirement. The composable_kernel backend, which is the default, supports MI200x, MI250x, MI300x, MI355x and RDNA 3/4 GPUs with fp16 and bf16, and both forward and backward head dimensions up to 256. The Triton backend covers CDNA (MI200, MI300) and RDNA GPUs with fp16, bf16 and fp32, and the README lists a broad feature set: causal masking, variable sequence lengths, arbitrary Q/KV lengths and head sizes, MQA/GQA, dropout, rotary embeddings, ALiBi, paged attention, and FP8 through the FlashAttention v3 interface.

The honest limitation is stated in the README itself: sliding window attention is currently a work in progress on the Triton backend. If your model uses a sliding window pattern, this backend is not the one to plan around. There is also a maintenance cost that the README flags directly. The Triton kernels come from the aiter package, included as a git submodule, and the README documents how to check out a specific aiter commit for testing or development. That means your build is pinned to a submodule revision as well as a FlashAttention revision. The full ROCm test suite, per the README, takes hours.

Where FlashAttention is the wrong choice

The clearest boundary is hardware. Turing GPUs such as the T4 and RTX 2080 are not covered by the main CUDA path; the README points to a separate flash-attention-turing repository maintained by a different author, described as supporting a core subset of FlashAttention features. If you are on Turing and need the full feature set, this repository is not the answer.

Windows is a second boundary. The README states that the code might work on Windows starting with v2.3.2 and references a few positive reports, but that Windows compilation still requires more testing. It then asks for help setting up prebuilt CUDA wheels for Windows. That is not a supported configuration, and treating it as one will cost you time in the build toolchain rather than in your model.

A third boundary is the build itself. This project compiles CUDA or HIP extensions against a specific toolkit version. If your deployment environment cannot run a compiler with the right toolkit, or if your PyTorch is older than 2.2, the install will fail before any kernel runs. The README's own recommendation of the NVIDIA PyTorch container is effectively an admission that matching these versions by hand is the hard part.

Finally, the version split is a trap for anyone who skims. A head dimension of 256 in the backward pass used to require A100/A800 or H100/H800 hardware; the README notes this now works on consumer GPUs as of flash-attn 2.5.5, provided there is no dropout. That qualifier matters. If you use dropout with head dim 256 on consumer hardware, the constraint still applies.

Alternatives and how their approach differs

The most direct alternative on the CUDA side is PyTorch's own scaled_dot_product_attention, which dispatches to fused attention kernels built into PyTorch itself. The difference in approach is packaging, not mathematics: PyTorch's version ships with the framework, so there is no separate compile step, no ninja dependency and no CUDA toolkit matching beyond what PyTorch already requires. The trade-off is that you get whatever kernel PyTorch bundles for your hardware and dtype rather than the FlashAttention-specific tiling and work partitioning described in the FlashAttention-2 paper. If your install keeps failing on toolkit mismatches, that is the escape hatch worth measuring against.

On the AMD side, the repository already contains its own alternative. The composable_kernel backend and the Triton backend implement the same FlashAttention-2 algorithm with different kernel-generation strategies: CK is the default and covers the listed MI and RDNA parts with head dimensions up to 256, while Triton covers a wider dtype range including fp32 and a longer feature list, at the cost of the sliding-window gap and the aiter submodule dependency. Choosing between them is a real decision, not a formality, and the README does not declare one universally better.

For anyone who wants the algorithm rather than the package, the FlashAttention-2 paper is linked from the README, and the Triton backend's kernels are readable Python rather than CUDA C++. Reading the Triton implementation is a more practical route to understanding the tiling than reading csrc/.

Licence, maintenance and what an upgrade costs

The repository is BSD-3-Clause. The README states that FlashAttention and FlashAttention-2 are free to use and modify under the LICENSE file, and asks users to cite and credit the project. BSD-3-Clause is permissive and does not carry the copyleft obligations of a GPL-style licence, but the citation request is a norm rather than a legal term, and whether your organization treats it as an obligation is a policy question rather than a licensing one. Nothing here is legal advice; read LICENSE and AUTHORS directly before you ship.

The maintenance picture is mixed in a way worth stating plainly. The repository is not archived, and the last push was on 2026-08-26, which is recent. But the most recent releases are all on the fa4-v4.0.0.beta line, with beta28 dated 2026-08-26. The v3 path is likewise described in the README as a beta release for testing and benchmarking before integration with the rest of the repository. So the actively changing surface is the beta surface, and the stable surface is the v2 package.

Upgrade cost follows from that. Moving from v2 to v3 changes the import name and the install directory. Moving to v4 changes the import path to flash_attn.cute and pulls a CuTeDSL implementation that targets a different set of GPUs. On ROCm, a Triton-backend upgrade can also move the pinned aiter submodule, which the README documents as a separate checkout step. Budget for re-running your own correctness checks after any of these moves, because the README's own test commands differ per backend and the full ROCm suite takes hours.

Editorial conclusion

Adopt FlashAttention if you train or serve transformer models on Ampere, Ada, Hopper or Blackwell NVIDIA GPUs, or on MI200/MI300/RDNA3/4 AMD GPUs under ROCm 6.0 or newer, and you can pin a CUDA or ROCm version that matches your PyTorch build. Do not adopt it if you are on Turing hardware and need the full feature set, on a platform where you cannot compile CUDA extensions, or on a CPU-only deployment. Before committing, verify three things: that your GPU is in the supported list for the specific package you install, that your CUDA toolkit is 12.0 or above for the v2 build and 12.3 or above for the v3 beta, and that your PyTorch is 2.2 or newer. The fa4-v4.0.0.beta28 tag is still a beta line, so treat the v4 CuTeDSL path as something to benchmark before it replaces a v2 or v3 build in production.

Frequently asked questions

Is FlashAttention faster?

The repository presents FlashAttention as a fast and memory-efficient exact attention implementation, and the papers linked from the README describe the IO-aware approach behind it. The README does not publish benchmark numbers, so the honest answer is that the design targets speed and memory at the same time as exactness, and you should measure it on your own hardware and sequence lengths.

How do I install FlashAttention?

The README gives pip install flash-attn --no-build-isolation for the v2 package, with python setup.py install as the source-build alternative. You need a CUDA or ROCm toolkit, PyTorch 2.2 or above, and the packaging, psutil and ninja Python packages.

How long does FlashAttention take to build?

The README states that compiling takes 3 to 5 minutes on a 64-core machine with the CUDA toolkit when ninja is installed and working. Without ninja it can take around two hours, because the build does not use multiple CPU cores.

What is the latest FlashAttention?

The most recent release tag listed in the repository is fa4-v4.0.0.beta28, dated 2026-08-26. FlashAttention-4 is the CuTeDSL implementation for Hopper and Blackwell GPUs, and the README describes both v3 and v4 as beta lines.

How do I use FlashAttention in PyTorch?

After installing, the v2 interface lives at src/flash_attention_interface.py. For the v4 package the README shows importing flash_attn_func from flash_attn.cute and calling it with q, k, v and a causal flag.

How do I install FlashAttention 2?

Install the flash-attn package with pip install flash-attn --no-build-isolation, or compile from source with python setup.py install. CUDA 12.0 or above is required on the NVIDIA path, and the README recommends the NVIDIA PyTorch container as a known-good toolchain.

Official sources

  1. Official README
  2. Project repository
  3. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/dao-ailab-flash-attention.svg)](https://hysenlabs.com/projects/dao-ailab-flash-attention)