Self-hosted service
mjun0812/flash-attention-prebuild-wheels avatar
mjun0812/flash-attention-prebuild-wheels

flash-attention-prebuild-wheels: prebuilt flash-attn wheels for Linux and Windows

Project brief: Provide with pre-build flash-attention 2 and 3 package wheels on Linux and Windows using GitHub Actions.

1,752 stars82 forksPythonBSD-3-Clause

At a glance

What is it?
mjun0812/flash-attention-prebuild-wheels publishes prebuilt flash-attn 2 and 3 wheels so you can skip a long source build, but the wheel you need depends on an exact Python, CUDA and PyTorch combination. Here is how the naming works, what to install, and where it breaks.
Who is it for?
Adopt it when your Python, CUDA, PyTorch and flash_attn combination already appears on the packages page or the search page, and when you are on Linux x86_64, Linux ARM64 or Windows x86_64.
Can I use it commercially?
Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 4 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What problem prebuilt flash-attn wheels solve

Building flash-attention from source is slow and heavy. The README says so directly: building it "takes a very long time and is resource-intensive." That matters most on Windows, where the CUDA toolchain setup is awkward, and in CI pipelines where a source build can exceed the job time limit. The README acknowledges that second case under Self-Hosted Runner Build, noting that some version combinations cannot be built on GitHub-hosted runners because of job time limitations.

The project's answer is to build the wheels once and publish them. It goes further than the upstream project by also building combinations of CUDA and PyTorch that are not officially distributed. The audience is therefore narrow but real: people who need flash_attn on a specific torch and CUDA pair, and who would rather download a file than run a compiler. It is not a library, not a runtime, and not a fork of flash-attention. It is a distribution channel for binary artifacts, plus the CI configuration used to produce them.

How the wheel naming scheme encodes your whole environment

Everything in this project hangs off one filename convention. The README gives the template:

bash
flash_attn-[flash_attn Version]+cu[CUDA Version]torch[PyTorch Version]-cp[Python Version]-cp[Python Version]-linux_x86_64.whl

The local version label is the part after the plus sign. Since v0.5.0 the wheels carry a label indicating the CUDA and PyTorch versions, so `pip list` shows `flash_attn==2.8.3+cu130torch2.9` instead of a bare `flash_attn==2.8.3`. That label is the mechanism that lets several builds of the same flash_attn release coexist in one index. It also means the filename is a compatibility contract: if your installed torch is not the one in the label, you are installing a wheel built against a different ABI.

The README's own example is Python 3.11, CUDA 12.4, PyTorch 2.5 and flash_attn 2.6.3, written as `flash_attn-2.6.3+cu124torch2.5-cp312-cp312-linux_x86_64.whl`. Note the mismatch between the prose and the filename: the text says Python 3.11, the tag says cp312. Treat the tag as authoritative and read it carefully, because that kind of slip is exactly the sort of thing that costs you a failed install.

Since v0.7.0 the Linux x86_64 wheels are built with the manylinux2_28 platform, which the README says keeps them compatible with old glibc versions (<=2.17). Since v0.8.0, Flash Attention 3 wheels are also published, under the `flash_attn_3` name, and they require Hopper (SM90) or newer GPUs and CUDA 12.3 or newer. That is a hardware floor, not a packaging detail: on an A100 or anything older, the FA3 wheel is the wrong artifact regardless of whether it installs.

Installing the right wheel and checking it landed

The README's install flow has three steps: work out your version combination, find the matching wheel on the search page, the packages page or the releases page, then install it. The direct install form points pip at a release asset URL:

bash
pip install https://github.com/mjun0812/flash-attention-prebuild-wheels/releases/download/v0.0.0/flash_attn-2.6.3+cu124torch2.5-cp312-cp312-linux_x86_64.whl

The `v0.0.0` in that path is a placeholder in the README, not a real release tag. Substitute the tag of the release you actually want, such as one of the recent ones, and keep the filename exactly as published. If you prefer to keep the artifact, download it first and install from disk:

bash
wget https://github.com/mjun0812/flash-attention-prebuild-wheels/releases/download/v0.0.0/flash_attn-2.6.3+cu124torch2.5-cp312-cp312-linux_x86_64.whl
pip install ./flash_attn-2.6.3+cu124torch2.5-cp312-cp312-linux_x86_64.whl

After installing, verify two things. First, that the local version label survived: run `pip list` and look for the `+cu...torch...` suffix, because that is how the README says you distinguish these builds from a plain build. Second, that the wheel matches the interpreter that will run your code, since a cp312 wheel will not load into a cp311 environment. The README does not document an uninstall or rollback procedure; if you need to revert, you are relying on ordinary pip behaviour.

Coverage numbers and what they do not promise

The README carries a coverage table: Linux x86_64 at 521 existing and 0 missing, Linux ARM64 at 174 and 0, Windows at 159 and 0, for 854 total and 100.0% coverage. Read that as coverage of the project's own build matrix, not of every combination a user might want. The matrix is defined in `create_matrix.py`, and the README's self-build section is explicit that "depending on the combination of versions, it may not be possible to build." A 100% figure against a self-defined matrix can coexist with a combination you need being absent entirely.

The build environments table shows how the artifacts get made: Linux x86_64 on GitHub-hosted `ubuntu-22.04` or a self-hosted runner with `ubuntu:24.04` or `manylinux_2_28_x86_64`, Linux ARM64 on `ubuntu-22.04-arm`, Windows x86_64 on `windows-2022`, a self-hosted `windows11`, or AWS CodeBuild. The repository also carries `build_linux.sh`, `build_linux_rocm.sh`, `build_windows.ps1`, `get_torch_cuda_version.py` and a `patches/` directory, which is consistent with a project that maintains its own build recipes rather than only mirroring upstream output. The ROCm script is present in the tree, but the README's package and coverage sections describe CUDA and PyTorch combinations, so do not assume a ROCm wheel is published just because a build script exists.

Where this is the wrong tool

The clearest failure case is a version combination that was never built. There is no fallback inside the project other than building it yourself, and the self-build path means forking the repository, optionally setting up a self-hosted runner, editing `create_matrix.py`, and pushing a `v*.*.*` tag to trigger the workflow. That is a real amount of work, and if you are going to run that pipeline anyway, you are maintaining a fork rather than consuming a wheel.

The second case is hardware. Flash Attention 3 wheels need Hopper (SM90) or newer and CUDA 12.3+. On older GPUs the FA3 artifact is unusable, and you should be looking at the FA2 wheels instead. The third case is glibc: the manylinux2_28 build is described as compatible with old glibc (<=2.17), which is a floor to check against your target image, not a blanket guarantee for any distribution.

The fourth case is governance. The README states that the project uses a self-hosted runner and AWS CodeBuild, and asks readers to sponsor to help maintain that infrastructure, with a note of thanks to a contributor who provided computing resources. That is honest about the cost model, and it also tells you the supply chain depends on donated and personally funded machines. If your organisation requires a vendor with a support contract, this is not that. The last push to the repository was on 2026-07-26, and the most recent release, v0.9.52, carries the same date.

How it differs from building flash-attention yourself

The alternative is the upstream Dao-AILab/flash-attention repository, which this project links to as the original. Upstream gives you source and its own build instructions; you compile against your local CUDA toolkit and torch, and the resulting binary is guaranteed to match the environment it was built in because you built it there. That is the real difference in approach: upstream trades time for exactness, this project trades exactness for time.

The trade is not free. A prebuilt wheel is compiled against one specific torch and CUDA pair, which is why the local version label exists and why the filename is so long. If your torch is a nightly, a custom build, or a version the matrix does not cover, the prebuilt wheel is either unavailable or a mismatch. Upstream also gives you the option of building with flags this project may not use, and the repository does carry a `patches/` directory, which suggests the build recipes are not always a straight compile of upstream source. If reproducibility against pristine upstream source matters to you, read that directory before trusting a wheel.

A second alternative, for the common case, is simply installing upstream's own published wheels where they exist for your combination. This project's stated reason to exist is the combinations upstream does not distribute, so if you are on a mainstream torch and CUDA pair, check upstream first.

Maintenance cost, licensing and what to verify

The upgrade cost is dominated by version matching, not by the package itself. Every time you bump torch or CUDA you need a new wheel, because the local version label binds the artifact to that pair. That makes torch upgrades a two-step operation: upgrade torch, then find and install the matching flash_attn wheel, then confirm the label in `pip list`. If no wheel exists for the new pair, you are back to the fork-and-build path. This is the ongoing tax of the approach, and it is worth writing down before you depend on it.

The repository is BSD-3-Clause, and the README includes a BibTeX entry asking that you cite the repository if you use it in research. The wheels themselves package flash-attention, which is a separate project with its own licensing and citation requirements; the README reproduces the two flash-attention BibTeX entries for the 2022 NeurIPS and 2024 ICLR papers. Redistributing and citing are questions for your own legal and research-compliance review, not something to settle from a README.

Before adopting, verify three concrete things: that your exact Python, CUDA and torch combination appears on the search page or in `doc/packages.md`; that the platform matches one of Linux x86_64, Linux ARM64 or Windows x86_64; and that your GPU generation is appropriate for the FA2 or FA3 wheel you picked. The repository also carries a `tests/` directory and a `docs/` directory alongside `doc/`, which is a slightly odd duplication worth a look if you plan to rely on the published documentation.

Editorial conclusion

Adopt it when your Python, CUDA, PyTorch and flash_attn combination already appears on the packages page or the search page, and when you are on Linux x86_64, Linux ARM64 or Windows x86_64. Do not adopt it if you need a combination that is not listed, if you are targeting an older glibc than the manylinux2_28 floor allows, or if you want a project with a formal support commitment; this is one maintainer's build infrastructure, funded partly by sponsors, and the README does not describe a support policy. Before installing, run pip list to record your exact torch and CUDA versions, open the search page or doc/packages.md to confirm a matching filename exists, and check that the local version label in that filename matches what pip reports. If nothing matches, the documented fallback is to fork the repository, edit create_matrix.py, and build the wheel yourself on GitHub Actions or a self-hosted runner.

Frequently asked questions

How do I install flash-attention from mjun0812/flash-attention-prebuild-wheels?

Pick the wheel whose filename matches your flash_attn, CUDA, PyTorch and Python versions, then install it either directly from the release URL with pip or by downloading the file with wget and running pip install on the local path. The README gives both forms.

Why does my project fail to build flash_attn and ask for a wheel-building requirement?

The README states that building flash-attention takes a very long time and is resource-intensive, which is the reason this repository publishes prebuilt wheels in the first place. Installing a matching prebuilt wheel avoids the source build entirely.

Why can't I install flash-attention from these prebuilt wheels?

The most likely cause is that no wheel exists for your exact combination of Python, CUDA, PyTorch and flash_attn, since the README notes that some version combinations may not be buildable at all. Check the search page or doc/packages.md for a filename matching your environment before assuming the install should work.

How do I check whether flash attention is installed correctly?

Run pip list and look for the local version label, for example flash_attn==2.8.3+cu130torch2.9. Since v0.5.0 the README says wheels carry that label indicating the CUDA and PyTorch versions, so its presence tells you which build you have. The README does not document a separate verification command.

Official sources

  1. Official README
  2. Project repository
  3. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/mjun0812-flash-attention-prebuild-wheels.svg)](https://hysenlabs.com/projects/mjun0812-flash-attention-prebuild-wheels)