xFormers: a review of the memory-efficient attention library for PyTorch
Hackable and optimized Transformers building blocks, supporting a composable construction.
At a glance
- What is it?
- xFormers is a set of composable Transformer building blocks from Meta research, best known for memory-efficient exact attention. It installs from PyTorch's own wheel index and is aimed at GPU researchers, not at CPU inference.
- Who is it for?
- Adopt xFormers if you train or run Transformer models on a supported NVIDIA or ROCm GPU and want exact attention without writing CUDA. Skip it if you are on CPU, on an older PyTorch, or need a documented rollback path, because the README does not describe one.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 6 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
What xFormers solves, and who it is actually for
The README describes xFormers as a toolbox that is "research first", containing components "not yet available in mainstream libraries like PyTorch". That framing matters. This is not a framework you build an application on top of. It is a bag of pieces: attention operators, fused softmax, fused linear, fused layer norm, fused dropout, fused SwiGLU, and the sparse and block-sparse attention variants. The components are described as domain-agnostic, and the README says the library is used by researchers in vision and NLP.
The concrete problem is memory and speed in the attention step. The README lists memory-efficient exact attention as "up to 10x faster" and stresses one word twice: the operation is exact attention, not an approximation. That distinction is the whole point for anyone who has avoided approximate attention schemes because they change model behaviour. If you are fine-tuning or pretraining a Transformer on a GPU and the attention tensor is what limits your batch size, xFormers is aimed at you. If you are serving a small model on a laptop CPU, nothing here is addressed to you.
How the attention kernel and the dispatch layer fit together
The public entry point named in the README is xformers.ops.memory_efficient_attention. The README states that the library contains its own CUDA kernels but "dispatches to other libraries when relevant", and the v0.0.35 release note is titled "Rely on upstream FA3". Read together, that tells you the attention path is not a single hand-written kernel. There is a dispatcher that picks an implementation, and at least one of those implementations is now upstream FlashAttention 3 rather than something xFormers maintains itself.
That design has a direct consequence for anyone pinning versions. The numerical behaviour and the supported GPU list of memory_efficient_attention can shift with the upstream dependency, not only with xFormers' own release. The repository also carries a third_party directory and a .gitmodules file, and the credits section lists Sputnik, GE-SpMM, Triton, CUTLASS and Flash-Attention among the projects used "either in close to original form or as an inspiration". So the code you install is a composition, and the composition is the product.
Installing xFormers with pip against the right CUDA wheel index
The README marks the pip path as recommended for Linux and Windows, and states it requires PyTorch 2.10.0. The wheels are not on the default PyPI index alone: the commands point pip at download.pytorch.org with a CUDA-specific index URL. Pick the line matching your CUDA version.
# [linux & win] cuda 12.6 version
pip3 install -U xformers --index-url https://download.pytorch.org/whl/cu126
# [linux & win] cuda 12.8 version
pip3 install -U xformers --index-url https://download.pytorch.org/whl/cu128
# [linux & win] cuda 13.0 version
pip3 install -U xformers --index-url https://download.pytorch.org/whl/cu130For AMD hardware the README offers one extra line, marked experimental and Linux only: rocm7.1 at https://download.pytorch.org/whl/rocm7.1. After installing, the README gives a single verification command, and says it reports what kernels are built and available.
python -m xformers.infoRun that before you change any model code. If the kernels you intend to call are not listed, the import will succeed and the operator will not do what you expect. The README also documents a development channel with pip install --pre -U xformers, and a source build for other PyTorch versions, which it warns "can take dozens of minutes".
Building from source is where the install story gets thin
The README's source instructions are short and the troubleshooting list underneath them is where the real constraints live. The build needs pytorch installed first, and the README suggests pip install ninja to make it faster. It recommends setting TORCH_CUDA_ARCH_LIST, giving "6.0;6.1;6.2;7.0;7.2;7.5;8.0;8.6" as a suggested value that is "slow to build but comprehensive". If the build runs out of memory, the documented lever is MAX_JOBS, with MAX_JOBS=2 as the example.
On Windows there is a specific failure the README names: the error "Filename longer than 260 characters", fixed by enabling long paths at the OS level and running git config --global core.longpaths true. Note what this list implies. Every item is something you discover after a failed build, which means the happy path is the wheel and the source path is for people who have a reason to leave it. The README does not document rollback, so if a source build produces a broken extension you are recovering with pip tooling, not with project instructions.
Where xFormers is the wrong tool
The requirement line is torch >= 2.10, and the README repeats PyTorch 2.10.0 as the requirement for stable wheels. If your environment is pinned to an older PyTorch, the stable wheel is closed to you and you are on the source build, with the build-time and CUDA-toolkit constraints that come with it. That is a real gate, not a footnote.
Second, the ROCm path is labelled experimental and Linux only in the README. Anyone on AMD hardware should treat that word as a statement about support expectations rather than a marketing qualifier. Third, the library is GPU-oriented by construction: the value proposition is CUDA kernels and dispatch to other GPU libraries, and the README's benchmark note describes an A100 on f16. Nothing in the README describes a CPU attention path, so if CPU is your target, the library's headline feature is not available to you. Finally, if you need a stable, slowly changing operator surface, note that the v0.0.35 release note is "Rely on upstream FA3". Depending on an upstream attention implementation means your operator's behaviour is partly governed by a project you do not pin through xFormers.
xFormers versus FlashAttention
The comparison is unusual because the two are not fully separate. The README credits Flash-Attention among the repositories used in xFormers "either in close to original form or as an inspiration", and the v0.0.35 release note states that xFormers relies on upstream FA3. So the difference is not kernel versus kernel. It is scope.
FlashAttention is an attention implementation. xFormers is a collection of building blocks, and memory_efficient_attention is one entry in it. The README lists sparse attention, block-sparse attention, fused softmax, fused linear, fused layer norm, fused dropout and fused SwiGLU alongside it. If you want a single attention operator and nothing else, going straight to the upstream implementation removes one layer of version coupling. If you want the other fused pieces, or the sparse variants, xFormers is the package that carries them, and it will route the dense attention case to the upstream kernel anyway.
Licence, releases and what upgrading costs you
The README states that xFormers has a BSD-style license as found in the LICENSE file, and that it includes code from the triton-lang/kernels repository. The repository metadata reports the licence as NOASSERTION, which means the automated classifier did not resolve a standard identifier. That is the kind of mismatch worth reading the LICENSE file over before you ship, and it is not something to resolve from a README sentence. This is a description of what the files say, not legal advice.
The release cadence visible in the changelog is tied to PyTorch and CUDA: v0.0.34 is subtitled "Stable wheels for PyTorch 2.10+", v0.0.33.post2 is "Wheels for PyTorch 2.9.1", and v0.0.35 is the FA3 change. Upgrades are therefore not free of the surrounding stack. Moving xFormers usually means moving PyTorch or CUDA with it, and the wheel index URL in your install command changes with the CUDA version. Budget for that as a coordinated upgrade rather than a patch bump. The last push to the repository was on 2026-09-15.
Editorial conclusion
Adopt xFormers if you train or run Transformer models on a supported NVIDIA or ROCm GPU and want exact attention without writing CUDA. Skip it if you are on CPU, on an older PyTorch, or need a documented rollback path, because the README does not describe one. Before installing, run python -m xformers.info after a stable wheel install and confirm the kernels you need are listed.
Frequently asked questions
How do I install xFormers?
The README recommends pip on Linux and Windows, pointing the index URL at download.pytorch.org with a CUDA-specific path such as cu126, cu128 or cu130, and states that PyTorch 2.10.0 is required. There is also a development channel using pip install --pre -U xformers.
How do I check which xFormers kernels are available after installing?
The README gives one command for this: python -m xformers.info. It reports information about the installation and which kernels are built and available, so it is the first thing to run before changing model code.
What is xFormers?
The README describes it as a toolbox of customizable, domain-agnostic Transformer building blocks, including memory-efficient exact attention, sparse attention, fused softmax, fused linear and fused layer norm. It is built for research and contains components the README says are not yet available in mainstream libraries like PyTorch.
How do I use xFormers?
The README points to xformers.ops.memory_efficient_attention as the main operator and links a Colab notebook that walks through a minGPT example. The components are meant to be used individually rather than through a single top-level API.
What is the xFormers library used for?
The README says the building blocks are domain-agnostic and that xFormers is used by researchers in vision and NLP. The stated goal is fast, memory-efficient Transformer components, including exact attention and several fused operators.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/facebookresearch-xformers)
Community notes