Model or dataset
lucidrains/x-transformers avatar
lucidrains/x-transformers

lucidrains/x-transformers: a full-attention transformer library with experimental features behind flags

A concise but complete full-attention transformer with a set of promising experimental features from various papers

5,949 stars522 forksPythonMIT

At a glance

What is it?
x-transformers is a PyTorch library that packs encoder, decoder, cross-attention and vision transformer variants into a small set of composable classes. Its value is in the feature flags, not in the defaults, and the README is the only real manual.
Who is it for?
Adopt x-transformers if you are prototyping attention variants in PyTorch and want encoder, decoder, cross-attention and vision wrappers behind one API, or if you want to read a working implementation of a paper rather than a description of one.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What x-transformers is for, and who it is written for

The README describes the project as a concise but fully-featured transformer, complete with experimental features from various papers. That sentence is the whole positioning. This is not a training framework and not a model zoo. It is a set of nn.Module classes that implement attention variants, and the person it is written for is someone who already knows what they want to build and does not want to write the attention block from scratch.

The repository layout supports that reading. The top level holds a package directory, a tests directory, and a long list of single-file training scripts: train_enwik8.py, train_parity.py, train_copy.py, train_belief_state.py, train_gpt_vae.py, train_with_muon.py and others. Each of those is a self-contained example rather than a framework entry point. There is no CLI, no config directory, no experiment tracker integration. If you want orchestration, you bring it.

The intended audience is narrow in a useful way. If you are reproducing a paper's attention mechanism, or ablating one component of a transformer, this library gives you the component and gets out of the way. If you are building a product on top of a pretrained model, you are probably looking at the wrong layer of the stack.

The three wrapper classes and how data flows through them

The architecture visible in the README is a wrapper-plus-layers split. TransformerWrapper handles token embedding, positional handling and the output projection back to vocabulary size. It takes num_tokens, max_seq_len and an attn_layers argument. The attn_layers argument is where the actual stack lives: Decoder for a GPT-like causal stack, Encoder for a BERT-like bidirectional stack, or a cross-attending Decoder for sequence-to-sequence work.

XTransformer is the second shape. It bundles an encoder and a decoder in one object with parallel parameter groups: dim, enc_num_tokens, enc_depth, enc_heads, enc_max_seq_len, and the dec_ equivalents. It also exposes tie_token_emb, which the README's example sets to True to tie the encoder and decoder embeddings.

ViTransformerWrapper is the third. It takes image_size and patch_size, turns an image into patches, and runs them through an Encoder stack. The README shows it used two ways: standalone for classification, returning a tensor of shape (1, 1000) for num_classes = 1000, and with return_embeddings = True to produce vectors that feed a decoder.

The multimodal path is worth reading closely because it is the least obvious data flow in the README. In the PaLI example, the vision transformer produces embeddings, and those embeddings are passed to XTransformer as src_prepend_embeds. The README states this will prepend image embeddings to encoder text embeddings before attention. So the image tokens and the prompt tokens end up in one sequence inside the encoder, rather than being fused by a separate projection module. That is a design choice with consequences: the encoder's positional handling now spans two modalities.

Installing x-transformers and running a first decoder

The README gives one install line. The package name on PyPI matches the project name, and the dependencies declared in pyproject.toml are einx, einops, loguru, packaging, torch and torch-einops-utils, with torch pinned at >=2.0. Python 3.9 or later is required. There is an optional examples extra and a separate flash-pack-seq extra that pulls flash-attn>=2.0.

bash
pip install x-transformers

After that, the smallest useful thing to run is a decoder-only model. The README's GPT-like example builds a TransformerWrapper around a Decoder with dim = 512, depth = 12 and heads = 8, wraps it in .cuda(), and feeds a random token tensor of shape (1, 1024). The comment on the final line states the output shape is (1, 1024, 20000), matching num_tokens.

python
import torch
from x_transformers import TransformerWrapper, Decoder

model = TransformerWrapper(
    num_tokens = 20000,
    max_seq_len = 1024,
    attn_layers = Decoder(
        dim = 512,
        depth = 12,
        heads = 8
    )
).cuda()

x = torch.randint(0, 256, (1, 1024)).cuda()

model(x) # (1, 1024, 20000)

One thing to notice in that snippet: the tensor is created with torch.randint(0, 256, ...) while num_tokens is 20000. The README does not explain the mismatch, and it does not matter for a smoke test, but it means the example is checking shapes rather than semantics. Do not read it as a statement about how to prepare real data.

For an encoder, the shape is the same but the call takes a mask. The README's BERT-like example passes mask = mask alongside the token tensor and documents the same (1, 1024, 20000) output.

Flash Attention and the flags that change behaviour

The feature the README spends the most words on is Flash Attention. The explanation given is that it processes the attention matrix in tiles, tracking only a running softmax and exponentiated weighted sums, and recomputes on the backward pass in a tiled fashion so memory stays linear in sequence length. It notes that PyTorch 2.0 exposes Tri Dao's CUDA kernel through torch.nn.functional.scaled_dot_product_attention, and that a mem_efficient variant uses the same design with a different tile traversal order.

The README is explicit about the trade-off, which is unusual and worth quoting in substance: the only reason to avoid Flash Attention is if you require operating on the attention matrix, and it names dynamic positional bias, talking heads and residual attention as cases that need it. That is a real boundary. If your variant reads the attention weights, the fused kernel is not available to you, and the library cannot paper over that.

The switch is attn_flash set to True. The README's wording is that you can use it by setting attn_flash to True and enjoy the immediate memory savings and increase in speed. Note that this is the README's claim about the kernel's behaviour, not a measurement of this library against a baseline.

Dropout flags follow the same pattern. TransformerWrapper accepts emb_dropout, and the attention layers accept layer_dropout, attn_dropout and ff_dropout. The README annotates layer_dropout as stochastic depth, meaning an entire layer is dropped rather than individual activations. The example sets all four to 0.1 in one model, which is a configuration choice, not a recommendation.

Where x-transformers is the wrong tool

The package classifier in pyproject.toml says Development Status :: 4 - Beta. That is the project's own label. It does not mean the code is broken, but it does mean you should not expect the API surface to be frozen, and the README does not document a deprecation policy, a changelog for breaking changes, or a rollback path for a version that regresses your training run.

The second limitation is structural. There is no trainer in this repository. The single-file scripts at the top level show how the author trains things, and they are examples, not supported entry points. If you need checkpoint resumption, gradient accumulation, mixed precision policy, or multi-node launch, you are assembling that yourself around the model classes. A framework that ships a Trainer gives you those for free and charges you in abstraction; x-transformers makes the opposite trade.

Third, the library will accept configurations that do not fit. The README's GPT-3 example uses dim = 12288, depth = 96, heads = 96 and attn_dim_head = 128, and the README itself says you would not be able to run it. Nothing in the constructor prevents you from writing that. Memory planning is your problem, and the only lever the README points at is attn_flash.

Finally, the README is the documentation. There is a Homepage URL in pyproject.toml that resolves to the PyPI project page, and the Repository URL points at a Codeberg mirror. Beyond that, feature behaviour is described in README sections with paper links, and anything not covered there is not covered anywhere you can read before installing.

Compared with Hugging Face Transformers

The obvious alternative is Hugging Face Transformers, and the difference is not quality but layer. Hugging Face ships pretrained weights, tokenizers, a Trainer with checkpointing and distributed support, and a configuration system with versioned model cards. Its attention implementations are deliberately conservative because they have to keep thousands of checkpoints loadable.

x-transformers ships no weights and no tokenizer. Its attention blocks are the point, and they change. If you want to fine-tune an existing checkpoint, Hugging Face is the shorter path by a wide margin. If you want to test whether a tiled attention variant, a different positional scheme or a cross-attention arrangement behaves differently on your data, x-transformers lets you change that in the constructor rather than by editing a modelling file inside a large library.

A second comparison worth making is against writing the block yourself. A standard multi-head attention module is short, and many people start there. What x-transformers adds is the accumulated set of variants and the wrapper classes that handle embedding, masking and output projection consistently across encoder, decoder and vision shapes. Whether that is worth a dependency depends on how many variants you plan to try. For one, write it yourself.

Maintenance, licence and what an upgrade costs

The repository is not archived, and the last push was on 2026-09-02. Recent releases listed are 2.16.1 on 2026-02-12, 2.16.0 on 2026-02-07 and 2.15.2 on 2026-02-03. The version declared in pyproject.toml is 2.28.4, which is ahead of the newest release in that list, so the packaging metadata and the release list are not in step. If you pin by version, check PyPI rather than trusting the repository file.

The licence is MIT, declared both in the LICENSE file and in the pyproject.toml classifier. MIT is permissive: it allows commercial use and modification with attribution and without a copyleft obligation on your own code. It also comes with no warranty, which matters here because the package is labelled Beta. This is a description of the licence text, not legal advice; if your organisation has a policy on dependencies, run it through that.

Upgrade cost is the part the README does not address. Because the API is constructor keyword arguments rather than configuration files, a breaking change shows up as a TypeError or a silently ignored argument at model construction time, which is at least early. What you cannot tell from the repository is which flags are considered stable. The practical approach is to pin the version in your own dependency file and read the README section for each flag you use when you bump it.

Editorial conclusion

Adopt x-transformers if you are prototyping attention variants in PyTorch and want encoder, decoder, cross-attention and vision wrappers behind one API, or if you want to read a working implementation of a paper rather than a description of one. Do not adopt it if you need a training framework with checkpointing and distributed launch handled for you, or if you need documented stability guarantees, since the package classifier marks it Beta and the README does not document a deprecation or rollback policy. Before you commit, verify three things: that torch>=2.0 is installed, that the flag you intend to use actually appears in the README section for that feature, and that your sequence lengths and head counts fit the memory you have, because the library will not stop you from configuring a model that does not fit.

Frequently asked questions

How do I install x-transformers?

The README gives a single command, pip install x-transformers. It requires Python 3.9 or later and torch>=2.0, along with einx, einops, loguru, packaging and torch-einops-utils. An optional flash-pack-seq extra pulls flash-attn>=2.0.

What is x-transformers by lucidrains?

It is a PyTorch library described in the README as a concise but fully-featured transformer with experimental features from various papers. It provides encoder, decoder, cross-attention and vision transformer wrappers rather than pretrained weights or a training framework.

Can I use Flash Attention in x-transformers?

Yes, by setting attn_flash to True, which the README says brings memory savings and speed. The README notes the one reason to avoid it is if you need to operate on the attention matrix, such as with dynamic positional bias, talking heads or residual attention.

Official sources

  1. Issues
  2. License: MIT
  3. lucidrains/x-transformers on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/lucidrains-x-transformers.svg)](https://hysenlabs.com/projects/lucidrains-x-transformers)