perceiver-pytorch: Perceiver and Perceiver IO in PyTorch
Implementation of Perceiver, General Perception with Iterative Attention, in Pytorch
At a glance
- What is it?
- lucidrains/perceiver-pytorch packages the Perceiver and Perceiver IO architectures as PyTorch modules, including a PerceiverLM variant and an experimental bottom-up attention path. It is a research-oriented implementation, not a trained model or a training framework.
- Who is it for?
- Adopt perceiver-pytorch when you want to modify the Perceiver architecture itself in PyTorch and can supply your own training pipeline and data; it gives you Perceiver, PerceiverIO, PerceiverLM and an experimental bottom-up variant as ordinary nn.Module classes. Do not adopt it expecting pretrained weights, a fine-tuning recipe or a maintained inference server, because the README documents none of those.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 8 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What perceiver-pytorch actually provides
The package is a set of nn.Module implementations, not a model zoo. From perceiver_pytorch you import Perceiver, PerceiverIO or PerceiverLM, each of which builds an untrained network when constructed. There are no pretrained checkpoints in the repository, no download helper, and no inference server. The README's usage example ends with model(img) returning a tensor of shape (1, 1000), which is a randomly initialized forward pass, not a prediction you can act on.
The audience is therefore narrow: researchers and engineers who want to experiment with the Perceiver family of architectures, change the latent array, swap the Fourier encoding, or attach the encoder to their own decoder head. If you want a working classifier, this repository gives you the skeleton and leaves the training to you. That is a deliberate scope choice by the author, and it is consistent with the rest of the lucidrains collection.
How cross attention and the latent array fit together
Perceiver's core idea, as implemented here, is to stop attending over the raw input. Instead the model keeps a small set of latents (num_latents, default 256 in the README example) and repeatedly pulls information out of the input through cross attention into that latent array. The input can be an image, a video, audio or a point cloud; the input_axis argument tells the model how many spatial axes to Fourier-encode, and fourier_encode_data controls whether the module does that encoding for you.
The depth argument is the number of cross attention plus self attention blocks. The README states the final attention shape as depth * (cross attention -> self_per_cross_attn * self attention), so depth 6 with self_per_cross_attn 2 gives six cross attention blocks and twelve latent self attention blocks. cross_heads is 1 in the paper, while latent_heads is 8, and the per-head dimensions are set separately by cross_dim_head and latent_dim_head. weight_tie_layers optionally reuses the same weights across depth, which cuts parameter count at the cost of expressiveness.
PerceiverIO extends this with a decoder that cross-attends from a separate query tensor into the latents, which is why the output sequence length is decoupled from the input length. The README example passes seq of shape (1, 512, 32) and queries of shape (128, 32), and gets back (1, 128, 100). PerceiverLM wraps PerceiverIO for language modelling, taking token ids and a boolean mask and returning per-position logits over num_tokens.
The experimental module, imported as perceiver_pytorch.experimental.Perceiver, adds bottom-up attention in the style of the Set Transformers Induced Set Attention Block. The README calls this experimental and shows only the import change, so treat it as a research branch rather than a supported second implementation.
Installing perceiver-pytorch and running a first forward pass
The README gives a single install line. It pulls in torch and einops, and pyproject.toml sets the floors at torch>=2.5 and einops>=0.8.0 with requires-python >= 3.9.
pip install perceiver-pytorchOnce installed, the smallest useful check is to build a Perceiver and push a random image through it. The README's own example uses a 224x224 RGB image with input_axis = 2, and the call returns a tensor shaped (1, 1000) because num_classes is 1000.
import torch
from perceiver_pytorch import Perceiver
model = Perceiver(
input_channels = 3,
input_axis = 2,
num_freq_bands = 6,
max_freq = 10.,
depth = 6,
num_latents = 256,
latent_dim = 512,
num_classes = 1000
)
img = torch.randn(1, 224, 224, 3)
print(model(img).shape) # (1, 1000)If your task has a variable-length output, switch to PerceiverIO and pass decoder queries. The README's example uses a 512-token sequence of dimension 32 and 128 queries, and the returned logits have one row per query. Setting seq_dropout_prob drops a fraction of the input tokens before encoding, which the README describes as structured dropout for saving compute and regularizing effects.
import torch
from perceiver_pytorch import PerceiverIO
model = PerceiverIO(
dim = 32,
queries_dim = 32,
logits_dim = 100,
depth = 6,
num_latents = 256,
latent_dim = 512,
seq_dropout_prob = 0.2
)
seq = torch.randn(1, 512, 32)
queries = torch.randn(128, 32)
logits = model(seq, queries = queries)
print(logits.shape) # (1, 128, 100)The repository also ships train_cifar10.py at the top level, which is the closest thing to an end-to-end example. There is no documented CLI for it in the README, so read the file before running it. The test suite lives in tests/ and is configured through pyproject.toml with pytest and pythonpath set to the repository root.
pip install -e ".[test]"
pytestWhat the implementation does not give you
The most important limitation is the absence of trained weights. Every constructor in the README returns a randomly initialized module, so any accuracy number you see from this package is one you produced yourself. There is no model card, no checkpoint URL and no evaluation script in the repository.
The README also does not document rollback, version pinning or migration between releases. The published releases listed are 0.8.8 from 2023-08-22, 0.8.7 from 2023-01-03 and 0.8.6 from 2022-12-05, while pyproject.toml declares version 0.9.0, so the version string in the source tree is ahead of the most recent release shown. Anyone pinning a version should confirm what PyPI actually serves rather than assuming 0.9.0 is published.
Perceiver is also the wrong tool for small, fixed-shape inputs where a plain Transformer or a CNN is simpler. The latent bottleneck trades raw capacity for a fixed compute budget, and if your input is already short, that bottleneck only removes information. The experimental bottom-up variant is explicitly labelled experimental in the README, so it should not be the basis of a production choice.
Finally, the README documents constructor arguments but not training dynamics. Nothing in the repository says how many steps convergence takes, what learning rate to use, or how the Fourier encoding interacts with normalization. Those are open questions you will answer on your own data.
How this differs from a standard Transformer implementation
The natural alternative is a conventional Transformer encoder, such as the reference nn.TransformerEncoder in PyTorch or a library like Hugging Face Transformers. The difference is where the attention is paid. A Transformer encoder attends over every pair of input tokens, so cost grows with the square of sequence length. Perceiver attends from a fixed latent array into the input, so the cost of the cross attention stage scales with num_latents times input length rather than input length squared.
The trade-off is that all information must pass through a fixed-width bottleneck. If the latent array is too small, detail is lost before the decoder ever sees it; if it is too large, you have given back the compute savings. A standard Transformer makes no such commitment. PerceiverIO's decoder queries add a second difference: output length is chosen at call time, which a plain encoder-decoder Transformer only achieves through autoregressive decoding or a separately designed head.
If your modality is text and your goal is a working model, Hugging Face Transformers with a pretrained checkpoint is the more direct route. perceiver-pytorch is for the case where you want the architecture itself and are prepared to train it.
Maintenance status and licence
The repository is not archived. The last push was on 2026-06-08, which is recent, but the published releases stop at 0.8.8 from 2023-08-22. That gap matters: the code has moved since the last release, and pyproject.toml already declares 0.9.0. If you install from PyPI you may get older code than the main branch contains, so decide deliberately between pip install perceiver-pytorch and installing from the repository.
The licence is MIT, declared both in the LICENSE file and in the pyproject.toml classifier. MIT is permissive and places few conditions on reuse, but it says nothing about the licensing of any dataset or pretrained component you bring to the training pipeline, and nothing here addresses patents or the underlying papers. Check those separately for your own use case.
Upgrade cost is dominated by the torch>=2.5 and einops>=0.8.0 floors. Raising the torch floor can force an upgrade across the rest of your stack, which is often more disruptive than the perceiver-pytorch change itself. Because the README does not document a deprecation policy, treat any minor version bump as potentially breaking and run pytest against your own integration before upgrading.
Editorial conclusion
Adopt perceiver-pytorch when you want to modify the Perceiver architecture itself in PyTorch and can supply your own training pipeline and data; it gives you Perceiver, PerceiverIO, PerceiverLM and an experimental bottom-up variant as ordinary nn.Module classes. Do not adopt it expecting pretrained weights, a fine-tuning recipe or a maintained inference server, because the README documents none of those. Before committing, check the pyproject.toml dependency floor (torch>=2.5, einops>=0.8.0), read tests/ to see which constructor arguments are actually exercised, and run train_cifar10.py on a small subset to confirm throughput on your hardware.
Frequently asked questions
Does perceiver-pytorch ship pretrained weights?
No. The README shows only constructor calls and a forward pass on random tensors, and there is no checkpoint download or model card in the repository. You have to train the network yourself.
How do I install perceiver-pytorch?
The README gives a single command, pip install perceiver-pytorch. The package requires Python 3.9 or later, torch 2.5 or later and einops 0.8.0 or later according to pyproject.toml.
What is the difference between Perceiver and PerceiverIO in perceiver-pytorch?
Perceiver takes an input and returns a fixed number of classes, while PerceiverIO takes decoder queries and returns one output row per query, so the output sequence length is flexible. The README imports both from perceiver_pytorch.
What language is used for PyTorch?
perceiver-pytorch is written in Python and its examples are Python, with dependencies on torch and einops declared in pyproject.toml. The package metadata classifies it as Python 3.9.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/lucidrains-perceiver-pytorch)