# autoregressive-diffusion-pytorch: a PyTorch implementation of autoregressive image generation without vector quantization

> lucidrains' package wraps the MAR architecture from Autoregressive Image Generation without Vector Quantization into a small set of PyTorch modules, with a flow-matching variant and a bundled image trainer. It is a research implementation, and the README is thin on sampling details, checkpointing and evaluation.

**lucidrains/autoregressive-diffusion-pytorch** — Implementation of Autoregressive Diffusion in Pytorch

- Repository: https://github.com/lucidrains/autoregressive-diffusion-pytorch
- Stars: 441 · Forks: 13
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/lucidrains-autoregressive-diffusion-pytorch

## The gap this fills: MAR-style generation without a codebook

Autoregressive image models usually need a discrete vocabulary. A VQ-VAE or similar tokenizer turns an image into integer codes, and the transformer predicts the next code. That works, but the tokenizer becomes a second model to train, and quantization error puts a ceiling on what the generator can produce. The paper this repository implements, Autoregressive Image Generation without Vector Quantization, removes that step: generation stays autoregressive over a sequence of positions, but each position is produced by a small diffusion process rather than a softmax over a codebook.

The README points at the official repository for the paper and at transfusion-pytorch as an alternative route, which tells you where this package sits. It is not the reference implementation. It is a reimplementation in the style lucidrains uses across his repositories: one class per idea, keyword arguments for dimensions and depth, and a training loop you can call as a function. The audience is people who read the paper, want to try the architecture on their own data, and would rather import a module than reconstruct the training objective from the appendix.

## How the model is wired: a transformer over positions, a diffusion loss per position

The README shows two entry points. AutoregressiveDiffusion takes dim_input, dim, max_seq_len, depth, mlp_depth and mlp_width, and it expects a tensor of shape (batch, sequence, dim_input). Calling the model on that tensor returns a loss, and sample(batch_size) returns a tensor of the same shape. So the object owns both the training objective and the sampler, and the sequence dimension is fixed at construction time by max_seq_len.

ImageAutoregressiveDiffusion sits on top. It takes a model config dict with dim, depth and heads, plus image_size and patch_size. It accepts a normal image tensor of shape (batch, channels, height, width), returns a loss when called, and sample() returns images in the same shape. The patch size determines how many positions the image is cut into, which is what max_seq_len would be in the flat version. The README does not describe the internal transformer beyond naming x-transformers as a dependency, and it does not document the noise schedule or the number of diffusion steps used at sampling time.

ImageAutoregressiveFlow and AutoregressiveFlow swap the diffusion objective for flow matching; the README says the rest of the usage is the same. The flow variant adds xm_candidates, described in the README as explorative modeling, which trains against the best of k candidates per timestep. That is the one place where the two variants differ in their arguments, and it is worth reading the cited paper on explorative modeling before setting the value.

## Installing autoregressive-diffusion-pytorch and running a first forward pass

The README gives a single install line. The package requires Python 3.8 or later and pulls in torch, x-transformers, einops, einx, ema-pytorch, torchdiffeq, tqdm and accelerate, with accelerate pinned at 1.6.0 or below. That pin is the first thing to check against your existing environment, because a project already on a newer accelerate will conflict.

```bash
pip install autoregressive-diffusion-pytorch
```

The smallest useful check is the flat sequence model. It builds a model, runs a forward pass to get a scalar loss, backpropagates, and then samples. The shapes in the assertions are the contract: whatever you feed in, sample() returns the same shape.

```python
import torch
from autoregressive_diffusion_pytorch import AutoregressiveDiffusion

model = AutoregressiveDiffusion(
    dim_input = 512,
    dim = 1024,
    max_seq_len = 32,
    depth = 8,
    mlp_depth = 3,
    mlp_width = 1024
)

seq = torch.randn(3, 32, 512)
loss = model(seq)
loss.backward()

sampled = model.sample(batch_size = 3)
assert sampled.shape == seq.shape
```

For images, ImageAutoregressiveDiffusion takes a config for the internal transformer plus image_size and patch_size. A 64 by 64 image with patch_size 8 gives 64 positions, so the sequence length is derived rather than passed.

```python
import torch
from autoregressive_diffusion_pytorch import ImageAutoregressiveDiffusion

model = ImageAutoregressiveDiffusion(
    model = dict(dim = 1024, depth = 12, heads = 12),
    image_size = 64,
    patch_size = 8
)

images = torch.randn(3, 3, 64, 64)
loss = model(images)
loss.backward()

sampled = model.sample(batch_size = 3)
assert sampled.shape == images.shape
```

The first real training run uses ImageDataset and ImageTrainer. The dataset takes a directory path and an image_size, and the trainer takes the model and dataset and is invoked by calling it. Nothing else appears in the README: no batch size, no learning rate, no number of steps, no checkpoint directory. The repository does contain a train_ar_flow_oxford.py script at the top level, which is the place to look for the arguments the README omits.

```python
from autoregressive_diffusion_pytorch import (
    ImageDataset,
    ImageAutoregressiveDiffusion,
    ImageTrainer
)

dataset = ImageDataset('/path/to/your/images', image_size = 128)
model = ImageAutoregressiveDiffusion(
    model = dict(dim = 512),
    image_size = 128,
    patch_size = 16
)
trainer = ImageTrainer(model = model, dataset = dataset)
trainer()
```

## Where the README goes quiet: sampling cost, checkpoints and evaluation

The README shows one result image, labelled as oxford flowers at 96k steps, and stops there. It does not report FID, IS or any other metric, and it does not describe how many diffusion steps sample() uses or how that number affects wall-clock time. For an autoregressive model with a diffusion process at every position, sampling cost is the question a practitioner asks first, and the documentation does not answer it. You will have to read the sampler in autoregressive_diffusion_pytorch/ to find out.

Checkpointing is also undocumented. ImageTrainer is called with no arguments beyond model and dataset in the README, and nothing describes where weights are written, whether training resumes, or what format is used. The repository does depend on ema-pytorch, which suggests an exponential moving average of the weights exists somewhere in the training path, but the README never mentions it. If you plan a run longer than a few hours, treat checkpoint handling as work you will do yourself.

There is a version mismatch worth noting. pyproject.toml in the repository declares version 0.4.0, while the most recent release listed on the repository page is 0.3.0. The README install line does not pin a version. Pin it explicitly in your own requirements if you need reproducibility across machines.

## When this is the wrong tool, and what to use instead

This is the wrong tool if you want a maintained inference library. There is no documented CLI, no serving path, no ONNX or TorchScript export, and no API stability statement beyond the Beta classifier in pyproject.toml. A diffusion model with a standard UNet and a scheduler from a library such as diffusers will get you to a working image generator with far less reading, because that ecosystem documents its samplers and its checkpoints. The trade-off is that you inherit the vector-quantization or continuous-latent pipeline the library assumes, which is exactly the thing this package removes.

The closer alternative is the official MAR repository linked from the README. Both implement the same paper. The difference is in scope and packaging: the official repository is a research codebase tied to the paper's experiments and its own training scripts, while this package is a pip-installable module with a generic ImageTrainer and a flow-matching variant that the paper does not cover. If your goal is to reproduce the paper's numbers, use the official code. If your goal is to drop the architecture into your own PyTorch training loop with your own data loader, this package is the shorter path. The README also points at transfusion-pytorch as an alternative route, which is a different answer to the same problem of combining autoregressive and diffusion objectives.

## Maintenance, licence and the cost of upgrading

The repository is not archived and the last push was on 2026-09-02, so the code is recent. The release history is uneven: 0.2.7 in September 2024, 0.2.8 in November 2024, then 0.3.0 in December 2025, with pyproject.toml already at 0.4.0. That pattern means releases cluster around periods of work rather than following a cadence, and you should expect the API to move with the papers it tracks. The flow-matching variant and the xm_candidates argument are evidence of that: new research gets folded into the same classes, and constructor signatures are the place where breakage will show up.

The licence is MIT, declared in both the LICENSE file and the pyproject.toml classifier. In practice that means you can use, modify and redistribute the code, including in closed products, provided the copyright notice and permission notice are kept. The dependency list is separate: torch, x-transformers, einops, einx, ema-pytorch, torchdiffeq and accelerate each carry their own licences, and accelerate is capped at 1.6.0 or below, which is a constraint you inherit. None of this is legal advice; check the terms yourself if the distinction matters to your organisation.

## Conclusion

Adopt it if you are reproducing or extending the MAR line of work in PyTorch and want the architecture without writing the transformer, the diffusion loss and the trainer yourself. Do not adopt it if you need a supported library with documented checkpointing, evaluation metrics or a stable API for production inference; the README does not document any of those, and the package is classified as Beta. Verify first that the installed version matches what you intend to run: pyproject.toml declares 0.4.0 while the most recent release listed for the repository is 0.3.0, so check which version pip resolves before you build on it.

## FAQ

### What is autoregressive diffusion in autoregressive-diffusion-pytorch?

It is the architecture from the paper Autoregressive Image Generation without Vector Quantization, where generation proceeds autoregressively over a sequence of positions but each position is produced by a diffusion process instead of a softmax over a discrete codebook. The package exposes it as AutoregressiveDiffusion for flat sequences and ImageAutoregressiveDiffusion for images treated as a sequence of patches.

### How do I install autoregressive-diffusion-pytorch?

The README gives a single command, pip install autoregressive-diffusion-pytorch. The package requires Python 3.8 or later and depends on torch, x-transformers, einops, einx, ema-pytorch, torchdiffeq, tqdm and accelerate, with accelerate pinned at 1.6.0 or below.

### Can I train autoregressive-diffusion-pytorch on my own images?

Yes. The README shows ImageDataset pointed at a directory of images with an image_size, and ImageTrainer wrapping that dataset and the model, invoked by calling the trainer. The README does not document batch size, learning rate, step count or where checkpoints are written, so those come from the train_ar_flow_oxford.py script in the repository or from your own loop.

### What is the difference between AutoregressiveDiffusion and ImageAutoregressiveFlow?

The README says the flow-matching variant is an improvised version, imported as ImageAutoregressiveFlow or AutoregressiveFlow, and that the rest of the usage is the same. The flow variant adds an xm_candidates argument, described in the README as explorative modeling, which trains against the best of k candidates per timestep.

### Does autoregressive-diffusion-pytorch report image quality metrics?

The README shows one sample image labelled as oxford flowers at 96k steps and reports no FID, IS or other metric. It also does not describe how many diffusion steps sample() uses, so sampling cost and quality both have to be measured by reading the sampler code and running it yourself.

## Sources

- [Issues](https://github.com/lucidrains/autoregressive-diffusion-pytorch/issues)
- [License: MIT](https://github.com/lucidrains/autoregressive-diffusion-pytorch/blob/main/LICENSE)
- [lucidrains/autoregressive-diffusion-pytorch on GitHub](https://github.com/lucidrains/autoregressive-diffusion-pytorch)
- [README](https://github.com/lucidrains/autoregressive-diffusion-pytorch/blob/main/README.md)
- [Releases](https://github.com/lucidrains/autoregressive-diffusion-pytorch/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/lucidrains-autoregressive-diffusion-pytorch
