# Infinity: bitwise autoregressive modeling for high-resolution image synthesis

> The ByteDance research project that replaces discrete visual tokens with bit-level prediction, ships 2B and 8B checkpoints under the MIT licence, and reports 1024x1024 generation in 0.8 seconds.

**FoundationVision/Infinity** — [CVPR 2025 Oral]Infinity ∞ : Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis

- Repository: https://github.com/FoundationVision/Infinity
- Stars: 1,588 · Forks: 93
- Language: Python
- License: MIT
- Published: 2026-10-06 · Updated: 2026-10-06 · Language: en
- Canonical page: https://hysenlabs.com/projects/foundationvision-infinity

## What problem the paper is attacking

Autoregressive image generation works by predicting a sequence of visual tokens one at a time, in the same shape as a language model predicting words. The catch has always been the vocabulary. A text model needs tens of thousands of tokens to cover a language. An image model needs a codebook to cover visual detail, and a codebook large enough to look good turns out to be enormously expensive in memory.

Infinity, published as arXiv 2412.04431 and accepted as an oral at CVPR 2025, attacks that cost directly. Instead of predicting an index into a fixed codebook, it predicts bits. The paper frames this as taking the visual vocabulary to its limit while scaling the transformer alongside it, and it reports that the combination sets a new record for autoregressive text-to-image models at the time.

The repository is a research release rather than a library. You will find training code, inference code, tokenizer code, and a model zoo of checkpoints. There is no pip package, no plugin, no API. The intended audience is someone who wants to read the implementation or fine-tune a checkpoint, not someone who wants to call a function.

The lineage is worth noting because it explains the design. The same group published Visual AutoRegressive Modeling, which received the NeurIPS 2024 Best Paper Award, and a separate BitVAE repository holds the image tokenizer training code. Infinity is the next step: apply the bitwise idea to the whole image rather than only to the token layout.

## The tokenizer: bitwise multi-scale residual quantization

The first component is a new quantizer called a bitwise multi-scale residual quantizer, and its stated benefit is memory reduction. That is what makes a very large vocabulary trainable at all: with `V_d = 2^32` or even `2^64`, a conventional codebook would be impossible to store, let alone train.

The idea of residual quantization is to refine an approximation in stages. An early stage captures coarse structure, and each later stage encodes the error left by the one before. Doing that across multiple scales means early tokens carry the large-scale layout of an image and later tokens carry progressively finer detail. That ordering is what makes next-token prediction work at all on images, because by the time the model reaches a late token, the hard global decisions are already made.

The tokenizer table in the model zoo reports reconstruction quality measured as rFID and PSNR, and the trend is the point. Moving from `V_d = 2^16` up to `2^32`, IN-256 rFID goes from 1.22 down to 0.61 and PSNR from 20.9 up to 22.7. At `2^64`, IN-256 reaches 0.33 rFID and 24.9 PSNR. On IN-512 the same progression gives rFID 0.31 down to 0.15 and PSNR 22.6 up to 26.4.

Two readings of that table. The optimistic one is that reconstruction quality simply improves with vocabulary size, which is what the scaling argument predicts. The cautious one is that this is a compressor benchmark, and rFID on a reconstruction task measures something narrower than whether generated images look good. Both can be true, and the honest summary is that the tokenizer works better as the codebook grows without the memory cost that would normally make that impossible.

## The classifier: why bitwise heads matter

The second component is the infinite-vocabulary classifier, and this is the part of the design that is easiest to grasp as a systems argument rather than a modelling one.

A conventional classifier over a codebook of size `2^d` predicts `2^d` logits, which means the output projection alone has vocabulary-size times hidden-size parameters. The paper's worked example makes the scale concrete: with `d = 32` and hidden size `h = 2048`, a conventional classifier needs 8.8 trillion parameters. The infinite-vocabulary classifier predicts `d` bits instead, and for the same configuration needs 0.13 million. That is a reduction of roughly seven orders of magnitude, and it is the reason the head does not become the bottleneck when the codebook grows.

The paper also gives a stability argument for why bits are better targets than indices. Because a conventional head works on continuous features and maps them to a hard index, a small perturbation near a decision boundary can flip the predicted index completely. The supervision signal then jumps discontinuously even though the underlying feature barely moved. Bit labels change gradually under the same perturbation, so the gradient stays steady during training.

Both halves of that argument hold up on inspection. The parameter count is arithmetic. The stability claim is a reasonable expectation about a smoother target, though how much it helps in practice depends on training details that the summary table cannot show.

## Bitwise self-correction and the autoregressive training gap

The third component addresses a problem specific to autoregressive models, and it is the least discussed of the three even though it may be the most broadly useful.

Autoregressive training uses teacher forcing. At each step the model is shown the ground-truth prefix and asked to predict the next token, which gives dense supervision across the whole sequence. The problem is that this is not what happens at inference. At inference the model consumes its own predictions, and a mistake in an early token propagates into every token after it. The model has learned to refine a correct prefix, not to recognise and repair an incorrect one. The paper describes the result as mistakes being propagated and amplified until the image falls apart.

Bitwise self-correction is the proposed mitigation. The framing is that a bit-level prediction target gives the model a way to notice and fix small errors earlier in the sequence than an index-level target would, because a partially wrong code decomposes into bits that are individually recoverable.

This is the piece most worth following in follow-up work, because if the description is accurate it applies to any autoregressive visual model, not just this one. It is also the piece where a reader should be most careful. The README summary does not include ablation numbers for it, so the claim as published is a direction of improvement rather than a quantified result, and anyone building on this should measure it for their own sequence length and tokenizer configuration rather than assume the benefit transfers unchanged.

## Reported numbers, and what they do and do not show

The README quotes specific comparative results, and it is worth reading them as claims from the paper rather than as a benchmark you can reproduce from the repository alone.

Against SD3-Medium on the GenEval benchmark, the paper reports an improvement from 0.62 to 0.73. On ImageReward the reported move is from 0.87 to 0.96, with a win rate of 66 percent. The paper also claims the model generates a 1024x1024 image in 0.8 seconds without extra optimization, which it describes as 2.6 times faster than SD3-Medium and fast enough to set a record at the time.

Two caveats belong next to those figures. First, throughput depends entirely on hardware, and the README does not state which GPU the 0.8 seconds was measured on. Second, GenEval and ImageReward are automatic metrics with known failure modes, and a large gap on ImageReward in particular tends to reflect both genuine quality and the metric's known leniency. Neither figure should be read as a general statement about image quality.

There is also a scope limit worth naming. The reported speed advantage comes from predicting a sequence of bits with a parallel-friendly head rather than a large categorical distribution, and the token count for a 1024x1024 image is still substantial. The claim is that Infinity generates faster than competing autoregressive and diffusion baselines at this resolution. It is not a claim that the model is small.

On that last point, the model sizes are given plainly: Infinity-2B and Infinity-8B. Both are available, with checkpoints at 512x512 and 1024x1024 for the 8B variant, and the 2B is the sensible starting point for anything that is not a dedicated research run.

## The repository, the dependencies, and the release situation

The tree is laid out the way a research codebase usually is. `train.py` and `trainer.py` handle training, `predict.py` handles single prediction, `conf.py` holds configuration for the documentation build, and `cog.yaml` is a container definition for the Replicate deployment. `infinity/` holds the model, `tools/` holds the notebooks, `scripts/` holds helper scripts, `data/` and `evaluation/` cover datasets and metrics, and `assets/` holds the figures used in the README. There is a `DockerFile` with an unconventional capital F.

The dependency list is short and mostly unremarkable, with two pins that matter. `torch` is fixed at 2.5.1 and `timm` at 0.9.6, so a newer PyTorch will not resolve cleanly without edits:

```bash
torch==2.5.1
timm==0.9.6
transformers
einops
```

`flash_attn` is also required, which is worth flagging because it needs a compiler and matching CUDA toolkit rather than a wheel from PyPI, and it is the usual reason a fresh environment fails on the first attempt. `httpx` is pinned at an older 0.20.0 release. The remaining dependencies are standard research tooling: `transformers`, `einops`, `timm`, `omegaconf`, `kornia`, `decord`, `opencv-python`, `imageio`, `wandb`, `pandas`, `seaborn`, `gputil` and others.

There is one thing the repository does not have, and it is the first thing to check for if you plan to depend on it: tagged releases. The release list is empty, so there is no version to pin to and no upgrade path to follow. The MIT licence is permissive and applies to the code as committed. The last push was 2026-04-16, and the update notes around that date describe work on a separate framework rather than changes to Infinity itself, so the sensible reading is that this is a finished research artifact that still builds, rather than a project on a weekly release train.

## Conclusion

Infinity is worth reading as architecture rather than as a checkpoint, because the parts are separable. The bitwise residual quantizer is a contribution to image compression, and the tokenizer numbers in the model zoo show it working: reconstruction quality keeps improving as the codebook size grows, with no collapse. The infinite-vocabulary classifier is a contribution to how a transformer head is parameterised, and the argument against the conventional softmax-over-codes head is a parameter-count argument you can check on paper. Bitwise self-correction is a contribution to autoregressive training, and it addresses the teacher-forcing discrepancy that has always been the weak point of next-token prediction.

The practical caveats are equally clear. There are no tagged releases, so version pinning is on you. Training an 8B model is not a weekend activity, and the token count for a 1024x1024 image means you need serious compute for fine-tuning even though inference is fast. The last push to the repository was 2026-04-16, and the work announced around then has been a follow-up framework rather than changes to this code, so read it as a stable published research artifact rather than an actively churning one. If you want to try it, the 2B checkpoint and the Hugging Face weights are the lowest-cost entry point.

## FAQ

### What is Infinity and who built it?

Infinity is a bitwise visual autoregressive modeling framework for generating high-resolution images, from FoundationVision and ByteDance. The paper is arXiv 2412.04431 and it was accepted as an oral at CVPR 2025. It replaces discrete visual tokens with bit-level prediction and was released with training code, inference code and checkpoints under the MIT licence.

### What is the infinite-vocabulary tokenizer in Infinity?

It is a bitwise multi-scale residual quantizer that encodes an image in stages, with early tokens carrying coarse structure and later tokens carrying finer detail. Its purpose is to cut memory usage so that a very large codebook is trainable at all. The model zoo reports vocabularies of 2 to the 16th, 24th, 32nd and 64th power, with reconstruction quality improving as the vocabulary grows.

### Why does Infinity use a bitwise classifier instead of a normal one?

A conventional classifier over a codebook of size 2 to the d power needs 2 to the d logits, so the output layer alone becomes enormous. The paper's example gives 8.8 trillion parameters for d equals 32 with a hidden size of 2048, against 0.13 million for the bitwise classifier. It is also argued to give steadier supervision, since bit labels change gradually where index labels can flip abruptly.

### How fast is Infinity at generating images?

The README reports a 1024x1024 image in 0.8 seconds without extra optimization, described as 2.6 times faster than SD3-Medium and the fastest autoregressive text-to-image model at the time. The hardware used is not stated, so treat the figure as a paper result rather than something to reproduce directly.

### Which Infinity checkpoints are available and what license applies?

Infinity-2B and Infinity-8B checkpoints are published on Hugging Face, with 512x512 and 1024x1024 variants for the 8B model, plus visual tokenizer checkpoints and the BitVAE tokenizer training code in a separate repository. The code is MIT licensed. Note that the repository has no tagged releases, so there is no version to pin against.

### What does bitwise self-correction do in the Infinity model?

It is the component aimed at the gap between teacher-forced training and inference in autoregressive models. During training the model sees the correct prefix; during generation it consumes its own output, and errors compound. Bitwise self-correction is proposed as a way for the model to detect and repair early errors rather than only refine a correct prefix.

## Sources

- [FoundationVision/Infinity on GitHub](https://github.com/FoundationVision/Infinity)
- [Issues](https://github.com/FoundationVision/Infinity/issues)
- [License: MIT](https://github.com/FoundationVision/Infinity/blob/main/LICENSE)
- [README](https://github.com/FoundationVision/Infinity/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/foundationvision-infinity
