# TurboDiffusion: 100-200x Acceleration for Video Diffusion on Consumer and Data Center GPUs

> TurboDiffusion is an Apache-2.0 Python framework that speeds up video diffusion inference by 100 to 200 times using a combination of sparse attention and consistency model distillation. It is designed for researchers and practitioners who need fast video generation on RTX 5090 or H100-class hardware.

**thu-ml/TurboDiffusion** — TurboDiffusion: 100–200× Acceleration for Video Diffusion Models

- Repository: https://github.com/thu-ml/TurboDiffusion
- Website: https://arxiv.org/pdf/2512.16093
- Stars: 3,851 · Forks: 277
- Language: Python
- License: Apache-2.0
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/thu-ml-turbodiffusion

## Video Diffusion at 100-200x: What TurboDiffusion Solves

Standard video diffusion inference is slow. The README gives concrete figures from a single RTX 5090: a 5-second clip from Wan2.1-T2V-1.3B-480P takes 184 seconds end-to-end with the original model. With TurboDiffusion, the same clip takes 1.9 seconds. For the larger Wan2.2-I2V-A14B-720P model doing image-to-video, the original takes 4,549 seconds; TurboDiffusion brings that to 38 seconds. These times measure end-to-end diffusion generation latency and explicitly exclude text encoding and VAE decoding.

The use case is practitioners who need to iterate on video generation at scale or who run inference on hardware where each generation cycle otherwise costs minutes. TurboDiffusion targets RTX 5090 and RTX 4090-class consumer GPUs as well as data center GPUs like the H100. The framework does not build new model architectures; it accelerates existing Wan-series video diffusion models from Wan-AI using three inference optimizations applied together.

## How TurboDiffusion Works: SageAttention, SLA, and rCM

TurboDiffusion combines three distinct acceleration techniques. SageAttention replaces the standard attention computation with an optimized kernel implementation that reduces attention latency. SLA, which stands for Sparse-Linear Attention, replaces the quadratic attention mechanism with a sparser approximation that cuts computation further. These two components target the attention operations inside the transformer blocks of the diffusion model.

The third technique is rCM, short for recursive Consistency Model. This is a timestep distillation method originally from NVlabs that compresses the many denoising steps of the standard diffusion process into 1 to 4 steps. The inference script exposes this via a --num_steps argument accepting values from 1 to 4, with a default of 4. A related --sigma_max parameter sets the initial sigma for rCM; the README notes that larger values such as 1600 reduce diversity but may improve quality, with 80 as the default.

The I2V (image-to-video) model also uses a two-checkpoint architecture: a high-noise model and a low-noise model, with a --boundary argument that sets the timestep boundary for switching between them, defaulting to 0.9. The high and low noise checkpoints are downloaded separately.

## Installing TurboDiffusion and Downloading Checkpoints

The required base environment is Python 3.9 or higher and PyTorch 2.7.0 or higher. The README recommends torch==2.8.0 specifically, noting that higher versions may cause out-of-memory errors. Installation from pip:

```bash
conda create -n turbodiffusion python=3.12
conda activate turbodiffusion

pip install turbodiffusion --no-build-isolation
```

The `--no-build-isolation` flag is required because the package compiles CUDA extensions at install time. For contributors or those who prefer source installation:

```bash
git clone https://github.com/thu-ml/TurboDiffusion.git
cd TurboDiffusion
git submodule update --init --recursive
pip install -e . --no-build-isolation
```

To enable SageSLA, which combines SageAttention with the SLA forward pass for additional speed, install the SpargeAttn package first:

```bash
pip install git+https://github.com/thu-ml/SpargeAttn.git --no-build-isolation
```

Checkpoints require downloading the shared VAE and text encoder from Wan-AI, then the model-specific checkpoint from TurboDiffusion's Hugging Face organization:

```bash
mkdir checkpoints
cd checkpoints
wget https://huggingface.co/Wan-AI/Wan2.1-T2V-1.3B/resolve/main/Wan2.1_VAE.pth
wget https://huggingface.co/Wan-AI/Wan2.1-T2V-1.3B/resolve/main/models_t5_umt5-xxl-enc-bf16.pth
```

The pyproject.toml lists a substantial set of runtime dependencies including triton, flash-attn, einops, imageio with ffmpeg support, transformers, and nvidia-ml-py. These are installed automatically via pip but flash-attn itself has its own CUDA compilation requirement.

## Available Models and GPU-Specific Checkpoint Selection

TurboDiffusion ships four checkpoint variants as of the last README update:

- TurboWan2.1-T2V-1.3B-480P: text-to-video, 1.3 billion parameters, best at 480p
- TurboWan2.1-T2V-14B-480P: text-to-video, 14 billion parameters, best at 480p
- TurboWan2.1-T2V-14B-720P: text-to-video, 14 billion parameters, best at 720p
- TurboWan2.2-I2V-A14B-720P: image-to-video, approximately 14 billion parameters, best at 720p

The README notes that all checkpoints technically support 480p and 720p; the "Best Resolution" designation indicates where each model was trained to perform best.

The GPU determines which checkpoint variant to download. For RTX 5090, RTX 4090, or similar GPUs, use quantized checkpoints (file names ending in -quant.pth) and add --quant_linear to the inference command. For GPUs with more than 40GB of memory such as the H100, use unquantized checkpoints (without -quant in the name) and omit --quant_linear.

For the I2V model, two checkpoint files are required: a high-noise checkpoint and a low-noise checkpoint. These cover different parts of the denoising trajectory and must both be specified in the inference command.

All checkpoints are hosted on Hugging Face under the TurboDiffusion organization. The README states the checkpoints are not finalized and will be updated to improve quality.

## What TurboDiffusion Does Not Handle

The prompt limitation is the most significant constraint for practical use. The README states directly: "The current models are only trained on long English prompts. If you use other types of prompts, please augment them to get better performance." Short prompts, non-English prompts, and mixed-language prompts are outside the training distribution. The README does not describe what augmentation means in practice or provide examples.

The hardware dependency chain is demanding. CUDA compilation is required at install time, torch 2.8.0 is recommended over newer releases due to documented OOM behavior, and the inference is only benchmarked on a single RTX 5090. The setup.py compiles CUDA extensions with targets covering sm_80 through sm_120a, but the README's inference guidance addresses only the RTX 5090/4090 and H100 classes. Whether A100 or other Ampere-class GPUs produce correct results is not addressed.

The repository has no GitHub releases. Version 1.0.0 is set in pyproject.toml, but the README notes the paper and checkpoints are not finalized. Teams relying on stable, frozen releases for reproducibility cannot treat the current state as production-ready. The repository is an active research release rather than a stable software package.

## The Original Wan Models Without Acceleration

The natural baseline is running Wan2.1 or Wan2.2 directly through their standard inference pipelines without TurboDiffusion. Wan-AI's models use a full diffusion denoising loop with many timesteps and standard attention. TurboDiffusion is a drop-in distillation of those models, not a separately trained architecture.

The difference in approach is the distillation step. The original Wan models require many denoising iterations; TurboDiffusion's rCM distillation compresses that to 1 to 4 steps. Without this distillation, adding SageAttention or SLA alone would not achieve a 100x speedup. The README's benchmark figures make this concrete: the original Wan2.2-I2V model at 4,549 seconds versus TurboDiffusion's 38 seconds on the same GPU is not achievable with attention optimization alone at the standard step count.

For users whose GPU is a high-end consumer card and who cannot wait minutes per generation, TurboDiffusion's distilled checkpoints are the only path to sub-minute inference on those models. Users who have access to very large GPU clusters and can parallelize standard inference across many devices, or who do not need the Wan-series models specifically, have other options. TurboDiffusion is specifically a distilled version of Wan models, not a general acceleration framework applicable to arbitrary video diffusion architectures.

## Conclusion

Practitioners who need fast text-to-video or image-to-video generation on RTX 5090 or H100-class hardware and can tolerate a torch 2.7.0 environment with CUDA source compilation should evaluate the Wan2.1 and Wan2.2 checkpoints. The models are trained exclusively on long English prompts, so quality on short or non-English input is undocumented. The README states the checkpoints and paper are not finalized and will be updated.

## FAQ

### Does TurboDiffusion support GPUs other than the RTX 5090 and H100?

The README describes two GPU classes: GPUs with more than 40GB of memory such as the H100, which use unquantized checkpoints without --quant_linear, and GPUs like the RTX 5090 and RTX 4090, which use quantized checkpoints with --quant_linear. The setup.py compiles CUDA targets including sm_80 through sm_120a, but performance and correctness on other GPU models is not documented.

### Why are TurboDiffusion models only trained on long English prompts?

The README states this as a current limitation without giving a reason: the models were trained on long English prompts only. It recommends augmenting other prompt types before use but does not document what augmentation method to apply.

### What does the --num_steps argument control in TurboDiffusion inference?

According to the inference script argument descriptions in the README, --num_steps controls the number of sampling steps and accepts values from 1 to 4, with 4 as the default for both T2V and I2V models. The default --sigma_max of 80 can be raised to values like 1600 to reduce diversity but potentially improve quality.

## Sources

- [Issues](https://github.com/thu-ml/TurboDiffusion/issues)
- [License: Apache-2.0](https://github.com/thu-ml/TurboDiffusion/blob/main/LICENSE)
- [Project website](https://arxiv.org/pdf/2512.16093)
- [README](https://github.com/thu-ml/TurboDiffusion/blob/main/README.md)
- [thu-ml/TurboDiffusion on GitHub](https://github.com/thu-ml/TurboDiffusion)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/thu-ml-turbodiffusion
