FastVideo: a post-training and inference framework for accelerated video diffusion
A unified inference and post-training framework for accelerated video generation.
At a glance
- What is it?
- FastVideo is an Apache-2.0 Python framework from hao-ai-lab that combines distillation-based post-training with a serving path for video diffusion models. It is powerful, but the install matrix is where most of the work lives.
- Who is it for?
- Adopt FastVideo if you already have a video diffusion checkpoint and want to cut its sampling steps through DMD2 or sparse distillation, or if you need sequence-parallel inference on H100, A100 or 4090 hardware. Do not adopt it if you want a one-line text-to-video API, if you are on a non-Linux NVIDIA platform expecting a prebuilt CUDA kernel wheel, or if you cannot afford to pin torch, transformers and tokenizers exactly as pyproject.toml specifies.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What FastVideo is actually for
FastVideo is not a model. It is the machinery around a model: the distillation recipes that compress a slow diffusion sampler into a few forward passes, and the inference stack that runs the result. The README describes it as "a unified post-training and real-time inference framework for accelerated video generation", and the repository layout backs that up. There is a fastvideo/ package, a separate fastvideo-kernel/ directory for custom CUDA kernels, and an examples/ tree split into distill, train, training, inference and serving. The audience is narrow and identifiable: people who have a diffusion transformer checkpoint, a GPU, and a latency budget they need to hit. If you want to type a prompt into a website and get a video back, this project is not aimed at you. If you want to take MiniMax-H3 or Wan and make it sample in four or eight steps instead of dozens, it is.
Distillation as the core mechanism, not a feature
The pipeline FastVideo describes has two halves that share weights. On the post-training side, the README lists full finetuning and LoRA finetuning for open video DiTs, a data preprocessing pipeline for video, image and text data, Distribution Matching Distillation (DMD2) for stepwise distillation, sparse attention via Video Sparse Attention, and sparse distillation that the project says achieves more than 50x denoising speedup. Causal distillation is handled through Self-Forcing. Training scales with FSDP2, sequence parallelism and selective activation checkpointing. On the inference side, the same accelerated checkpoint is served with sequence parallelism, several attention backends, and both a CLI and a Python API.
The design consequence is that FastVideo is a pipeline you enter at the checkpoint stage, not at the prompt stage. The released FastH3 8-Step V2 checkpoint is described as an eight-forward data-free DMD2 distillation of MiniMax-H3 with 80% Video Sparse Attention. That single sentence tells you the whole shape: a teacher model, a distillation objective, a sparsity ratio, and a step count. If your workflow does not include a teacher model you are willing to distill from, most of the repository is irrelevant to you, and you are left with the inference half.
Installing FastVideo with uv and running a first example
The README recommends uv over Conda and gives the environment creation commands directly. The torch backend is selected through an environment variable rather than a package extra, which is the first thing to get right.
# Create and activate a new uv environment
uv venv --python 3.12 --seed
source .venv/bin/activate
# Install FastVideo on NVIDIA CUDA 12
UV_TORCH_BACKEND=cu126 uv pip install fastvideoOn CUDA 13 the README says to use UV_TORCH_BACKEND=cu130 instead. If you are on an NVIDIA DGX Spark (GB10, ARM64 with CUDA 13), the README states there is no prebuilt ARM wheel for the FastVideo CUDA kernel, so the install is editable from source and compiles that kernel for you:
UV_TORCH_BACKEND=cu130 uv pip install -e .Apple Silicon takes a different route entirely. FastVideo runs FastMetal-QAD through an MLX runtime, installed as an extra:
uv pip install -e '.[mlx]'After that you download FastVideo/FastMetal-1.3B-QAD from Hugging Face and follow the Apple Silicon guide. For a first real run on NVIDIA hardware, the README points at examples/inference/basic/basic_fasth3_8step.py for the FastH3 8-Step V2 checkpoint. Expect the first invocation to spend most of its time downloading weights rather than sampling.
Where the install breaks, and what the README does not cover
The dependency block in pyproject.toml is unusually opinionated, and that is the main practical risk. torch is pinned to exactly 2.12.0, not a range. transformers has a floor of 5.15.0, with a comment explaining that GLM-Image's autoregressive encoder first ships in 5.0.0 and the floor was raised for a vision interpolation helper that the MiniMax-H3 Qwen3-VL encoder cross-checks against. tokenizers is capped below 0.23 because the 0.23 release renamed RobertaProcessing binding arguments, which breaks CLIP-style tokenizer loading. flashinfer-python is gated on sys_platform == 'linux' because there are no wheels for other platforms and the source build shells out to nvcc unconditionally.
Every one of those constraints is a real failure mode someone hit. Together they mean FastVideo does not coexist easily with an existing environment. If you have another project pinning transformers to a 4.x line, you are looking at two virtual environments, not one. The README does not document a rollback path for a failed distillation run, and it does not state memory requirements per model. The support matrix page is where hardware assumptions and optimization compatibility are said to live, and that page is linked rather than reproduced.
FastH3, FastWan and the checkpoint release cadence
FastVideo ships reference checkpoints rather than only code, which changes how you evaluate it. FastH3 Preview v1, released 2026-08-27, is described as an open-weight 4-step sparse-distilled MiniMax-H3 model for synchronized video-and-audio generation, developed with Nuva Lab and the NVIDIA FastGen team. FastH3 8-Step V2 followed on 2026-09-15. FastWan-QAD, released 2026-06-23, is quoted as generating 5 seconds of video in 1.8 seconds end to end, with FP8 and 1.3B variants on Hugging Face. CausalWan2.2 I2V A14B Preview arrived 2025-11-19. The FastMetal-QAD family covers 1.3B, 5B and 14B sizes for Mac.
This cadence is the strongest argument for the project and also the thing to plan around. A checkpoint tied to a specific distillation recipe and a specific sparsity ratio is not a general-purpose model, and swapping the checkpoint usually means revisiting the example script that drives it. The naming is dense: VSA, QAD, DMD2, Self-Forcing, Preview versus V2. Budget time for reading the cookbook entries rather than guessing from filenames.
FastVideo against a plain diffusers pipeline
The obvious alternative is to skip FastVideo and run the upstream model through diffusers directly, which pyproject.toml already depends on at version 0.38.0 or later. The difference in approach is real. A plain diffusers pipeline gives you the model as published, with its original step count and its original attention implementation, and you manage sampling yourself. FastVideo gives you a distillation training loop, sparse attention kernels, sequence parallelism across devices, and a set of pre-distilled checkpoints that trade some fidelity for a large step reduction.
Neither is strictly better. Diffusers is the safer choice when you need the exact published model behaviour, when you are evaluating a checkpoint rather than deploying it, or when your hardware is not on FastVideo's supported list. FastVideo is the choice when the published step count is the problem you are trying to solve. The cost of the FastVideo route is that you inherit its dependency pins and its kernel compilation, and that a distilled checkpoint is a different artifact from the one the model authors released.
Licence, upgrades and the maintenance picture
FastVideo is Apache-2.0, and the package metadata declares the same classifier. That is permissive and compatible with commercial deployment of the framework code. It does not automatically cover the model weights. The FastH3 and FastWan checkpoints live on Hugging Face under the FastVideo organisation, and their own licences are separate documents that the README does not summarise here. Check the model card before shipping anything generated by those checkpoints. This is a factual boundary, not legal advice.
On maintenance, the repository is not archived and the last push was on 2026-09-18, days before this writing. The version in pyproject.toml is 0.2.1, one patch above the v0.2.0 release tagged on 2026-06-04. Upgrade cost is dominated by the pinned dependencies rather than by FastVideo itself: a torch bump, a transformers bump or a tokenizers bump each has the potential to break the tokenizer loading path that the comments in pyproject.toml describe. Treat the environment as something to rebuild from the documented commands rather than to patch in place.
Editorial conclusion
Adopt FastVideo if you already have a video diffusion checkpoint and want to cut its sampling steps through DMD2 or sparse distillation, or if you need sequence-parallel inference on H100, A100 or 4090 hardware. Do not adopt it if you want a one-line text-to-video API, if you are on a non-Linux NVIDIA platform expecting a prebuilt CUDA kernel wheel, or if you cannot afford to pin torch, transformers and tokenizers exactly as pyproject.toml specifies. Before committing, verify three things: that the model you intend to run appears in the support matrix, that your CUDA version matches the UV_TORCH_BACKEND value in the install command, and that the example script for that model exists under examples/inference. On a DGX Spark the README is explicit that the install is editable from source, and on Apple Silicon it points at the MPS guide rather than the default pip path.
Frequently asked questions
What is FastVideo?
FastVideo is an Apache-2.0 Python framework for post-training and real-time inference on accelerated video generation models. It covers distillation recipes such as DMD2 and sparse distillation, plus a serving path with sequence parallelism and multiple attention backends.
How do I install FastVideo?
The README recommends uv: create an environment with uv venv --python 3.12 --seed, activate it, then run UV_TORCH_BACKEND=cu126 uv pip install fastvideo for CUDA 12, or UV_TORCH_BACKEND=cu130 for CUDA 13. Apple Silicon uses uv pip install -e '.[mlx]' and the MPS guide instead.
Does FastVideo support Apple Silicon Macs?
Yes. The README states that FastVideo runs FastMetal-QAD through an MLX runtime, installed as the mlx extra, with the FastVideo/FastMetal-1.3B-QAD checkpoint and the Apple Silicon guide covering setup.
What is the relationship between FastVideo and MiniMax-H3?
FastH3 is FastVideo's distilled MiniMax-H3 line. FastH3 Preview v1 is described as an open-weight 4-step sparse-distilled model for synchronized video-and-audio generation, and FastH3 8-Step V2 as an eight-forward data-free DMD2 checkpoint with 80% Video Sparse Attention.
Which GPUs does FastVideo support?
The README lists H100, A100 and 4090, and states support for Linux, Windows and macOS. It also notes that FastH3 runs on Apple Silicon through MLX and on NVIDIA DGX Spark through CUDA 13, including two-Spark inference.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/hao-ai-lab-fastvideo)