Causal Forcing: autoregressive diffusion distillation for real-time interactive video
[ICML 2026] Official codebase for "Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation" & Causal Forcing++
At a glance
- What is it?
- Causal Forcing is a Tsinghua-led training method that turns Wan-based autoregressive video diffusion into few-step generators. It ships checkpoints, a three-stage training pipeline and a CLI, and it stops at 81 frames.
- Who is it for?
- Use Causal Forcing if you work inside the Wan ecosystem, have the GPU budget for a three-stage distillation run, and need few-step interactive generation where the frame-wise 1-step or 2-step checkpoints under causal-forcing++/ are the payoff. Skip it if you need long-form video without the separate long-video extension, or if you expect a hosted endpoint or a ComfyUI node, neither of which the README documents.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 6 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap Causal Forcing is built to close
Autoregressive video diffusion models generate frames in sequence, which is what makes interactive use possible: you can feed a new prompt or an action between chunks and the model keeps going. The problem is speed. A full diffusion sampler needs many denoising steps per frame, and a self-forcing style pipeline that trains on its own rollouts is expensive to run and easy to destabilize. Causal Forcing targets that specific bottleneck. The README frames the work as using Causal ODE or Causal Consistency Distillation as what it calls a theoretically correct initialization for asymmetric DMD, and the goal is few-step generation at real-time latency. The audience is narrow and technical: researchers and engineers who already work with Wan-family video diffusion, who have multi-GPU training access, and who want to distill a slow teacher into a 1, 2 or 4-step student. It is not a consumer video tool and it is not a hosted API. If you want to type a prompt into a web page, this repository is the wrong layer of the stack.
How the three-stage pipeline actually fits together
The training path is staged, and each stage produces the input for the next. Stage 1 is autoregressive diffusion training on the base model. Stage 2 has two options: Causal ODE, which is the original Causal Forcing route, or Causal Consistency Distillation, which the README labels as the Causal Forcing++ route. Stage 3 is asymmetric DMD, the distribution-matching step that compresses the model down to few-step sampling. The repository layout matches this: the top level contains train.py, trainer/, pipeline/, model/, wan/, configs/ and prompts/, plus three separate data-curation scripts, get_causal_ode_data_chunkwise.py, get_causal_ode_data_framewise.py and get_causal_ode_data_kv_optimized.py. Those scripts exist because the ODE route needs paired data generated ahead of time. Causal Forcing++ exists largely to remove that step: the README says it replaces ODE with causal consistency distillation to eliminate ODE data curation and improve performance. The architecture also comes in two shapes. Chunk-wise models denoise blocks of frames; frame-wise models work one frame at a time and, per the README, natively unify text-to-video and image-to-video in a single model rather than needing separate conditioning paths. That unification is the more interesting design decision, because it means I2V is not a bolt-on task head but a consequence of the frame-wise formulation.
Installing Causal Forcing and running a first inference
The README states that the inference environment is identical to Self Forcing, and the installation section opens with a conda environment on Python 3.10. The environment name in the create command and the name in the activate command differ in the README as written, so use one consistent name when you run it.
conda create -n causal_forcing python=3.10 -y
conda activate causal_forcingThe dependency set is pinned in requirements.txt. Notable entries are torch>=2.4.0, diffusers==0.31.0, transformers>=4.49.0, numpy==1.24.4, pydantic==2.10.6 and av==13.1.0. Several pins are exact rather than minimums, so expect resolution conflicts if you install into an environment that already carries a different diffusers or numpy. setup.py registers the package as causal_forcing at version 0.0.2 with find_packages().
pip install -r requirements.txt
pip install -e .Weights come from Hugging Face under zhuhz22/Causal-Forcing, split by variant: chunkwise/causal_forcing.pt, framewise/causal_forcing.pt, causal-forcing++/framewise-2step.pt, causal-forcing++/framewise-1step.pt and chunkwise/longvideo.pt. The README points to inference.py for the CLI path and demo.py plus demo_utils/ for the interactive demo. Prompts live under prompts/, and configs/ holds the training configurations. The README does not spell out the exact inference.py flags in the section reproduced here, so check the script's own argument parser before assuming a flag name. What you should expect after a successful run is a generated clip from a text prompt, or from an image plus prompt for the frame-wise models, at the step count of the checkpoint you chose.
The 81-frame ceiling and why the README calls long-video comparisons unfair
The clearest limitation is stated plainly: Causal Forcing does not natively support videos longer than 81 frames, and the README groups it with CausVid and Self Forcing on that point. The authors go further and warn that using the 5-second trained model directly as a baseline for long video generation is extremely unfair. That is an unusually direct caveat, and it should govern how you read any comparison table that pits Causal Forcing against a minute-scale generator. The repository does ship a path around the limit: chunkwise/longvideo.pt is described as a minute-long autoregressive video generator, and the README credits Rolling Forcing support for minute-level generation. But that is an extension, not the default behaviour, and the README treats it as orthogonal to the base training method. The second constraint is compute. This is a distillation pipeline with three stages, and Stage 2 in the ODE route requires generating paired data before training can start. The release notes mention a 3x speedup for chunk-wise ODE data curation and a 3x faster consistency distillation infrastructure, which tells you what the baseline cost looked like before those optimizations. If you only have one GPU, or you want to fine-tune on a small custom dataset, this is a poor fit.
Causal Forcing versus Self Forcing and the initialization debate
Self Forcing is the direct reference point, and the README makes a specific claim: Causal Forcing significantly outperforms Self Forcing in both visual quality and motion dynamics while keeping the same training budget and inference efficiency. The mechanism behind that claim is the initialization. Self Forcing initializes from autoregressive diffusion directly, while Causal Forcing inserts a causal ODE or causal consistency distillation stage in between. The repository's own FAQ, added on 2026-02-28, addresses exactly this question of which initialization is better between AR diffusion and causal ODE distillation, which suggests it was a recurring point of confusion rather than a settled one. Causal Forcing++ pushes further by dropping ODE for causal consistency distillation and releasing what the README calls the first 1-step and 2-step frame-wise models. The 2-step frame-wise checkpoint is described as even better than the 4-step model, which is counterintuitive enough that you should verify it on your own prompts rather than take it as given. Outside this family, the broader alternative is simply not distilling at all: run a full multi-step sampler and accept the latency. That trades throughput for the freedom to skip a three-stage training pipeline entirely.
Adoption, licensing and what to check before committing
The licence is Apache-2.0, which is permissive and includes an explicit patent grant. That matters here because the codebase builds on Wan, and the README links a separate repository, shengshu-ai/minWM, for the HY1.5-TI2V-8B model and the action-conditioned world model, with weights hosted under MIN-Lab/minWM. Those are different artifacts under their own terms, so the Apache-2.0 grant on this repository does not automatically extend to every checkpoint you might download. Read the model card for whichever weights you use. On maintenance, the last push was on 2026-08-28, and the news log runs from 2026-02-02 through 2026-07-23 with releases, optimizations and external adoptions recorded along the way. Upgrade cost is the real risk. requirements.txt pins diffusers==0.31.0, numpy==1.24.4 and pydantic==2.10.6 exactly, so moving to a newer diffusers or numpy will require testing rather than a version bump. The README does not document a migration path between checkpoint versions or a rollback procedure for a failed training run.
Who should adopt Causal Forcing, and who should not
Adopt it if you are already inside the Wan ecosystem, you have the GPU budget for a three-stage distillation run, and your target is interactive or real-time generation where a few seconds of video at low step counts is the product. The frame-wise 1-step and 2-step checkpoints under causal-forcing++/ are the most distinctive artifacts here, and the frame-wise formulation's native T2V and I2V unification is the part most likely to save you engineering time. Do not adopt it if you need long-form video out of the box; the 81-frame limit is explicit, and the long-video path is a separate extension with its own checkpoint. Do not adopt it if you want a hosted endpoint or a ComfyUI node, because the README documents neither. Before you commit, verify three things in order: that inference.py accepts the arguments you need for your checkpoint variant, that the pinned diffusers and numpy versions coexist with the rest of your environment, and that the 2-step frame-wise model actually holds up on your prompt distribution rather than only on the paper's examples. The last one is the only claim in the README that a single afternoon of your own inference runs can settle.
Editorial conclusion
Use Causal Forcing if you work inside the Wan ecosystem, have the GPU budget for a three-stage distillation run, and need few-step interactive generation where the frame-wise 1-step or 2-step checkpoints under causal-forcing++/ are the payoff. Skip it if you need long-form video without the separate long-video extension, or if you expect a hosted endpoint or a ComfyUI node, neither of which the README documents. Verify first that inference.py accepts the arguments your checkpoint variant needs, and confirm the exact pins on diffusers==0.31.0 and numpy==1.24.4 resolve in your environment before you start a training run, because the README documents no rollback path.
Frequently asked questions
What are causal forces in the context of Causal Forcing?
The project does not define the phrase as a standalone term. Its README describes Causal Forcing as using Causal ODE or Causal Consistency Distillation as the initialization for asymmetric DMD, which is the mechanism behind the name.
How do I install Causal Forcing?
The README creates a conda environment on Python 3.10, then installs requirements.txt and the package via setup.py. The README states the inference environment is identical to Self Forcing.
Does Causal Forcing support image-to-video generation?
Yes. The README says the frame-wise models natively unify T2V and I2V, and I2V support was added on 2026-02-11. The HY1.5-TI2V-8B model is hosted separately under MIN-Lab/minWM.
What is the difference between Causal Forcing and Causal Forcing++?
Causal Forcing++ replaces the ODE stage with causal consistency distillation, which the README says eliminates ODE data curation and improves performance. It also releases the first 1-step and 2-step frame-wise models.
How many frames can Causal Forcing generate?
The README states it does not natively support videos longer than 81 frames, similar to CausVid and Self Forcing. A separate chunkwise/longvideo.pt checkpoint is offered for minute-level generation.
What licence does Causal Forcing use?
The repository is licensed under Apache-2.0. Note that some linked checkpoints, such as those under MIN-Lab/minWM, are separate artifacts with their own terms.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/thu-ml-causal-forcing)