# Wan2.2: AI Video Generation with Mixture-of-Experts and Cinematic Aesthetics

> Wan2.2 is an Apache-2.0 video generation suite from the Wan-AI team that introduces a Mixture-of-Experts architecture, a new VAE with a 16x16x4 compression ratio, and a 5B hybrid model that generates 720P video at 24fps from both text and image prompts on consumer GPUs. It also includes specialist models for speech-driven video and character animation.

**Wan-Video/Wan2.2** — Wan: Open and Advanced Large-Scale Video Generative Models

- Repository: https://github.com/Wan-Video/Wan2.2
- Website: https://wan.video
- Stars: 17,598 · Forks: 2,258
- Language: Python
- License: Apache-2.0
- Published: 2026-09-21 · Updated: 2026-09-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/wan-video-wan2-2

## What Wan2.2 Introduces Over Wan2.1

Wan2.2 is the second major release from the Wan-AI team at Alibaba Cloud. The primary architectural change is the introduction of a Mixture-of-Experts (MoE) framework into video diffusion models. In MoE, different expert subnetworks specialize in different parts of the task, which allows the overall model to have more capacity than a dense model of the same compute cost. The README states that Wan2.2 separates the denoising process across timesteps using specialized expert models.

Beyond architecture, Wan2.2 was trained on substantially more data: the README states +65.6% more images and +83.2% more videos compared to Wan2.1. This increase is cited as the reason for improved generalization across motion, semantics, and aesthetics.

The third major addition is aesthetic labeling in the training data. Wan2.2 incorporates detailed labels for lighting, composition, contrast, and color tone, which the README states enables more controllable cinematic style generation. For users whose output quality depends heavily on visual style, this is a meaningful difference from Wan2.1.

The last push to the repository was on September 21, 2026, and the README's Latest News section lists updates through November 2025, with the most recent being the Animate-14B Diffusers integration.

## The Model Family: T2V, I2V, TI2V, S2V, and Animate

Wan2.2 ships with five distinct model types. The T2V-A14B generates video from text prompts. The I2V-A14B generates video from an image with an optional text prompt. The TI2V-5B is the hybrid model that accepts either a text prompt or an image, and is the recommended starting point for consumer GPU users.

The S2V-14B (Speech-to-Video) model generates cinematic video driven by audio input. It was released on August 26, 2025, with inference code, model weights, and a technical report. On September 5, 2025, text-to-speech synthesis support via CosyVoice was added, allowing a user to provide text and have the pipeline generate both audio and the corresponding video. The README notes the model is available through wan.video, a ModelScope Gradio space, and a HuggingFace Gradio space.

Wan2.2-Animate-14B generates character animation and replacement, described as holistic movement and expression replication. It was released on September 19, 2025, and integrated into Diffusers on November 13, 2025. A HuggingFace Space and a ModelScope Studio are available for trying it without a local installation.

The A suffix in T2V-A14B and I2V-A14B indicates the MoE architecture variant. The base 14B model weights differ from Wan2.1's 14B models and are not interchangeable.

## Installing Wan2.2 and Choosing the Right Requirements File

The repository requires Python 3.10 or later, consistent with Wan2.1. The base dependencies are in requirements.txt:

```bash
git clone https://github.com/Wan-Video/Wan2.2
cd Wan2.2
pip install -r requirements.txt
```

Unlike Wan2.1, Wan2.2 splits some dependencies into separate requirements files for the specialist models. To use the Speech-to-Video model:

```bash
pip install -r requirements_s2v.txt
```

To use the Animate model:

```bash
pip install -r requirements_animate.txt
```

The base requirements.txt includes torch, torchvision, torchaudio, diffusers, transformers, accelerate, flash_attn, and imageio with ffmpeg. The torchaudio dependency added in Wan2.2's base requirements is absent from Wan2.1's requirements, reflecting the audio generation capabilities.

Model weights are downloaded separately from Hugging Face at the Wan-AI organization or from ModelScope. The repository includes an INSTALL.md file with detailed installation instructions, though the content of that file is not reproduced in the main README.

ComfyUI support was confirmed on July 28, 2025, with documentation available in both Chinese and English on the ComfyUI documentation site. T2V, I2V, and TI2V models were integrated into Diffusers on the same date.

## The TI2V-5B Model: Hybrid Input at 720P and 24fps

The TI2V-5B is the most consumer-accessible model in Wan2.2. It uses the Wan2.2-VAE with a 16x16x4 compression ratio, which the README states enables efficient high-definition output. The model supports both text-to-video and image-to-video generation from a single set of weights, distinguishing it from models that require separate weights for each direction.

At 720P resolution and 24fps, it is positioned as one of the fastest models at that resolution and quality level, according to the README, which states it can serve both industrial and academic use cases simultaneously. A HuggingFace Space using the TI2V-5B model was opened on July 28, 2025.

The older Wan2.1 T2V-1.3B model produced 480P output. TI2V-5B at 720P is a meaningful resolution increase, though the 5B parameter count places it above the T2V-1.3B in VRAM requirements. The README does not give a specific VRAM figure for TI2V-5B, but states it runs on consumer-grade cards like the RTX 4090, consistent with a 10-16 GB range.

The Wan2.2-VAE's 16x16x4 compression ratio is specific to this release. Wan2.1's VAE had different compression characteristics optimized for 480P and 1080P encoding rather than for the frame rate and resolution target of TI2V-5B.

## Speech-to-Video: Wan2.2-S2V-14B and the CosyVoice Path

The S2V-14B model takes audio as input and generates video that corresponds to the spoken content or musical material. The technical report is linked from the README. The initial release supported audio files as input; on September 5, 2025, text-to-speech synthesis via CosyVoice was added as a separate step, so the pipeline can now accept text, synthesize speech, and then generate video.

This distinguishes it from I2V models, which generate video from a static image. S2V generates video that is driven by the timing and content of audio, which is relevant for producing talking-head videos, musical visualizations, or any case where the audio determines the visual rhythm.

The CosyVoice dependency is not in the base requirements.txt; the requirements_s2v.txt file covers it separately. CosyVoice is a separate project maintained by the FunAudioLLM team. Users who want the full text-to-speech-to-video pipeline need to install both files.

Practically, the S2V-14B model requires more VRAM than the TI2V-5B, though the README does not specify an exact requirement. At 14B parameters with the MoE architecture, it is not a consumer GPU model.

## Limitations and the Cases Where Wan2.1 Remains Relevant

Wan2.2 does not include a VACE (video creation and editing) equivalent as of the current repository state. VACE was a notable feature of Wan2.1 and allowed video editing alongside generation from a single model. Users who need video editing capabilities may need to stay with Wan2.1's VACE model until Wan2.2 adds a comparable feature.

The community ecosystem around Wan2.1 is larger at this point. Projects like Helios, ATI, Wan-Move, AniCrafter, MagicTryOn, and DriVerse are built on Wan2.1 weights and have more extensive tooling, fine-tuning examples, and documentation. Community projects listed in the Wan2.2 README are fewer and more recent.

The MoE architecture introduces a structural dependency: the T2V-A14B and I2V-A14B models are not drop-in replacements for the Wan2.1 14B models. Any existing inference pipeline, fine-tune, or ComfyUI workflow built for Wan2.1 14B models requires updating to handle the MoE variant.

Flash_attn remains a hard dependency in the base requirements. It requires a CUDA GPU and can fail to build on older CUDA toolkit versions. The LightX2V community project provides step-distilled and quantized Wan2.2 models that may reduce VRAM requirements, but that is a separate codebase not maintained by the Wan-AI team.

## License, Distribution, and Integration with Diffusers

Wan2.2 is released under Apache-2.0, the same license as Wan2.1. The code and, by implication, the model weights are available for commercial use with attribution. The Hugging Face model pages should be checked for any additional terms attached to specific model weights.

All three main model types (T2V, I2V, and TI2V) are available in Diffusers as of July 28, 2025. The Diffusers model IDs are Wan-AI/Wan2.2-T2V-A14B-Diffusers, Wan-AI/Wan2.2-I2V-A14B-Diffusers, and Wan-AI/Wan2.2-TI2V-5B-Diffusers. The Animate-14B model was added to Diffusers later, on November 13, 2025, at Wan-AI/Wan2.2-Animate-14B-Diffusers. Using the Diffusers integration avoids dealing with the repository's generate.py script directly and fits into existing Diffusers-based pipelines.

The repository has no GitHub releases; versioning is tracked through the README's Latest News section rather than tagged releases. Updates are frequent: the README shows entries from July through November 2025, covering new model releases and integration announcements.

## Conclusion

Wan2.2 suits engineers and researchers who need an open-source video generation model with 720P output at 24fps on consumer hardware, or who want speech-to-video or character animation capabilities under a permissive license. The MoE architecture means the 14B models carry more capacity without proportionally more compute, but consumer GPU users will find the TI2V-5B the most practical entry point. Before deploying, check that the target GPU can handle the model's VRAM requirement, confirm that flash_attn builds on your CUDA version, and review which requirements file (base, s2v, or animate) matches the tasks you need.

## FAQ

### Is WAN 2.2 free?

The code and model weights are released under Apache-2.0, which permits free use, including commercial use with attribution. The models are available to download from Hugging Face and ModelScope at no cost.

### How do I install WAN 2.2 locally?

Clone the repository, enter the directory, and run pip install -r requirements.txt. Python 3.10 or later is required. For the Speech-to-Video or Animate models, install their respective requirements files (requirements_s2v.txt or requirements_animate.txt) separately. Model weights must be downloaded from Hugging Face or ModelScope under the Wan-AI organization.

### Is there a Wan2.2 VAE?

Yes. Wan2.2 introduces the Wan2.2-VAE with a 16x16x4 compression ratio, which enables the TI2V-5B model to generate 720P video at 24fps. It differs from the Wan2.1 VAE, which was designed for quality at 480P and 1080P resolutions.

### Does Wan2.2 generate images?

Wan2.2 is focused on video generation. The TI2V-5B model accepts either a text prompt or an image as input to generate video. The README does not document a standalone text-to-image mode in Wan2.2, unlike Wan2.1 which included text-to-image as a listed task.

## Sources

- [Issues](https://github.com/Wan-Video/Wan2.2/issues)
- [License: Apache-2.0](https://github.com/Wan-Video/Wan2.2/blob/main/LICENSE)
- [Project website](https://wan.video)
- [README](https://github.com/Wan-Video/Wan2.2/blob/main/README.md)
- [Wan-Video/Wan2.2 on GitHub](https://github.com/Wan-Video/Wan2.2)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/wan-video-wan2-2
