Model or dataset
Wan-Video/Wan2.1 avatar
Wan-Video/Wan2.1

Wan2.1: A Multi-Task Open-Source Video Generation Suite from Wan-AI

Wan: Open and Advanced Large-Scale Video Generative Models

17,029 stars3,593 forksPythonApache-2.0

At a glance

What is it?
Wan2.1 is an Apache-2.0 suite of video foundation models from the Wan-AI team at Alibaba Cloud, supporting text-to-video, image-to-video, video editing, text-to-image, and video-to-audio generation. Its 1.3B parameter model runs on consumer-grade GPUs with 8.19 GB VRAM, while the 14B models target workstations and servers with higher memory.
Who is it for?
Wan2.1 is a reasonable starting point for researchers and engineers who need an open-source video generation model with a permissive license and flexible task support. The T2V-1.3B model makes it accessible on consumer hardware, though four minutes to generate a five-second 480P clip on an RTX 4090 without optimization is a hard constraint for any production use case.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 22, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Wan2.1 Does and the Problem It Addresses

Wan2.1 is a collection of video foundation models that handle multiple generation tasks from a single repository. The core tasks are text-to-video, image-to-video, text-to-image, video editing (through VACE), first-last frame to video (FLF2V), and video-to-audio. These span a range of generation directions, making it possible to use one codebase for several different production needs rather than maintaining separate models for each task.

The project targets two kinds of users: researchers who want access to open-source video generation weights they can fine-tune or study, and engineers who want a model they can run locally without a cloud API dependency. The Apache-2.0 license makes commercial use permissible, which separates it from models with restrictive licenses.

Wan2.1 is also notable as the first video model, according to its documentation, that can generate both Chinese and English text embedded inside video frames. The README describes this as enhancing practical applications where on-screen text is part of the content.

Task Coverage: VACE, FLF2V, and the Model Range

The model suite covers five distinct generation tasks. Text-to-video (T2V) generates clips from a text prompt. Image-to-video (I2V) animates a still image according to a text prompt. VACE (Video Creation and Editing) is an all-in-one model that handles both generation and editing in a single model. FLF2V (First-Last Frame to Video) takes two still images, one for the start and one for the end of the clip, and generates the frames in between. Video-to-audio generates audio to accompany an existing video.

The model comes in two sizes. The T2V-1.3B model is the consumer-GPU variant; the README states it requires 8.19 GB VRAM and can generate a five-second 480P clip on an RTX 4090 in approximately four minutes without quantization or other optimization techniques. The T2V-14B and I2V-14B models are larger and target workstations with higher memory.

Wan2.1 integrates with the Diffusers library for both T2V and I2V, and with ComfyUI. Integration with Diffusers was announced on March 3, 2025, and ComfyUI support on February 27, 2025. These integrations let users run Wan2.1 through frameworks they may already be using rather than directly through the repository's generate.py script.

Installing Wan2.1 Locally

The project requires Python 3.10 or later, as specified in pyproject.toml. The dependencies are listed in requirements.txt and include torch 2.4.0 or later, torchvision, diffusers, transformers, and flash_attn. The flash_attn dependency requires a CUDA-capable GPU and can fail to build on systems without one; the README does not document a CPU-only fallback.

To install from source:

bash
git clone https://github.com/Wan-Video/Wan2.1
cd Wan2.1
pip install -r requirements.txt

The requirements include packages that can be large and version-sensitive:

text
torch>=2.4.0
torchvision>=0.19.0
diffusers>=0.31.0
flash_attn

After installing the Python dependencies, model weights must be downloaded separately from Hugging Face at the Wan-AI organization or from ModelScope. The repository does not include weights, and the README does not provide inline download commands; it points to the Hugging Face and ModelScope pages for the Wan-AI organization.

The pyproject.toml declares an optional dev group with pytest, black, flake8, isort, mypy, and huggingface-hub for development use. The huggingface-hub CLI inclusion suggests the intended weight download path is through that tool, though exact download commands are in INSTALL.md rather than in the README or pyproject.toml.

The T2V-1.3B Model and Consumer GPU Accessibility

The 1.3B parameter model is the main argument for consumer GPU accessibility. At 8.19 GB VRAM, it fits on a 10 GB or 12 GB GPU, which covers RTX 3080 and RTX 4070 class cards. The README states explicitly that performance is comparable to some closed-source models, though it does not name specific models or cite a benchmark source for that claim.

The generation speed constraint is significant: four minutes per five-second 480P clip on an RTX 4090 is a rough baseline. The README notes this is without optimization techniques like quantization, implying that quantized variants can be faster, but Wan2.1 itself does not ship a quantization pipeline. The LightX2V community project integrates Wan2.1 with acceleration techniques, but that is a separate codebase.

The 14B models are a different category. They target workstations and cloud instances with 20 GB or more of VRAM, and the README does not give generation speed figures for them. For anyone who needs higher resolution, longer clips, or faster generation than the 1.3B model provides, the 14B models require a hardware upgrade that removes the consumer GPU accessibility argument.

Wan-VAE: Encoding and Decoding High-Resolution Video

Wan-VAE is the video autoencoder included with Wan2.1. The README describes it as encoding and decoding 1080P videos of any length while preserving temporal information. This matters because many video generation models use VAEs that were designed for image generation and impose length limits or lose temporal coherence at longer durations.

The VAE is listed as a separate component that could serve as a foundation for video and image generation tasks beyond Wan2.1 itself. Several community projects mentioned in the README build on Wan2.1 and its VAE: Helios, which achieves minute-scale video synthesis on a single H100 GPU, and Video-As-Prompt, which uses the I2V-14B model for semantic-controlled generation.

Wan2.2, the successor, introduces a different VAE with a 16x16x4 compression ratio and focuses on 720P at 24fps. Wan2.1's VAE is designed for quality at 480P and 1080P rather than for the specific compression ratio that enables Wan2.2's frame rate.

Licensing and Third-Party Dependencies

The repository is licensed under Apache-2.0, which permits commercial use, modification, and distribution with attribution. The pyproject.toml and LICENSE.txt both confirm this. The license applies to the code in the repository.

Model weights carry separate terms. The README notes that the Beatrice v1 voice changer example in the related software section has a custom license, but more directly, video generation models trained on third-party data often carry terms that restrict specific uses. The README for Wan2.1 does not specify any such restrictions on the model weights themselves, but users deploying the models for commercial content generation should verify the current terms on the Hugging Face model pages.

The flash_attn dependency has its own license. Its build process requires a CUDA toolkit and can add significant time to the initial installation. The transformers, diffusers, and torch dependencies are all permissively licensed. The dashscope dependency in requirements.txt points to Alibaba Cloud's SDK, which may have separate terms for any API calls it enables, though the README does not document how or whether dashscope is used at inference time.

Maintenance Status and Where Wan2.1 Stops

The last push to the Wan2.1 repository was on March 5, 2026. The project has no GitHub releases; versioning is tracked through the README's Latest News section, which lists updates up through May 14, 2025 (the VACE release). The gap between the last documentation update and the last code push suggests maintenance activity continued after the last announced feature.

For users who need current development, the Wan-AI team has released Wan2.2 as a successor, with a different architecture (Mixture-of-Experts), higher training data volume, and 720P at 24fps support on consumer hardware. Wan2.1 remains the appropriate choice when the 14B I2V model's community ecosystem is relevant: the community projects listed in the README (including Helios, ATI, AniCrafter, DriVerse, Wan-Move, and MagicTryOn) are all built on Wan2.1 models and have more extensive fine-tuning and tooling coverage than Wan2.2 at this point.

The project does not include training code. Fine-tuning any of the Wan2.1 models requires external tooling. The README does not document a training pipeline, and the pyproject.toml dev dependencies include only test and linting tools, not training dependencies.

Editorial conclusion

Wan2.1 is a reasonable starting point for researchers and engineers who need an open-source video generation model with a permissive license and flexible task support. The T2V-1.3B model makes it accessible on consumer hardware, though four minutes to generate a five-second 480P clip on an RTX 4090 without optimization is a hard constraint for any production use case. Anyone evaluating it should first check whether the 1.3B or 14B model matches their hardware, whether they need VACE for editing or FLF2V for first-last frame control, and whether the dependency on flash_attn builds cleanly on their target system. The last push to the repository was on March 5, 2026.

Frequently asked questions

What is Wan2.1?

Wan2.1 is an open-source suite of video generation models from the Wan-AI team at Alibaba Cloud, licensed under Apache-2.0. It supports text-to-video, image-to-video, video editing via VACE, first-last frame to video, and video-to-audio generation.

How do I install Wan2.1 locally?

Clone the repository, enter the directory, and install dependencies with pip install -r requirements.txt. Python 3.10 or later is required. Model weights must then be downloaded separately from Hugging Face or ModelScope under the Wan-AI organization.

Is Wan2.1 open source?

Yes. The repository is released under the Apache-2.0 license, which permits commercial use, modification, and distribution with attribution. The license applies to the code; model weight terms should be verified separately on the Hugging Face model pages.

What is Wan2.1 VACE?

VACE is an all-in-one model for video creation and editing released as part of Wan2.1 on May 14, 2025. It handles both generation and editing tasks in a single model. The repository includes inference code and model weights for VACE separately from the base T2V and I2V models.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Wan-Video/Wan2.1 on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/wan-video-wan2-1.svg)](https://hysenlabs.com/projects/wan-video-wan2-1)