echomimic_v3: the Windows install is a netdisk link with a published passcode
[AAAI 2026] EchoMimicV3: 1.3B Parameters are All You Need for Unified Multi-Modal and Multi-Task Human Animation
At a glance
- What is it?
- echomimic_v3 is Ant Group's 1.3B parameter audio-driven human animation model, built on Alibaba's Wan 2.1 base and shipped in two variants with different audio encoders and different memory floors. The Windows path is a shared drive link, and the dependency file pins TensorFlow exactly.
- Who is it for?
- echomimic_v3 is worth trying if you have a CUDA GPU with 12 GB or more and want audio-driven human animation you can run locally, and the Flash variant is the one to start with because it drops the face mask requirement, generates in 8 steps and stays under 768 by 768. Three things to know before you start.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Activity is slowing. The repository last received commits 6 months ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The documented Windows install is a shared drive link with its passcode printed
There are two installation sections and they are not comparable.
Linux gets a proper source path: a conda environment, then the requirements file.
conda create -n echomimic_v3 python=3.10
conda activate echomimic_v3pip install -r requirements.txtWindows gets one sentence and a hyperlink. It points at a Baidu Pan share described as a one-click installation package, and the passcode is given in the text beside it as `glut`. The sentence also calls it the Quantified version, which reads like a partial translation of a Chinese term rather than a product name you would search for.
So on Windows the documented route is a mirrored archive whose contents cannot be inspected from the documentation, gated behind a credential that is published in the same document. Anyone who wants to know what is inside that archive, or who needs to verify it, has to download it first.
On Linux the equivalent inspection is trivial: the dependencies are a text file, the code is in the repository, and the entry points are visible. `app.py` and `app_mm.py` are the two Gradio interfaces, `infer_preview.py` and `infer_flash.py` are the command-line paths, and `run_flash.sh` wraps the Flash variant.
If you are on Windows and want a reproducible install, the Linux instructions plus WSL are the shorter path to understanding what you are running.
Three exact pins in an otherwise floating requirements file, one of them TensorFlow
The dependency list is mostly unpinned, which for a diffusion stack means you get whatever your resolver finds. Five entries carry a floor, and three are held exactly:
torch>=2.1.2
diffusers>=0.30.1
transformers>=4.46.2
moviepy==2.2.1
tensorflow==2.15.0
retina-face==0.0.17`tensorflow==2.15.0` is the one to think about. A full TensorFlow install sits in the same environment as PyTorch, at an exact version, in a project whose stated minimum CUDA is 12.1 and whose tested GPUs include an A100 and an RTX 4090D. TensorFlow is not listed as a direct requirement of anything else visible here, so it arrives through another package, most plausibly the `mmgp` entry near the bottom of the list.
The cost is concrete. Two deep learning frameworks share one environment, one of them pinned, and the pin has to hold against whatever pulls it in. On a machine with a recent driver this is where an install goes wrong, not in the diffusion stack.
`retina-face==0.0.17` is the second exact pin, and it is the face detection library. That matters because the Flash variant advertises no face mask requirement, while the earlier half-body work in this series depended on landmark conditioning. The detection dependency is still in the environment either way.
`moviepy==2.2.1` is the third. Video assembly is the last step of the pipeline and the pin is cheap, but it is another entry that will need attention when something else in the stack moves.
Preview and Flash are not interchangeable because the audio encoder changes
Two weight sets ship for this model, and the difference is more than file size. The model preparation table pairs each variant with its own audio encoder:
`wav2vec2-base`, from the 960h checkpoint, is labelled the audio encoder for preview. `chinese-wav2vec2-base` is labelled the audio encoder for Flash.
So switching variants means swapping the audio encoder as well as the weights. The English-tuned encoder goes with the preview set and the Chinese one goes with Flash, which tells you something about which language each variant was tuned for.
The memory and quality story moves in the opposite direction to the naming. Preview is the heavier set: the tested GPUs are an A100 with 80 GB, an RTX 4090D with 24 GB and a V100 with 16 GB, so the floor is a 16 GB card. Flash is the lighter one, updated on 2026-01-22 and described as 8-step high-quality generation, no face mask required, a 12 GB VRAM requirement, and support for up to 768 by 768 resolution.
That 12 GB figure is the number most people will act on, and it arrived with a specific front end attached: the update note ties it to the GradioUI in `app_mm.py` and to a third-party tutorial. A separate note from the same day says 16 GB VRAM is enough through a community ComfyUI node, which is a different figure for a different integration.
So the three published numbers, 12 GB, 16 GB and 24 GB, are three different entry points rather than one requirement stated three ways.
The 1.3B figure comes from Alibaba's Wan 2.1 base, and the weights live under another org name
The parameter count in the title is not this project's own model. The base model row in the preparation table is `Wan2.1-Fun-V1.1-1.3B-InP`, taken from the Alibaba PAI organisation on HuggingFace and labelled simply as the base model.
So EchoMimicV3 is a 1.3B model because the Wan 2.1 image-to-portrait checkpoint it builds on is 1.3B. Anything that changes the memory profile or the licence surface of the result starts there.
The weights themselves are a different matter. Both the HuggingFace and the ModelScope entries sit under `BadToBest/EchoMimicV3`, not under the Ant Group organisation that publishes the code. The Flash weights are a subdirectory of that same repository, at `echomimicv3-flash-pro`, and the ModelScope path is the parallel copy.
That means there are two artefacts to keep in step and no single source of truth. The code is on GitHub under `antgroup/echomimic_v3`, the weights are on two model hubs under a third name, and a Gradio studio also runs on ModelScope. Nothing in the repository ties a commit to a weight revision.
The repository has no GitHub releases at all, which removes the mechanism that would normally connect the two. The update log is therefore the only index of what changed, and its most recent entry is the 2026-01-22 Flash update on HuggingFace.
For reproducibility, treat the weight repository as the versioned artefact and record its revision yourself.
No GitHub releases, and the last push was 2026-03-18
This repository has no GitHub releases. There is no tag to check out, no changelog entry per version and no published binary. What exists is the update log in the documentation, dated entries running from the paper appearing on arXiv in July 2025 through to the Flash update in January 2026.
The entries themselves are informative. The paper went public on 2025-07-08. Codes and models were released on 2025-08-08, models again on ModelScope the next day, a Gradio demo on ModelScope on 2025-08-21, acceptance at AAAI 2026 on 2025-11-09, and the Flash weights update on 2026-01-22.
Two entries on the same day, 2025-08-12, carry the memory claims: one says 12 GB VRAM is all you need to generate video and points at the GradioUI, the other says 16 GB works through the community ComfyUI node. Both credit community contributors by name.
The last push to the repository is dated 2026-03-18, which is the most recent activity signal available and is nearly seven months before the latest dated update in the log. Since there are no releases, that date and the model hosts are the only two things you can check when you want to know whether what you downloaded is current.
The series context is in the same log: V1 was audio-driven portrait animation through editable landmark conditioning, V2 moved to simplified semi-body human animation and appeared at CVPR 2025, and V3 is the unified multi-modal and multi-task entry with an arXiv identifier of 2507.03905 and an AAAI 2026 acceptance.
Three video decoders, one ONNX runtime and two face paths in one environment
The requirements file installs video decoding three separate ways: `decord`, `imageio[ffmpeg]` and `imageio[pyav]`. Any one of them can read a video file, and having all three present means the code is free to pick whichever handles a given clip, with the fallback invisible until a specific file fails.
`onnxruntime` is there as well, which is not a library this kind of model obviously needs and is another sign of a dependency arriving through another package rather than being chosen.
The rest of the file is a conventional diffusion stack spread out: `timm` for vision backbones, `einops` and `tomesd` for token manipulation, `safetensors` for weight loading, `accelerate`, `diffusers` and `transformers`, `omegaconf` for configuration, `albumentations` and `scikit-image` for image work, `librosa` for audio, `ftfy` for text cleanup, `func_timeout` for bounded calls, `Pillow`, `numpy`, `SentencePiece` and `tensorboard` for inspection.
`mmgp` sits at the end of the list and is the most likely source of both the TensorFlow pin and part of the multi-GPU story. `decord` and `albumentations` together suggest the training data path is present in this environment rather than inference only, which is consistent with a `datasets/` directory and a `config/` directory at the top level.
The code itself is split into `src/`, two Gradio apps, two inference scripts and a shell wrapper, with assets for the group imagery. Two legal files sit beside them, a `LICENSE.txt` and a separate `LEGAL.md`.
Tested on CentOS 7.2 and Python 3.10, with 3.11 also listed as supported
The environment section names a tested matrix, and it has some age in it.
- Tested System Environment: Centos 7.2/Ubuntu 22.04, Cuda >= 12.1
- Tested GPUs: A100(80G) / RTX4090D(24G) / V100(16G)
- Tested Python Version: 3.10 / 3.11CentOS 7.2 is the older of the two operating systems and reached end of life years ago, so the matrix spans a distribution from a previous generation and a current LTS. That is not a blocker, since the Ubuntu 22.04 entry is the modern one, but it does mean the compatibility claim is broader than the platforms are.
The Python entry is the sharper inconsistency. Both 3.10 and 3.11 are listed as tested, and the installation command creates the environment with `python=3.10`. So the documented path pins the older of the two supported interpreters with no reason given.
The GPU list is worth reading as a memory ladder rather than a support list. The 16 GB V100 is the floor the preview weights imply, the 24 GB RTX 4090D is the comfortable case, and the 80 GB A100 is the data centre case. Flash was later described as needing 12 GB, which is below every GPU in this list, so the published floor for the light variant comes from a later update rather than from the tested matrix.
CUDA 12.1 or newer is the only driver-side constraint stated, and it is consistent with a torch floor of 2.1.2.
Editorial conclusion
echomimic_v3 is worth trying if you have a CUDA GPU with 12 GB or more and want audio-driven human animation you can run locally, and the Flash variant is the one to start with because it drops the face mask requirement, generates in 8 steps and stays under 768 by 768. Three things to know before you start. The Linux path is the supported one, with a conda environment and a requirements file; the Windows path is a third-party netdisk share with a passcode printed in the documentation, which is not something to build a pipeline on. Preview and Flash are not interchangeable, because they pair with different audio encoders and different weight files. And the last push to this repository was 2026-03-18, with no GitHub releases at all, so the model hosts rather than the repository are where any update will appear.
Frequently asked questions
How much VRAM does echomimic_v3 need?
Three figures are published for three entry points: 12 GB for the Flash weights at up to 768x768, 16 GB through the community ComfyUI node, and a tested GPU list of V100 16 GB, RTX 4090D 24 GB and A100 80 GB for the preview weights.
What is the difference between the echomimic_v3 preview and Flash weights?
They pair with different audio encoders. Preview uses `wav2vec2-base` and Flash uses `chinese-wav2vec2-base`. Flash also needs no face mask, generates in 8 steps and caps at 768x768, while preview is the heavier set.
How do I install echomimic_v3 on Windows?
The documented route is a one-click installation package hosted on a netdisk share with the passcode printed beside the link, described as the Quantified version. Linux is the from-source path, with a conda environment and `pip install -r requirements.txt`.
Where are the echomimic_v3 weights hosted?
On HuggingFace and ModelScope under `BadToBest/EchoMimicV3`, with the Flash weights in the `echomimicv3-flash-pro` subdirectory, rather than under the Ant Group organisation that publishes the code. The repository has no GitHub releases.
What base model does echomimic_v3 build on?
The 1.3B parameter figure comes from `Wan2.1-Fun-V1.1-1.3B-InP` from Alibaba PAI, which the model preparation table labels as the base model. The repository also pins `tensorflow==2.15.0` and `retina-face==0.0.17` exactly in its requirements file.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/antgroup-echomimic-v3)