Open-source project
antgroup/echomimic avatar
antgroup/echomimic

EchoMimic: audio-driven portrait animation you can steer with landmarks

[AAAI 2025] EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditioning

4,300 stars466 forksPythonApache-2.0

At a glance

What is it?
The V1 release from Ant Group's Terminal Technology Department animates a still portrait from a speech clip, and lets you hold parts of the face still by editing facial landmarks. Apache-2.0, Python, four inference scripts.
Who is it for?
EchoMimic V1 is a research release that happens to be usable: audio alone gets you a talking portrait, and the landmark conditioning path is the part with real editorial value because a wrong mouth shape can be corrected at the source instead of retrained.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 20, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What the first EchoMimic release is trying to do

The repository title states the whole idea: lifelike audio-driven portrait animations through editable landmark conditioning. Input is a still image of a person and a clip of speech. Output is a video of that person appearing to say the words, with head motion and mouth shapes following the audio.

The interesting part is the word editable in that subtitle. Most audio-driven talking head systems take audio and produce motion, and whatever the face does is what you get. EchoMimic routes part of the signal through facial landmarks, which means a specific region of the face can be held in place or corrected. Pin the mouth and let the head move, or pin the eyes while the mouth speaks. That is an unusual amount of control to expose from a research codebase, and it is the reason the gallery is split into separate Audio Driven, Landmark Driven, and Audio plus Selected Landmark Driven sections rather than one undifferentiated set of clips.

The project comes from the Terminal Technology Department at Alipay, Ant Group, and the author list marks Zhiyuan Chen and Jiajiong Cao as equal contributors with Yuming Li and Chenguang Ma as corresponding authors. The paper appeared on arXiv in July 2024 and the update log records acceptance by AAAI 2025 on 2024.12.10. The code is Python under Apache-2.0, and the project page is hosted under the Ant Group GitHub Pages space at antgroup.github.io.

Four inference scripts, and the file layout behind them

The repository root is small enough to read in one screen, and almost every entry maps to a decision a user has to make. There is `infer_audio2vid.py` for the original audio driven path, `infer_audio2vid_acc.py` for the accelerated audio path, `infer_audio2vid_pose.py` which adds pose and audio driving, and `infer_audio2vid_pose_acc.py` which is the accelerated version of that. The naming convention is consistent enough to guess right, but it is worth reading because the `_acc` suffix is the difference between a wait measured in minutes and a wait measured in seconds.

Alongside those four sit `webgui.py`, which the update log credits to community contributions, `demo_motion_sync.py` for matching audio to existing motion, and the supporting directories `src/` for library code and `configs/` for the configuration files the scripts read. There is also a `requirements.txt` and an `assets/` directory for sample material.

What is absent matters as much as what is present. There is no Dockerfile, no `setup.py`, no `pyproject.toml` and no directory of pretrained checkpoints in the tree. Weights live on HuggingFace and ModelScope under the BadToBest namespace, and the scripts expect you to have them in place locally. For a project with 119 open issues, a good number of them are likely about exactly this part of the setup.

The acceleration work, with the numbers the project actually published

The update log is the most informative document in the repository, and the entry worth reading twice is dated 2024.07.17. Accelerated models and pipeline for audio plus selected landmarks were released, with the stated effect that inference speed improves by 10x, from roughly seven minutes for 240 frames to roughly 50 seconds for 240 frames on a V100 GPU. A parallel entry dated 2024.07.25 covers the audio driven pipeline.

That is a specific, dated, hardware named claim, which is rarer and more useful than a marketing sentence. It also explains the two-track repository: the unaccelerated scripts stay in the tree because the accelerated variants depend on different model weights, not because the slower path is preferred.

The other timeline entries explain the project's shape. Audio driven codes and models shipped 2024.07.09, the paper went public 2024.07.12, WebUI and Gradio versions followed the same day, pose and audio driven codes came 2024.07.13, a ComfyUI integration was contributed 2024.07.14, and the two Gradio demos on HuggingFace and ModelScope were ready on 2024.07.23. An A100 hosted version appeared on HuggingFace on 2024.08.02. The last update entry is 2024.12.10, though the repository itself was pushed to on 2026-04-07, which is recent enough to mean it has not been abandoned even though the public timeline stopped a while back.

Where you can run it without setting anything up

For a model like this, the hosted options are not a consolation prize, they are the first thing to try. The README links a Gradio demo on HuggingFace under fffiloni, a Gradio demo on ModelScope, and later an A100 backed version on HuggingFace hosted for ModelScope users. Model weights are on both platforms as well, so you can follow a hosted run first and only then decide whether the local install is worth the dependencies.

There is a video installation tutorial too, linked from the update log for 2024.07.13, which covers the local route end to end. And because ComfyUI integration arrived early, on 2024.07.14, anyone already working in node based image and video pipelines can reach the model without touching the Python scripts at all.

What none of these routes document is a video quality or fidelity comparison against other talking head systems, and the README does not claim one. Judging output quality from the gallery alone is guesswork, which is a good reason to spend ten minutes on a hosted demo before installing a pinned torch range.

V1, V2 and V3, and why this repository stopped moving

The README frames EchoMimic as a series rather than a single project. This repository is EchoMimicV1. EchoMimicV2 is described as towards striking, simplified, and semi-body human animation and lives in its own repository. EchoMimicV3 is described as 1.3B parameters are all you need for unified multi-modal and multi-task human animation, and it lives in another.

That framing explains a lot of the V1 repository's shape. Semi-body and multi-modal are V2 and V3 territory, so the V1 tree stays focused on portrait animation with audio, landmarks and optional pose. If you are starting a project now and have no attachment to the AAAI paper, the later versions are the more likely destination, and V1's value is as the clearest, smallest surface for understanding the landmark conditioning idea.

The three repositories are maintained by the same group, so choosing between them is a scope decision rather than a quality judgement. Portrait only, then, means this one. Upper body and simpler conditioning, then V2. Unified multi task work, then V3.

Dependencies and the pins that decide whether it installs

The dependency file is the part most likely to cost you an afternoon, and the pins are tight. `torch` is held between 2.0.1 and 2.2.2, `torchvision` between 0.15.2 and 0.17.2, and `torchaudio` between 2.0.2 and 2.2.2, which together mean a single torch family version across the stack. `diffusers` is held at exactly 0.24.0, a hard pin rather than a floor, and `transformers` requires at least 4.38.2. `mediapipe` is listed, which is a hint about where facial landmark extraction happens in the pipeline.

The rest is a standard diffusion video assembly: `einops`, `omegaconf` for the configs directory, `opencv-python`, `av` pinned to 11.0.0, `moviepy` at 1.0.3, `ffmpeg-python`, `torchmetrics`, `torchtyping`, `facenet_pytorch`, `tqdm` and `gradio` for the web UI.

Two practical readings follow. The upper torch bound means a newer CUDA stack or a workstation GPU with a newer bundled torch will not work without an edited dependency file, and since the README does not document a tested combination beyond that range, expect to resolve that yourself. And `mediapipe` plus `facenet_pytorch` together mean the setup pulls in a face detection and recognition stack even if you only want motion transfer, which is worth knowing before you install into an existing environment.

Editorial conclusion

EchoMimic V1 is a research release that happens to be usable: audio alone gets you a talking portrait, and the landmark conditioning path is the part with real editorial value because a wrong mouth shape can be corrected at the source instead of retrained. Start from `infer_audio2vid.py` for the plain audio case, switch to `infer_audio2vid_acc.py` for the accelerated pipeline that the update log dates to 2024.07.17, and use `infer_audio2vid_pose.py` or its `_acc` sibling when you also need upper body motion. Check the pinned `diffusers==0.24.0` and the torch range before installing, and if you want the current line of work rather than the paper version, read `echomimic_v2` and `echomimic_v3` instead of this repository.

Frequently asked questions

What hardware do you need to run EchoMimic locally?

The repository names V100 as the GPU behind its published speed figures, roughly 50 seconds for 240 frames with the accelerated pipeline, and HuggingFace hosts an A100 backed demo. Nothing in the README sets a minimum, so treat the hosted A100 space as the reference point for what someone considered comfortable and budget accordingly.

Which EchoMimic script should I use first?

Start with `infer_audio2vid.py`, the plain audio driven path released on 2024.07.09. Switch to `infer_audio2vid_acc.py` for the accelerated audio pipeline dated 2024.07.25, or to `infer_audio2vid_pose.py` and `infer_audio2vid_pose_acc.py` when you also want pose driven motion. `webgui.py` wraps the same work in a Gradio interface.

What is the difference between EchoMimicV1, V2 and V3?

V1, in this repository, animates portraits from audio with editable landmark conditioning. V2, in the separate echomimic_v2 repository, moves toward striking, simplified, semi-body human animation. V3, in echomimic_v3, is described as a 1.3B parameter model for unified multi-modal and multi-task human animation. All three come from the same Ant Group team.

Can I use EchoMimic without installing Python dependencies?

Yes. Gradio demos are hosted on HuggingFace and ModelScope, including an A100 backed version, and a ComfyUI integration by smthemex has been available since 2024.07.14. Model weights are published on HuggingFace and ModelScope, so a hosted run is a reasonable way to see the output before building the local environment.

What does editable landmark conditioning let me control?

It lets you hold selected facial landmarks in place while the rest of the face animates from audio, so a specific mouth shape or eye movement can be corrected at the source. The gallery separates Audio Driven, Landmark Driven, and Audio plus Selected Landmark Driven results, which shows the same model driven three different ways.

Official sources

  1. antgroup/echomimic on GitHub
  2. Issues
  3. License: Apache-2.0
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/antgroup-echomimic.svg)](https://hysenlabs.com/projects/antgroup-echomimic)