Open-source project
MeiGen-AI/MultiTalk avatar
MeiGen-AI/MultiTalk

MeiGen-AI/MultiTalk: audio-driven multi-person conversational video generation

[NeurIPS 2025] Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

3,002 stars496 forksPythonApache-2.0

At a glance

What is it?
MultiTalk is a Python research framework that turns a reference image, a prompt and multi-stream audio into a talking-head video with lip motion aligned to each speaker. It is a NeurIPS 2025 release built on Wan, and the README documents multi-GPU inference, low-VRAM 480p on a single RTX 4090, and LoRA acceleration rather than a hosted service.
Who is it for?
Adopt MultiTalk if you have a CUDA GPU, a Python environment you can pin to numpy<2, and a reason to generate multi-person conversation, singing or cartoon footage from separate audio streams. Skip it if you need a hosted editor, a ComfyUI graph that the repository itself ships, or a service with documented rollback and versioning.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 135 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap MultiTalk fills: several speakers, one video, separate audio tracks

Most audio-driven avatar models assume one face and one audio track. MultiTalk is aimed at the case where a clip contains more than one person and each person has their own audio stream, which the README calls multi-stream audio input. The stated output is a video whose characters interact according to a text prompt, with lip motion aligned to the audio. The repository lists single-person support as well, and the examples directory is split accordingly: examples/single/ and examples/multi/ hold JSON job files, with names like examples/single_example_1.json, examples/single_example_tts_1.json, examples/multitalk_example_1.json and examples/multitalk_example_tts_1.json.

The audience is narrower than the demo videos suggest. This is a research codebase from a NeurIPS 2025 paper, published under Apache-2.0, with a technique report on arXiv and weights on Hugging Face. It is for engineers and researchers who are comfortable running a diffusion pipeline from a shell, not for someone looking for a web app. The README does mention a Gradio interface, and app.py sits at the top level, but the primary path documented for generation is the script.

What the pipeline actually takes in and puts out

The input contract is three things: a reference image, a prompt, and audio. The audio can be multi-stream, meaning one track per speaker rather than a single mixed file. The README also notes support for input audios with TTS, and the repository contains a kokoro/ directory alongside a dashscope dependency in requirements.txt, which points to text-to-speech being generated inside the pipeline rather than only supplied from outside.

The output side is where the concrete numbers are. The README claims 480p and 720p output at arbitrary aspect ratios, and video generation up to 15 seconds. Those are the two constraints worth internalising before you plan anything: resolution is flexible, duration is not unlimited. Long video generation in this project means 15 seconds, not an hour.

Underneath, the framework is built on Wan, which is why there is a wan/ directory at the top level next to src/ and weights/. The README describes several acceleration and quality mechanisms layered on top: multi-GPU inference, teacache acceleration, APG, low-VRAM inference, INT8 quantization through optimum-quanto, SageAttention2.2, and LoRA options including FusionX and lightx2v. Two of those deserve a plain reading. APG is described as alleviating color error accumulation in long video generation, which is an admission that colour drift is a real failure mode in this class of model. TeaCache is described as increasing speed by approximately 2 to 3x. The LoRA path is described as requiring only 4 to 8 steps, and the July 2025 note says the CFG strategy was updated to 2 NFE per step for the FusionX LoRA. Those are the levers you pull when a run is too slow, and they trade against each other.

Installing MultiTalk and running a first generation

The repository does not ship a packaged installer. The dependency list lives in requirements.txt, and the install path is the usual one for a research repo: create an environment, install the requirements, then put model weights in the weights/ directory. The requirements file pins numpy below 2, which matters because a stray numpy 2.x in the same environment will break the stack.

Start by installing the declared dependencies from the repository root:

bash
pip install -r requirements.txt

What you should see is pip resolving opencv-python, diffusers, transformers, tokenizers, accelerate, gradio, xfuser, optimum-quanto and the rest. Note the exact pins: numpy>=1.23.5,<2 and optimum-quanto==0.2.6. Do not let a resolver upgrade either one.

Weights are not in the repository. The README points to the Hugging Face model page for MeiGen-MultiTalk, and the weights/ directory is where the code expects them. Check the model card for the file list rather than guessing filenames.

The generation entry point is generate_multitalk.py at the top level, and the examples/ directory holds JSON job files that describe a run. The README does not print a full invocation for the script, so the honest starting point is to inspect the script's arguments and mirror one of the example JSON files:

bash
python generate_multitalk.py

Compare the flags it prints against examples/multitalk_example_1.json for a multi-person job, or examples/single_example_1.json for one speaker. The TTS variants, examples/multitalk_example_tts_1.json and examples/single_example_tts_1.json, are the ones to look at if you want the pipeline to synthesise speech instead of reading your own audio files.

If you prefer a browser, app.py is a Gradio application and gradio>=5.0.0 is in requirements.txt. The README lists Gradio among the supported surfaces but does not document its launch command, so treat app.py as the reference.

For a constrained machine, the README states that low-VRAM inference enables 480P video generation on a single RTX 4090, and that multi-GPU inference is supported. Those two paths are the ones to reach for before you buy hardware.

Where MultiTalk stops being the right tool

The 15-second ceiling is the first hard boundary. If your use case is dubbing a long interview or generating a sustained scene, the README's own feature list does not promise that, and the APG note about colour error accumulation in long video generation tells you the authors hit the problem at the edge of the supported range.

The second boundary is the release cadence. The repository's own news feed shows the team moving on: InfiniteTalk in August 2025 for video dubbing, LongCat-Video-Avatar in December 2025, and LongCat-Video-Avatar-1.5 in May 2026, which the README describes as replacing Wav2Vec2 with Whisper-Large for more accurate lip synchronization, adding step distillation to 8 steps, and supporting both single-stream and multi-stream audio. That last point matters directly: the successor claims the same multi-stream capability MultiTalk introduced, plus stylised domains and longer video. If you are starting fresh today, MultiTalk is not the newest thing from the same group.

The third boundary is operational. There are no retrieved releases, so there is no versioned artefact to pin and no changelog to diff. Upgrading means pulling main and re-reading the news list. The README does not document rollback, and it does not document a supported dependency matrix beyond requirements.txt. On a project where a single numpy major version can break the environment, that is a real cost.

Finally, the README does not document a dataset release, despite dataset being a common search term around this project. If your goal is to train on the same data, the repository layout does not show it.

MultiTalk compared with InfiniteTalk and LongCat-Video-Avatar

The most useful comparison is inside the same family, because the difference is one of task shape rather than quality. InfiniteTalk, released in August 2025, is described in the README as a paradigm for video dubbing: infinite-length video-to-video generation and image-to-video generation. MultiTalk, by contrast, generates a video from a reference image and multi-stream audio, capped at 15 seconds. If your input is already a video and your problem is replacing or adding speech, InfiniteTalk matches the input format. If your input is a still image and several audio tracks, MultiTalk matches it.

LongCat-Video-Avatar-1.5 is the closer competitor and the harder one to argue against. Per the README it swaps Wav2Vec2 for Whisper-Large to improve lip synchronisation, distils inference to 8 steps, and covers single-stream and multi-stream audio. It also claims generalisation to anime, animals and complex real-world conditions. MultiTalk's own feature list already mentions cartoon characters and singing, so the overlap is substantial. The distinguishing factor for a new project is that LongCat-Video-Avatar-1.5 is the later release from the same group, with a stated accuracy improvement on the component that matters most for a talking-head model.

Outside the family, the realistic alternative is a general video diffusion stack that you condition on audio yourself. That is a much larger engineering job and the README gives no reason to think it would be cheaper, so it is only worth naming as the option you take when you need control the framework does not expose.

Licence, maintenance and what an upgrade costs you

The repository is Apache-2.0, with LICENSE.txt at the top level. That is a permissive licence, and it is the most predictable part of the project. It does not resolve the separate question of the model weights: those live on Hugging Face at MeiGen-AI/MeiGen-MultiTalk, and their terms are stated there, not in this repository. If you plan to ship generated output commercially, the weights page is the document to read, not LICENSE.txt. Nothing here is legal advice.

Maintenance is the weaker signal. The last push was on 2026-05-22, and that push coincides with the LongCat-Video-Avatar-1.5 announcement in the news list. The repository is not archived, so it is not formally retired, but the visible activity is the team pointing users at a newer framework rather than expanding this one. There are no retrieved releases, which means there is no tag to upgrade between.

The practical upgrade cost follows from that. You cannot pin a version, so reproducibility depends on your own snapshot of the checkout and of the weight files. requirements.txt gives you floors and a few exact pins, not a lockfile. Expect to re-test after any pull, particularly around numpy, optimum-quanto and the LoRA acceleration paths, since the README records that the CFG strategy for FusionX LoRA changed after the initial release.

Running MultiTalk on a single GPU versus a cluster

The README gives two explicit operating points and they imply different workflows. The low-VRAM path is 480P on a single RTX 4090. The other is multi-GPU inference through xfuser, which is in requirements.txt at version 0.4.1 or later. Choosing between them is mostly a question of whether you are iterating on prompts or producing final clips.

On one GPU, the LoRA acceleration options are what make iteration tolerable. The README states that the lightx2v and FusionX LoRAs require only 4 to 8 steps, and that TeaCache increases speed by approximately 2 to 3x. Stacking those is the difference between a usable loop and a slow one. The cost is quality: a distilled LoRA at 4 to 8 steps is not the same model as the full-step path, and the README does not publish a comparison between them.

On multiple GPUs, xfuser handles the split. The README does not document the launch configuration for it, so the argument names are something you read out of generate_multitalk.py rather than from the documentation. That is a recurring pattern in this repository: the capability is announced, the invocation is left to the code.

INT8 quantization through optimum-quanto and SageAttention2.2 are the other two memory and speed levers, both added in July 2025. optimum-quanto is pinned to exactly 0.2.6 in requirements.txt, which suggests the integration is version-sensitive. Treat that pin as load-bearing.

Editorial conclusion

Adopt MultiTalk if you have a CUDA GPU, a Python environment you can pin to numpy<2, and a reason to generate multi-person conversation, singing or cartoon footage from separate audio streams. Skip it if you need a hosted editor, a ComfyUI graph that the repository itself ships, or a service with documented rollback and versioning. Before you commit, check the weights directory and the Hugging Face model page for the checkpoint files, confirm your GPU memory against the low-VRAM path the README describes for 480P on a single RTX 4090, and read generate_multitalk.py to see which flags your build actually accepts.

Frequently asked questions

How do I use MeiGen MultiTalk?

Install the dependencies from requirements.txt, place the weights from the MeiGen-MultiTalk Hugging Face page in the weights/ directory, then run generate_multitalk.py using one of the JSON files under examples/ as a model. The README also lists a Gradio interface, with app.py at the top level.

What is MultiTalk?

MultiTalk is an audio-driven multi-person conversational video generation framework released under Apache-2.0 and presented as a NeurIPS 2025 paper. Given multi-stream audio, a reference image and a prompt, it generates a video with interactions following the prompt and lip motions aligned with the audio.

What is the difference between MultiTalk and InfiniteTalk?

InfiniteTalk, released in August 2025, is described as a paradigm for video dubbing with infinite-length video-to-video and image-to-video generation. MultiTalk generates from a reference image and multi-stream audio, with video generation supported up to 15 seconds.

How do I use MultiTalk AI?

The documented path is the Python codebase: install requirements.txt, obtain the weights from Hugging Face, and run generate_multitalk.py against an example JSON job file. The README does not describe a hosted online version of MultiTalk.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. MeiGen-AI/MultiTalk on GitHub
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/meigen-ai-multitalk.svg)](https://hysenlabs.com/projects/meigen-ai-multitalk)