# MultiTalk's newest news item announces its successor under a different organisation, and one LoRA is spelled two ways

> A NeurIPS 2025 framework that generates multi-person conversational video from multi-stream audio, a reference image and a prompt, at 480p or 720p and up to fifteen seconds. The repository is an inference tree rather than a package, with model weights committed and no install command anywhere in the README.

**MeiGen-AI/MultiTalk** — [NeurIPS 2025] Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

- Repository: https://github.com/MeiGen-AI/MultiTalk
- Website: https://meigen-ai.github.io/multi-talk/
- Stars: 3,002 · Forks: 496
- Language: Python
- License: Apache-2.0
- Published: 2026-09-24 · Updated: 2026-09-24 · Language: en
- Canonical page: https://hysenlabs.com/projects/meigen-ai-multitalk

## Multi-stream audio is the input shape, and fifteen seconds is the ceiling

The model is described by its conditioning, and the conditioning is the interesting part. It takes a multi-stream audio input, a reference image and a prompt, and generates a video containing interactions that follow the prompt, with lip motions kept consistent with the audio.

Multi-stream is the load-bearing word. Two streams is two speakers, which is what separates this from a single-person lip-sync model: the problem is not one mouth moving correctly but several people taking turns in one frame. The feature list states the scope in both directions, with support for single and multi-person generation.

The other four claims are bounded, which makes them checkable. Output is 480p or 720p at arbitrary aspect ratios, so the resolution is a choice rather than a fixed target. Long video generation is supported up to fifteen seconds, which is a hard cap rather than a soft one. Generalisation covers cartoon characters as well as realistic ones, and singing as well as speech, so the lip motion model is not restricted to conversational audio.

Character control is by prompt, and the README describes it as directing virtual humans through prompts. There is no separate control interface described anywhere in the document, so prompt text is the whole of the control surface.

## The last commit follows a news item that points somewhere else entirely

The news section is dated and long, and its final entry is the most consequential one for anyone deciding what to build on.

Dated May 21, 2026, it announces LongCat-Video-Avatar-1.5, described as an upgraded open-source framework for audio-driven human video generation. The listed changes are concrete: the speech encoder is replaced, moving from one acoustic model to Whisper-Large for more accurate lip synchronisation; inference is cut to eight steps through step distillation; the model is generalised to stylised domains including anime and animals; and both single-stream and multi-stream audio inputs are supported.

The detail to notice is where the artefacts live. The code link for that release is under a `meituan-longcat` organisation and the weights are under a HuggingFace organisation of the same name, while this repository and its weights are under `MeiGen-AI`. Two organisation names across one line of work.

The timeline makes the timing unambiguous. That news entry is dated May 21, 2026 and the last push to this repository is dated May 22, 2026, one day later. So the final activity in this tree is the announcement of its successor.

The earlier entry in the same lineage is dated December 16, 2025 and covers the first version of that model, with the same two organisations again split between the announcement repository and the code and weights.

## One LoRA is spelled FusionX and FusioniX ten days apart

Two news entries from July 2025 describe the same acceleration route and spell its name differently.

The earlier entry, dated July 1, says MultiTalk supports input audio produced by text to speech, together with FusioniX and lightx2v LoRA acceleration, which it says requires only four to eight steps, and a Gradio interface. The later entry, dated July 11, says MultiTalk supports INT8 quantization and a named attention implementation, and updates the CFG strategy to two NFE per step for FusionX LoRA.

So across ten days the same asset is written `FusioniX` once and `FusionX` once. Both entries link to files on a HuggingFace user account whose path contains the lowercase form, which suggests the lowercase spelling is the one in the filenames.

The two entries are describing different things, though, and it is worth keeping them apart. The July 1 entry is about LoRA based acceleration and step count. The July 11 entry is about quantization, an attention implementation, and a classifier-free guidance schedule change for a particular LoRA.

Taken with the July 1 entry, the practical shape of the acceleration story is that two different LoRAs are both described as reaching usable quality in four to eight steps, one of them alongside text-to-speech audio input. Neither entry states what the baseline step count was, so the reduction is not quantified.

## TeaCache is the speed claim, and 8 GB of VRAM comes from someone else's integration

The June 14, 2025 entry is the one with numbers in it. It announces MultiTalk with multi-GPU inference, a caching acceleration, a scheduler, and low-VRAM inference, and states that low-VRAM mode enables 480p video generation on a single RTX 4090.

The caching method is named and quantified: it is described as increasing speed by approximately two to three times. The scheduler is a linked paper, and its stated purpose is to alleviate colour error accumulation in long video generation, which is a specific failure mode rather than a general quality claim.

The remaining entries supply the origin story. Weights and inference code were released on June 9, 2025, four months after the project page and technique report on May 29.

Then there is the Community Works section, which credits a third-party project for making MultiTalk run on very low VRAM hardware, stated as 8 GB, and for combining it with the capabilities of another project. So the headline memory number in the README, a single 4090, is the project's own claim, and the 8 GB figure arrives through someone else's integration rather than from this repository.

That distinction matters for planning, since an integration maintained elsewhere moves on its own schedule.

## The examples split single from multi, and the text to speech path is a local model

The examples directory is organised by how many people are speaking, which is the clearest signal of what the model is for. There is a `multi/` subdirectory and a `single/` one, and six configuration files sit alongside them.

The multi-person set has three numbered examples plus a text-to-speech variant. The single-person set has one numbered example plus a text-to-speech variant. So the multi case is the one with more coverage, and the text-to-speech entry point is documented for both.

Those configurations are JSON files. A run is therefore driven by a JSON document describing the audio streams, the reference image and the prompt, which is the shape you would want for batch work anyway.

The text-to-speech support has a local component. A `kokoro/` directory sits at the repository root, which is a speech synthesis model shipped in the tree rather than called over a network, so the July 1 entry's support for text-to-speech input audio is implemented by running a model you already have.

There is a cloud dependency as well: the requirements include a DashScope client, which is a hosted model service. So the tree carries both a local synthesis model and a hosted service client, and the README does not say which one the examples use.

## One dependency is capped below numpy 2 and one is pinned to an exact release

The requirements file has eighteen entries and two of them are pinned more tightly than the rest.

The first is numpy, with a floor and a ceiling: at least 1.23.5 and below 2. Every other entry in the file is floor-only or unpinned, so numpy is the single line in this project that declares which major series it belongs to, and it declares the older one.

The second is the quantization library, pinned to an exact release. That entry is the same library named in the July 11, 2025 news item as the route for INT8 quantization, so the quantized path and one specific version of that library are the same decision.

The rest of the list explains what the pipeline does. A diffusion library and a transformers stack with a tokenizer floor and an accelerate floor for model loading. A video-generation parallelisation library, which is the multi-GPU path from the June 14 entry. A loudness normalisation library, which is the audio correctness dependency for lip synchronisation. Gradio at a floor of 5.0.0, which is the interface behind the demo application. A hosted service client. And a familiar diffusion supporting cast: OpenCV, imageio with its ffmpeg companion, scikit-image, ftfy, easydict and loguru.

Two of those, OpenCV and imageio-ffmpeg, exist to read and write video, and the image processing library is there for the frame crops the pipeline needs. That is a research inference environment rather than a library dependency set, and it is why eighteen lines is a reasonable size for it.

## No package metadata, no install command, and model weights in the tree

This repository is an inference tree rather than a distributable package.

There is no project metadata file and no setup script anywhere in the top level. The licence file is named with a `.txt` suffix rather than the bare name used by most projects. What is at the root instead is an application file for the demo, a generation script, a source directory, an assets directory, the requirements file, a speech model directory, a directory named for the video backbone the work builds on, and a weights directory.

That last one is the item to note. Model weights are committed in the repository tree as well as published on HuggingFace, so a clone carries the model and a checkout is not a lightweight operation.

The README matches that shape. It has no installation section, no requirements walkthrough and no usage instructions. What it does have is a feature list, a dated news log, a project page, a paper, a model card, a community section and a table of nine embedded videos arranged in a grid.

The project page and the paper are the routes for anything the README omits. The paper identifier is 2505.22647, released as a technique report on May 29, 2025, five months before the weights, and the venue noted in both the title and the repository description is NeurIPS 2025.

There are no GitHub releases, so the repository history is the only version record.

## Conclusion

MultiTalk fits research on talking-head and multi-speaker generation where the conditioning is already a set of separate audio streams, since that is the shape the model is built around rather than an option. Three things to check before you plan around it. The repository is not a package: there is no project metadata file and no install command in the README, so you are cloning a tree, building an environment from an eighteen line requirements file, and pointing it at committed weights. That file pins numpy below version 2 and one quantization library to an exact release, so the environment is narrower than the rest of the dependency list suggests. And the newest news entry, dated one day before the last commit, points at a successor framework whose code and weights sit under a different organisation entirely, which is the signal about where this line of work now lives.

## FAQ

### multitalk vs infinite talk

Both are released by the same organisation and they are different tasks. MultiTalk is audio-driven multi-person conversational video generation from multi-stream audio, a reference image and a prompt, at 480p or 720p and up to fifteen seconds. InfiniteTalk, announced August 19, 2025, is described as a new paradigm for video dubbing, supporting infinite-length video-to-video and image-to-video generation, with models, code, a Gradio interface and ComfyUI released.

### how to use multitalk ai

The README documents no installation or run command. The tree names the entry points: a generation script for inference, an application file for the Gradio demo, a source directory, a weights directory, and JSON example configurations under examples/ split into single and multi person sets. Installation, environment setup and invocation are not covered in the document.

### What resolution and video length does MultiTalk support?

The feature list states 480p and 720p output at arbitrary aspect ratios, and long video generation up to 15 seconds. The June 14, 2025 news entry adds that low-VRAM inference enables 480p generation on a single RTX 4090, and the community section credits a third-party integration for running it on 8 GB of VRAM.

### what is multi talk

The phrase is ambiguous in general use. In this repository MultiTalk is the name of a video generation framework presented at NeurIPS 2025, described as audio-driven multi-person conversational video generation, given multi-stream audio, a reference image and a prompt, and producing a video whose lip motions follow the audio.

## Sources

- [Issues](https://github.com/MeiGen-AI/MultiTalk/issues)
- [License: Apache-2.0](https://github.com/MeiGen-AI/MultiTalk/blob/main/LICENSE)
- [MeiGen-AI/MultiTalk on GitHub](https://github.com/MeiGen-AI/MultiTalk)
- [Project website](https://meigen-ai.github.io/multi-talk/)
- [README](https://github.com/MeiGen-AI/MultiTalk/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/meigen-ai-multitalk
