Model or dataset
lipku/LiveTalking avatar
lipku/LiveTalking

LiveTalking: a real-time streaming digital human engine you host yourself

Real time interactive streaming digital human

9,661 stars1,536 forksPythonApache-2.0

At a glance

What is it?
LiveTalking drives a talking avatar from text or audio and pushes it out over WebRTC, RTMP or a virtual camera. This is what the Apache-2.0 release actually runs, what the install costs, and where it stops.
Who is it for?
Adopt LiveTalking if you need a self-hosted avatar pipeline with a WebRTC endpoint and you have an NVIDIA GPU plus a prepared avatar bundle, because nothing renders without wav2lip.pth in models/ and an avatar unpacked into data/avatars/. Do not adopt it if you expect a turnkey cloud product, since the open release ships the wav2lip, musetalk and ultralight models while the commercial tier adds wav2lipls and the performance work.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 17 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What LiveTalking solves, and who ends up running it

A talking avatar demo is easy to record and hard to serve. The hard part is keeping the mouth in sync with audio that arrives at unpredictable moments, while a viewer is watching over a live transport. LiveTalking is the serving half of that problem. The README describes it as a "real time interactive streaming digital human engine" and lists the pipeline in one line: user text or audio, an optional LLM reply, TTS synthesis, lip-sync inference, then audio and video pushed out to the viewer.

The README names the audiences directly. Virtual livestream hosts that run unattended, digital customer service backed by an enterprise knowledge base, recorded or live teaching, voice assistants that call the /human endpoint, exhibition screens, and batch short-video production through /human plus /record. Those are different products sharing one runtime, and the shared runtime is what the repository actually contains: a Flask and aiortc server, a registry for pluggable modules, and a set of inference models.

The project is Apache-2.0 and the last push was on 2026-09-13, with v2.0.4 released on 2026-06-20. That matters less than the split the README draws in its own commercial table: the open release carries wav2lip, musetalk and ultralight, while the paid tier adds wav2lipls and the performance and motion work. If you are evaluating LiveTalking, you are evaluating the open model set, not the vendor's best one.

The pipeline: sessions, TTS, lip-sync inference and the output layer

The README's architecture section splits the system into four layers plus a plugin mechanism, and the data flow diagram in assets/dataflow.png is the authoritative picture. The API layer accepts /human for text and /humanaudio for an audio file. Every connection gets a unique sessionid, which is how multiple viewers are kept apart.

The logic layer is where the choices live. An LLM engine generates replies, either directly or through an OpenAI-compatible gateway, and the README gives OrcaRouter as an example with the flag --llm_provider orcarouter. TTS is modular, with EdgeTTS, GPT-SoVITS, CosyVoice and Tencent Cloud named as options. Alongside the speech, the system extracts acoustic features such as a Mel spectrogram, and those features are what the lip-sync model consumes.

The rendering layer takes those features, runs Wav2Lip or MuseTalk or another model to produce the mouth region, then smooths that region back onto the original high-definition video. That last step is the reason a full-body video can be reused: the model is not generating a person, it is replacing a mouth and compositing.

The output layer is where LiveTalking differs from a video generator. WebRTC gives the low-latency browser path, RTMP pushes to standard live platforms, and a virtual camera exposes the result as a system camera device. The plugin system is a decentralized registry in registry.py, so TTS, Avatar and Output modules can be extended without patching the core. That is a real architectural commitment, and it is also why the repository has separate tts/, streamout/ and avatars/ directories rather than one inference script.

Installing LiveTalking and driving it from the browser

The README states that the project has been tested on Ubuntu 22.04 with Python 3.12, PyTorch 2.9.1 and CUDA 12.8. The install sequence creates a conda environment, installs the matching PyTorch wheels, then the requirements file.

bash
git clone https://github.com/lipku/LiveTalking.git
conda create -n livetalking python=3.12
conda activate livetalking
pip install torch==2.9.1 torchvision==0.24.1 torchaudio==2.9.1 --index-url https://download.pytorch.org/whl/cu128
cd LiveTalking
pip install -r requirements.txt

The README warns that if nvidia-smi does not report CUDA 12.8, you should pick the matching PyTorch build from the PyTorch previous-versions page instead of copying that index URL. requirements.txt is mostly light (python-dotenv, pyyaml, numpy, scipy, einops, flask, aiortc, aiohttp_cors, soundfile, librosa, openai, edge_tts, dashscope, diffusers, accelerate). The local ASR lines for SenseVoice and FunASR are commented out and marked optional, so a default install does not pull them.

Models are not in the repository. The README points to a Quark drive and a Google Drive folder, and gives two placement steps: copy wav2lip256.pth into models/ and rename it to wav2lip.pth, then unpack wav2lip256_avatar1.tar.gz and copy the whole folder into data/avatars/. Skipping either step leaves the server without a model or an avatar to render.

Start the service with the transport and avatar selected on the command line.

bash
python app.py --transport webrtc --model wav2lip --avatar_id wav2lip256_avatar1

The README notes the server needs TCP port 8010 open and UDP ports 1-65536, which is the practical cost of WebRTC: the signaling port is fixed, the media ports are not. Then open http://serverip:8010/index.html in a browser, click the connect control, and type into the text box. Submitting text is the first real use, and what you should see is the avatar speaking the text back with synchronized lips. The same server also exposes /avatar.html for generating an avatar from an uploaded video and /admin.html for session monitoring and global configuration.

Concurrency is two different bottlenecks, and the logs are the only honest measure

The README's performance section makes a distinction that is easy to miss when sizing a deployment. Every video stream is compressed on the CPU, and compression cost rises with resolution. Every lip-sync inference runs on the GPU. The consequence is stated plainly: when nobody is speaking, concurrency is bounded by CPU; when everyone is speaking at once, it is bounded by GPU.

That means a capacity number from a single-user demo tells you almost nothing. The README gives the check to run instead. In the backend log, inferfps is the GPU inference frame rate and finalfps is the final streaming frame rate, and both need to be at least 25 for the result to count as real time. If inferfps is fine and finalfps is not, the bottleneck is downstream of the model, which usually points at CPU compression or the network path rather than the GPU.

The published figures are model and card specific. wav2lip256 is listed at 60 FPS on an RTX 3060 and 120 FPS on an RTX 3080Ti. musetalk is listed at 42 FPS on an RTX 3080Ti, 45 on an RTX 3090 and 72 on an RTX 4090. The README recommends an RTX 3060 or better for wav2lip256 and an RTX 3080Ti or better for musetalk. Those numbers are single-stream inference rates, not a promise about how many viewers a card serves once you multiply by CPU compression and bandwidth.

Where LiveTalking is the wrong tool

The most concrete limitation is the dependency on prepared assets. The quick start requires a .pth checkpoint renamed to a fixed filename and an avatar archive unpacked into a fixed directory, both downloaded from third-party drives. There is no documented CLI that trains or exports an avatar from scratch in the README; avatar creation is a web page at /avatar.html with its own API document, and the README does not describe what happens when the upload fails or how long generation takes.

Second, the GPU requirement is not optional. The README's own wording is that each lip-sync stream consumes GPU, and the recommended cards start at an RTX 3060. On a CPU-only machine there is no described path to real-time output.

Third, if what you want is a finished video file, this is the wrong layer. Batch production exists through /human plus /record, but the repository is organized around live transports and sessions. A pipeline that only needs to write MP4s is carrying WebRTC, session management and streaming code it will never use.

Fourth, the maintenance split is a real consideration. The README's commercial table lists wav2lipls and the enhancement work as commercial-tier features, so the open release is not the vendor's fastest configuration. Anyone whose product depends on the newest model should read that table as a product boundary rather than a footnote. Finally, section 8 of the README states that videos built on this project and published on Bilibili, WeChat Channels, Douyin and similar platforms must carry the LiveTalking watermark and mark. That is a distribution condition attached to the open release, and it is easy to miss because it sits at the bottom of a long document.

How it differs from OpenAvatarChat, MuseTalk and hosted avatar services

The related searches around this project mostly point at two kinds of alternative: other open digital human stacks and hosted avatar services. The difference is structural rather than a matter of feature lists.

MuseTalk appears in LiveTalking's own topic list and is one of the models LiveTalking can run. That is the key distinction: MuseTalk is a lip-sync model, and LiveTalking is the server around a model. Choosing MuseTalk gives you inference code; choosing LiveTalking gives you sessions, a TTS layer, transports and an admin page, with MuseTalk as one selectable backend via --model. If your problem is only mouth synthesis inside an existing pipeline, the model is the smaller and more appropriate dependency.

OpenAvatarChat is another full-stack avatar project that appears in the same searches. LiveTalking's distinguishing choice is the plugin registry in registry.py, which lets TTS, Avatar and Output modules be added without modifying the core, and the transport set: WebRTC, RTMP and virtual camera in one server. A stack that only serves a browser page is not interchangeable with one that can also push RTMP to a live platform.

Hosted services such as HeyGen sit at the opposite end. They remove the GPU, the model download and the port configuration entirely, and they charge for it. LiveTalking's trade is the reverse: you own the inference, you own the avatar assets, you run the GPU, and you keep the session data. For an unattended 24-hour stream or a customer service deployment where the knowledge base cannot leave your network, that trade is the reason to pick it. For a one-off marketing clip, it is a large amount of infrastructure for a small output.

Licence, upgrades and what maintenance actually costs

LiveTalking is Apache-2.0. That permits commercial use and modification, and the repository ships a LICENSE file at the top level. It does not remove the README's watermark requirement for published videos, and it does not cover the model weights, which are distributed separately from third-party drives with their own provenance. Treating the code licence as clearance for the checkpoints is a mistake; the README says nothing about the weights' terms. This is a description of what the repository states, not legal advice.

Upgrade cost is shaped by the dependency list. requirements.txt pins almost nothing: torch is installed separately with an explicit version, soundfile is pinned to 0.12.1 and websockets to 12.0, and the rest float. The README's tested combination is Ubuntu 22.04, Python 3.12, PyTorch 2.9.1 and CUDA 12.8, so an upgrade that moves any of those four should be treated as a coordinated change rather than a pip bump. The Dockerfile in the repository is a poor guide here: it targets an older CUDA 11.6.1 base image, creates a Python 3.10 environment named nerfstream, installs PyTorch 1.12.1, and its COPY paths reference directories (../nerfstream, ../python_rtmpstream) that do not match the current repository layout. The README's Docker section instead points at AutoDL and UCloud images. Anyone planning to build from that Dockerfile should expect to rewrite it.

Configuration is split between config.py, config.yaml and a .env file. .env.example lists the credentials the optional integrations expect: TENCENT_APPID, TENCENT_SECRET_KEY, TENCENT_SECRET_ID, DASHSCOPE_API_KEY, ORCAROUTER_API_KEY, DOUBAO_API_KEY, AZURE_SPEECH_KEY and AZURE_TTS_ENDPOINT. If you use EdgeTTS only, none of those are needed. The upgrade surface is therefore three things: the PyTorch and CUDA pairing, the model checkpoints in models/, and whichever TTS and LLM providers you have wired in.

Editorial conclusion

Adopt LiveTalking if you need a self-hosted avatar pipeline with a WebRTC endpoint and you have an NVIDIA GPU plus a prepared avatar bundle, because nothing renders without wav2lip.pth in models/ and an avatar unpacked into data/avatars/. Do not adopt it if you expect a turnkey cloud product, since the open release ships the wav2lip, musetalk and ultralight models while the commercial tier adds wav2lipls and the performance work. Verify first that your GPU holds the inferfps and finalfps numbers above 25 for your chosen model, and confirm with the maintainer that the watermark requirement in section 8 of the README applies to your distribution.

Frequently asked questions

What is LiveTalking and what does it do?

LiveTalking is a real-time interactive streaming digital human engine written in Python. It takes text or audio, optionally generates a reply through an LLM, synthesizes speech with a TTS engine, runs lip-sync inference, and pushes the result out over WebRTC, RTMP or a virtual camera.

How do I install and run LiveTalking?

Clone the repository, create a Python 3.12 conda environment, install the matching PyTorch wheels and requirements.txt, then download wav2lip256.pth into models/ renamed as wav2lip.pth and unpack wav2lip256_avatar1.tar.gz into data/avatars/. Start it with python app.py --transport webrtc --model wav2lip --avatar_id wav2lip256_avatar1 and open port 8010.

What GPU does LiveTalking need?

The README recommends an RTX 3060 or better for wav2lip256 and an RTX 3080Ti or better for musetalk, and states that each lip-sync stream consumes GPU while video compression consumes CPU. It gives inferfps and finalfps in the backend log as the check, and both need to be at least 25 for real-time output.

Can I run LiveTalking with Docker?

The README's Docker section points to prepared AutoDL and UCloud images rather than giving build steps. The Dockerfile in the repository targets an older CUDA 11.6.1 base with PyTorch 1.12.1 and references directories that do not match the current layout, so it does not reflect the documented install.

Does LiveTalking work with MuseTalk?

Yes. The README lists ernerf, musetalk, wav2lip and Ultralight-Digital-Human as supported digital human models, and the model is selected on the command line with the --model flag. The README's performance table lists musetalk at 42 FPS on an RTX 3080Ti and 72 FPS on an RTX 4090.

Official sources

  1. License: Apache-2.0
  2. lipku/LiveTalking on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/lipku-livetalking.svg)](https://hysenlabs.com/projects/lipku-livetalking)