# LiveAvatar: training code is the one empty box on the list

> An ECCV 2026 Spotlight paper implementation that generates an audio-driven talking avatar in real time, as a LoRA on someone else's 14B base model. Inference, checkpoints, a web interface and single-GPU support are all published; training is not. The news log has been edited, with one performance entry commented out and a later one claiming a larger number.

**Alibaba-Quark/LiveAvatar** — [ECCV 2026 Spotlight] Implementation of "Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length"

- Repository: https://github.com/Alibaba-Quark/LiveAvatar
- Stars: 2,448 · Forks: 286
- Language: Python
- License: Apache-2.0
- Published: 2026-09-28 · Updated: 2026-09-28 · Language: en
- Canonical page: https://hysenlabs.com/projects/alibaba-quark-liveavatar

## Training code is the one empty box, and three others sit beside it

The to-do list is the clearest statement of scope in the project, and it is split into a completed early phase and a later phase with four items still open. In the completed group are the paper, a demo website, checkpoints on a model hub, a web interface, experimental real-time streaming on the target hardware, distribution-matching distillation down to four sampling steps, and a form of pipeline parallelism. The later group has four unchecked boxes: a user interface for easy streaming interaction, text-to-speech integration, training code, and a version 1.2. Everything else in that group is marked done, including single-GPU inference, multi-character support, two stages of inference acceleration, and an audio codec integration. So the project can generate and can be served and can be made faster, and it cannot be trained. There is no code, no configuration and no recipe for producing the adapter it ships.

## A performance entry has been commented out and the surviving one claims a bigger number

The news log has been edited rather than appended to. One entry, dated in mid-January, is present as a comment: it reports inference speed boosted to one and a half times peak and two times average, reaching a stable thirty or more frames per second on the target hardware. The entry above it, from a day later, reports low-precision quantization, further compilation, and a library attention implementation, claiming two and a half times peak and three times average, and a stable forty-five or more frames per second on the same multi-card setup. The later entry also adds that the fixes surpass the teacher model on qualitative metrics, which is a stronger claim than the speed one and carries no measurement. Nothing on the page explains the deletion, and the surviving entry's numbers are roughly double the commented-out ones for the same hardware.

## The hardware floor moved twice, and the two figures describe different paths

A December entry announces single-GPU inference and states the requirement directly: one GPU with 80GB of video memory is enough, and the point of the release is that you no longer need five of the target data centre cards. A January entry then says low-precision quantization enables inference on 48GB cards. So the documented floor dropped by a third, and the two statements are describing the same goal, an avatar running in real time on one card, by two different routes. Neither entry says what quality is lost at the lower figure, and the January entry's claim of surpassing the teacher model appears in the same paragraph as the reduced-memory path, so it is not clear whether the quality claim is about the low-precision configuration or the full-precision one. For anyone budgeting hardware, that is the sentence to ask about.

## You download two sets of weights, and one belongs to another organisation

The model table has two rows. The first is a 14B base model hosted by a different organisation, described as the base model. The second is the project's own adapter, described as their adapter model, hosted under the project's own account. So the released artefact is a low-rank adapter over someone else's base checkpoint, and a working setup needs both downloads placed in one directory. The instructions add a note inline for readers in mainland China, telling them to set an environment variable pointing at a mirror before running the hub's command-line client, which is the only distribution-specific instruction in an otherwise region-neutral guide. The example assets follow the same pairing: an anchor image identifies the avatar, and each character is a portrait plus a matching audio file, with a handful of still images for other demos.

## Two caps, one exact pin and one duplicated line among forty-five requirements

The requirements file is forty-odd lines long and mostly open-ended, with a few deliberate exceptions. The transformer library is bounded at both ends, from one version up to and including another, which is unusual for a research repository where the whole point is to track a moving target. The adapter library is pinned to one exact version. The array library is capped below its next major. The computer vision library appears twice, once with a minimum version and once bare, so the second line can only be meaningful by accident. Beyond those, the list pulls in a segmentation model, a face analysis library, a distributed training library, a speech codec library, a conformer implementation, a hydra-style configuration system, a training framework, a world vocoder, a video reader, an inference runtime, a speech recognition model, a tensor serialiser, and both a hub client and a mirror client. Whatever this project does, it does it through other people's packages.

## The install pins torch to one version while the requirements file accepts two years of them

Read the two files together and they disagree about the framework. The install step pins the tensor library and its vision sibling to exact versions and asks for them from a wheel index named for one CUDA release:

```bash
pip install torch==2.8.0 torchvision==0.23.0 --index-url https://download.pytorch.org/whl/cu128
```

while the requirements file asks only for anything from a much older minimum with no ceiling. So a user who follows the documented install gets a newer framework than the requirements file would have chosen, and a user who runs the requirements file first and the install step second gets the newer one anyway. There is a related mismatch one step earlier: the optional CUDA step installs a toolkit version whose number is lower than the CUDA release in the wheel index used two steps later. Neither is fatal, but together they mean the page does not describe one coherent environment.

## The attention library you need depends on your card, and the wheel path encodes your torch version

The third install step has a branch, and the branch is about hardware generation rather than about convenience. If you are on a data centre card from the Hopper generation, the page recommends a third-generation attention library and gives a wheel index URL whose path segment encodes both the CUDA release and the exact tensor library version it was built against. Otherwise it tells you to install a second-generation library at a pinned version with build isolation turned off. The isolation flag matters, because without it the extension compiles against whatever is already in the environment rather than a throwaway one, which is the only way the result matches the installed framework. The wheel index is pinned to a specific torch version, so anyone who changed the framework version first will not find wheels at that path.

## The release promise is still sitting in a comment at the top of the page

Directly under the title there is a line of HTML that renders as nothing. It is a commented-out heading saying the code will be open source in early December, which matches a December news entry announcing the same commitment, and the repository now contains inference code, so the promise has been kept and the comment left behind. Two other formatting choices are worth knowing about before you read the page as a record. Three of the badges are link elements with no text, so the build status, the paper and the hub links are invisible except as clickable areas. And the model identity is established entirely through a chain of five such empty links, meaning nothing on the page tells you at a glance which paper or which checkpoint you are looking at. The visible content is a demo link, a project page link, and a summary paragraph.

## Conclusion

LiveAvatar is worth evaluating if you want to see and run a streaming avatar system rather than train one, and you have the hardware for it. Three things decide whether you can. Work out your memory floor first, because the project has moved that number twice and the two figures describe different paths: single-GPU inference is documented as needing an 80GB card, while a later update says low-precision quantization brings it down to 48GB. Accept that this is a LoRA on a base checkpoint you download from a different organisation, so you are pulling two sets of weights rather than one. And read the install steps in order rather than combining them, because the page installs one CUDA toolkit version and then asks for a PyTorch build from a different CUDA wheel index, and the attention library you need depends on whether your card is Hopper. If your question is whether you can fine-tune it, the answer on the project's own list is that you cannot yet.

## FAQ

### what is live avatar

Live Avatar is an algorithm and system co-designed framework for real-time, streaming, infinite-length interactive avatar video generation, published as an ECCV 2026 Spotlight paper. It is powered by a 14B-parameter diffusion model with 4-step sampling, reports 45 frames per second on multi-card H800 hardware, and supports block-wise autoregressive processing for videos of 10,000 seconds or more.

### live avatar vs heygen

This repository does not make that comparison. What it states about itself is that it is a research implementation with inference code only, since training code appears as an unchecked item on its own list, and that a working setup needs a 14B base checkpoint hosted by another organisation plus the project's own adapter.

### Is there a free AI avatar?

That depends on which one you mean, and this repository says nothing about hosted services. What it describes is self-hosted research code: the checkpoints are published, but running it in real time needs either a single GPU with 80GB of memory, or after later optimisation a 48GB card, or several data centre GPUs for the full path.

### Can I train or fine-tune LiveAvatar?

Not with what has been published. The project's own list has training code unchecked, alongside text-to-speech integration and an interface for easy streaming interaction, while the paper, the demo site, the checkpoints, the web interface, single-GPU inference and two stages of acceleration are all marked done.

### What do I need installed to run LiveAvatar?

A conda environment on Python 3.10, PyTorch and its vision sibling at exact versions from a wheel index for one CUDA release, either the third-generation attention library for data centre cards or version 2.8.3 built without isolation elsewhere, the requirements file, and the ffmpeg binary from the system package manager. The checkpoints are then downloaded into one directory.

## Sources

- [Alibaba-Quark/LiveAvatar on GitHub](https://github.com/Alibaba-Quark/LiveAvatar)
- [Issues](https://github.com/Alibaba-Quark/LiveAvatar/issues)
- [License: Apache-2.0](https://github.com/Alibaba-Quark/LiveAvatar/blob/main/LICENSE)
- [README](https://github.com/Alibaba-Quark/LiveAvatar/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/alibaba-quark-liveavatar
