Model or dataset
PunithVT/ai-avatar-system avatar
PunithVT/ai-avatar-system

ai-avatar-system: a talking-head stack whose own files admit to three rough edges

🎭 AI Avatar / digital human platform — upload a photo, clone a voice, talk to any face in real time with lip-sync video. Open-source, self-hosted. Claude · Whisper · Chatterbox · MuseTalk.

519 stars129 forksPythonMIT

At a glance

What is it?
A self-hosted avatar platform that streams speech recognition, an LLM, text-to-speech and lip-sync into a browser over WebSocket. The interesting reading is in the inconsistencies: three different clone lengths, a data-flow diagram that contradicts the streaming claim, and a compose file carrying a comment about its own past env bug.
Who is it for?
This is a full-stack talking-head application rather than a model, and the parts worth reading first are the ones the project's own files admit to. The voice clone length appears as five seconds, ten seconds and ten to sixty seconds in three different places, so treat any of them as approximate.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 18 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The voice clone length is stated three times with three different numbers

How much audio a voice clone needs is the first thing anyone evaluating this would ask, and the page answers it three different ways.

The introduction says clone a voice from a five-second audio clip. The summary of what makes the project different says ten seconds of audio is all you need. The features table says record ten to sixty seconds for zero-shot cloning through Chatterbox Multilingual.

Five, ten, and a range starting at ten. The third form is compatible with the second, since a range that begins at ten contains ten. The first is not compatible with either.

The engine is named consistently in all three, which at least means the disagreement is about duration rather than about which model does the cloning.

This matters more than a documentation nit because the shorter figure is the one in the sentence a first-time reader is most likely to act on. Someone who uploads five seconds on the strength of the introduction and gets poor output has no way to know that the page elsewhere says ten, let alone that a third place suggests recording closer to a minute.

Everything else in the feature set is stated with a number attached, which makes this one the outlier rather than part of a pattern of loose phrasing.

The data-flow diagram waits for a complete response before the first chunk

The headline feature is token streaming: the model streams live tokens while text-to-speech and lip-sync run per sentence, and the first video chunk plays before the model finishes its reply. The one-line pipeline at the top puts a figure on it, under two to four seconds to first video chunk on an AWS GPU.

The data-flow diagram shows something different. The language model step produces full response text. Only after that does the flow split into sentences, and only then does each sentence go through the engine and out over the WebSocket to the browser.

Full text first, then sentence splitting, is a batch step. A stream that emits complete text before the first sentence is split cannot begin video work on sentence one, because sentence one is not known until the model stops.

So the diagram and the claim disagree about where the first video chunk comes from. Either the diagram is aspirational and simplified, or the per-sentence streaming happens later in the stack than the diagram shows. The project structure block does describe a WebSocket handler with sentence streaming, so the capability exists somewhere below the layer this diagram draws.

The diagram also stops partway through the architecture picture above it, so the layer where that would have to happen is the part that is not drawn.

The pipeline diagram says XTTS, the feature list says Chatterbox

Two different text-to-speech engines appear in the two diagrams, and neither matches the other.

The header pipeline names Chatterbox as the text-to-speech stage. The features table lists Chatterbox as the voice cloning engine and describes a fallback chain of Chatterbox in a separate optional virtual environment, then edge-tts with its free neural voices, then gTTS, ending with a promise of never being silent.

The per-sentence data-flow diagram names XTTS at every sentence step instead.

XTTS does not appear anywhere in the features table, and Chatterbox does not appear in the data-flow diagram. The two diagrams were evidently written at different times against different engines.

The fallback chain adds a third consideration. Chatterbox is described as optional and isolated in its own virtual environment, which is what lets a missing or broken dependency fall through to a cloud service rather than failing. So a deployment without Chatterbox installed silently changes which synthesis engine produces the voice, and the diagrams disagree about what that engine is.

For a project whose central claim is latency, which engine runs is not a cosmetic detail, since three neural speech engines have different cold-start and per-sentence costs.

The compose file documents its own past bug: 13 keys, 56 settings dropped, no error

The backend service in the development compose file carries a long comment explaining why it loads the whole environment file. It is the most informative thing in the repository.

The comment says the backend previously used an explicit allowlist of thirteen keys, the way the production compose file already does. That allowlist silently dropped fifty-six of the sixty-nine settings the application defines. Named casualties include the avatar engine setting, the face restore switch, the language model provider, the CORS origins, the Whisper model and the text-to-speech provider.

The failure mode is the important part. A setting changed in the environment file simply had no effect in the container, with no error to explain why. And since both the readme and the setup guide begin by telling you to copy the example environment file, the file was definitely present and definitely being edited by anyone following the instructions.

Thirteen plus fifty-six is sixty-nine, so the arithmetic in the comment is consistent, which suggests the number was counted rather than estimated.

The fix is one line, loading the file wholesale, and the comment explains the decision rather than leaving it as an unexplained change. This is what a maintained compose file looks like when the maintainer bothered to write down why.

Postgres and Redis bind to localhost, the backend API binds to everything

Both install paths end the same way, by copying an environment file and starting compose:

bash
git clone https://github.com/PunithVT/ai-avatar-system.git
cd ai-avatar-system
cp .env.example .env          # add your ANTHROPIC_API_KEY (or OPENAI_API_KEY)
docker compose up -d

Docker with Compose v2 or newer is the recommended route; the manual alternative is Python 3.10 or newer, Node.js 18 or newer, FFmpeg, PostgreSQL and Redis.

The three services in that compose file are not exposed the same way. Postgres publishes on the loopback interface only, and Redis does the same, so neither is reachable from another machine.

The backend publishes its port as a plain host-to-container mapping, which binds to every interface on the host. That is correct for a web API in development and is also the one service whose exposure the compose file does not constrain.

The datastores carry resource limits and the backend does not, at least in the visible portion. Postgres is capped at one gigabyte of memory and Redis at five hundred and twelve megabytes, while a backend fanning out to Whisper, a language model, text-to-speech and a persistent lip-sync worker has no stated ceiling here. Both datastores do have real health checks, using the database's own readiness query and a ping, and the backend is documented as depending on those conditions rather than merely on container start.

Redis runs allkeys-lru under a 384 megabyte ceiling

The Redis service is not started with defaults. The command line appends persistence, sets a memory ceiling of three hundred and eighty-four megabytes, and sets the eviction policy to evict any key rather than only keys with an expiry.

All-keys eviction on a broker is a sharp choice. Celery is in the stack, since Flower is one of the documented services, which means Redis holds queued tasks. An LRU policy under memory pressure will evict the least recently used key of any kind, and a queued task that nobody has touched yet is close to the least recently used thing in the database.

The volume persists the data directory, so this is not about restarts. It is about a busy afternoon: enough concurrent conversation state plus enough queued speech and lip-sync work to cross three hundred and eighty-four megabytes, at which point the cache and the queue compete for the same budget and the queue can lose.

The five hundred and twelve megabyte container limit above that is the second constraint, and a Redis that is being evicted inside its own memory limit has already lost.

None of this is visible from the readme. It is one line in a compose file, and it is the kind of line that explains a production incident rather than causing one.

MuseTalk renders the mouth at 256 while the avatar is stored at 512

The example environment file contains the most specific technical explanation in the project, attached to the optional face restore switch, which is off by default.

The note says the lip-sync engine regenerates the mouth at 256 by 256 while the avatar is stored at a resolution setting of 512. The consequence is that the mouth is about half the resolution of the face around it. The optional restoration step, which uses GFPGAN, exists to sharpen the composited frame back up.

The note also says the step costs real time per frame and advises measuring the frame rate trade before leaving it on. That is the useful part, because a face restoration pass on every generated frame is exactly the kind of option that looks free in a feature table and dominates the budget in practice.

Below that, the avatar engine section begins enumerating choices. The first is the default, at roughly thirty frames per second on sixteen to twenty-four gigabytes of memory. The second is described as higher fidelity and then breaks off partway through the memory requirement it needs. The two engines are also the two options the comparison table scores on lip-sync quality, with the proprietary one attributed to a competing product rather than to this project.

So the engine choice is the project's main differentiator, and the note describing its second option is the sentence the page cuts off.

The comparison table is the project's own, and MuseTalk is someone else's engine

The page includes a head-to-head table against three named alternatives, scoring each on real-time conversation, lip-sync, voice cloning, barge-in, local model support, web app maturity and licensing. Every cell in its own column is a check mark and most cells in the other three are crosses.

It is a project's assessment of its competitors, in its own readme, with no external verification. The closing sentence makes the pitch explicit: this is the one you can deploy as a real multi-user web service rather than a demo.

Two things in the table are worth reading carefully. The licence row lists the project's own as MIT, one competitor as MIT, another as Apache, and Duix-Avatar as custom, which is the row where the comparison has the most consequence and the least detail behind it.

And the lip-sync row claims MuseTalk, a third-party engine that this project vendors under its own backend directory. The lip-sync capability is therefore not the project's work, and the repository structure block shows the engine checked in beneath the application with its own worker script. That does not weaken the product, but it does mean the comparison is between integration depth around a shared engine rather than between engines.

Editorial conclusion

This is a full-stack talking-head application rather than a model, and the parts worth reading first are the ones the project's own files admit to. The voice clone length appears as five seconds, ten seconds and ten to sixty seconds in three different places, so treat any of them as approximate. The data-flow diagram waits for a complete response before splitting it into sentences, which is not what the streaming feature claims. And the compose file carries a long comment describing a past bug where an explicit key allowlist silently dropped most of the application settings. None of that makes it unusable, but it is a codebase documenting its own rough edges, which is more useful to a reader than a feature table.

Frequently asked questions

What does ai-avatar-system do?

It turns an uploaded face photo into a real-time talking avatar. A turn runs through speech-to-text, a language model, text-to-speech and a lip-sync engine, with video chunks sent to the browser over a WebSocket as each sentence finishes.

How much audio does ai-avatar-system need to clone a voice?

The page gives three answers: a five-second clip in the introduction, ten seconds in the feature summary, and a ten to sixty second recording in the features table. All three name Chatterbox Multilingual zero-shot cloning as the engine.

Can ai-avatar-system run without a cloud service?

Yes. There is a local mode using local storage, local Whisper and a local model through Ollama, and the language model provider can be set to ollama with the OpenAI-compatible client pointed at a local endpoint on port 11434.

What does ai-avatar-system need before it starts?

Either Docker with Docker Compose v2 or newer, which is the recommended path, or a manual stack of Python 3.10 or newer, Node.js 18 or newer, FFmpeg, PostgreSQL and Redis. Both documentation pages begin by copying the example environment file and adding an API key.

What does the FACE_RESTORE setting do in ai-avatar-system?

It is an optional face restoration pass, off by default, using GFPGAN. The example environment explains that the lip-sync engine regenerates the mouth at 256 by 256 while the avatar is stored at 512, leaving the mouth at about half the resolution of the surrounding face.

Official sources

  1. Issues
  2. License: MIT
  3. PunithVT/ai-avatar-system on GitHub
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/punithvt-ai-avatar-system.svg)](https://hysenlabs.com/projects/punithvt-ai-avatar-system)