Model or dataset
met4citizen/TalkingHead avatar
met4citizen/TalkingHead

The TalkingHead package publishes a modules folder, and its default voice is a Google Cloud call

Talking Head (3D): A JavaScript class for real-time lip-sync using full-body 3D avatars.

1,583 stars354 forksJavaScriptMIT

At a glance

What is it?
TalkingHead is a browser JavaScript class that renders a full body 3D avatar through ThreeJS and lip syncs it in real time. It depends on one package, publishes only its modules folder to npm, needs an avatar carrying a Mixamo rig and viseme blend shapes, and speaks through Google Cloud until you replace the TTS engine.
Who is it for?
TalkingHead earns adoption when you already have, or can get, a Ready Player Me style avatar with a Mixamo rig and the ARKit and Oculus viseme shapes, because that requirement is the one thing the class cannot work around and nothing in the package supplies it. Three things are worth checking before you start.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 8 days ago.
What is it written in?
Mainly JavaScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The published package is a modules folder, and the assets stay in the repository

Install it from npm and you receive three entries: the README, the LICENSE, and everything under `modules`. That is the whole published surface, declared in the manifest rather than produced by a build step, and it means the parts of the repository that make the demos work are not part of the package. The avatar library, the animation and pose folders, the fonts, the audio samples, the images, the four example pages, the tests directory and the site configuration all stay in git. A consumer of the published package supplies the avatar, and the avatar is the part with hard requirements.

Two more manifest entries shape every import. The package is an ES module with a single `.mjs` entry point, and the dependency list is short enough to read in one breath:

json
"main": "modules/talkinghead.mjs",
"type": "module",
"files": [
  "README.md",
  "LICENSE",
  "modules/*"
],

Three.js is the only runtime dependency, at `^0.180.0`, which matches the stated renderer: the class draws through ThreeJS and WebGL. Nothing else is installed on your behalf. The model, the audio pipeline and the speech service all arrive as your own code.

A GLB file is not an avatar, and the two descriptions disagree about why

The stated requirement is a model with a Mixamo-compatible rig plus both ARKit and Oculus viseme blend shapes, and the class renders full-body GLB avatars with Mixamo FBX animations on top. An arbitrary 3D file will load and then stay mute, because the viseme shapes it needs are named conventions rather than geometry a model happens to contain. A rig built for a different animation pipeline does not substitute either, and the animations you load have to match the rig you load.

The two descriptions in the repository disagree about how much of that is the author's doing. The repository description says the class lip syncs using full-body 3D avatars, while the package manifest narrows the same sentence to Ready Player Me full-body 3D avatars. Ready Player Me avatars ship with the Mixamo rig and the viseme sets, which is what makes the manifest the more accurate of the two and the repository description the sentence that will mislead someone sourcing a model elsewhere. Building an avatar from scratch is the subject of Appendix A, and the `blender/` directory at the root is where the rigging side of that work sits.

Six built-in lip-sync languages are a module count, not a capability limit

Six languages work out of the box: English, German, French, Brazilian Portuguese, Finnish and Lithuanian. That number counts language modules shipped inside the package, and it does not describe what the class can reach, because a module only does one job: turning text timing into viseme targets. Any TTS service returning word-level timestamps can drive it, which is why the ElevenLabs WebSocket API is named as an integration point.

The larger count comes from skipping that step entirely. If the engine emits viseme IDs or blend shape data directly, no language module is required, and the Microsoft Azure Speech SDK is the named example, extending lip sync to more than 100 languages. Both routes have a page waiting in the examples folder, `azure-blendshapes.html` for direct shape output and `azure-audio-streaming.html` for the streaming case, beside `minimal.html` and `mp3.html`. Four example files, none of them published. The class also knows a set of emojis and converts them into facial expressions, so text input carries more than speech.

The default voice is a paid network call, and nothing in the package holds a credential

Google Cloud TTS is the default, which makes the out-of-the-box configuration a billing decision you inherit before writing a line of your own. Nothing in the dependency list installs a speech engine or supplies a key: the single runtime dependency is Three.js. Whatever credential gets the request through comes from the application you write.

A free path exists, as a separate package. HeadTTS is an add-on module described as free and open-source English TTS with Kokoro neural voices, and the module table that lists it is also where the commercial alternatives are named. The second module, HeadAudio, carries the warning that belongs to this whole category of stack: realtime speech-to-speech over WebRTC through the OpenAI Realtime API costs much more than standard AI text tokens, and it points at `gpt-realtime-mini` pricing as the first thing to check. A fully in-browser build exists as a third-party app instead, pairing the class with HeadTTS, whisper-web and WebLLM on Llama 3.2, with no API and no account, plus a note that a desktop Chrome or Edge is wanted for WebGPU performance.

npm test exits 1 by design, and the demos are screen captures from a browser

The only script in the manifest is `test`, and its body prints a line saying no test is specified and exits with status 1. That is the placeholder npm generates when a package is created, and it is still there. A `tests/` directory does sit at the root, but nothing wires it to that script, and the directory is not among the published entries either.

What occupies the place of a suite is a browser. Every demo video is presented as a real-time screen capture from Chrome running the TalkingHead test web app, with no post-processing, so what the videos show is a person watching the class run in a page. The evidence for the lip-sync accuracy claim is a close-up view in English and Finnish, with the engines named alongside it: GPT-3.5 and Microsoft text-to-speech. Another pairs OpenAI's Whisper with the `speakAudio` method over two MP3 tracks. Verification here is visual and manual, which is also why the closest thing to a regression test lives in `examples/` and in the test app rather than in a runner.

The branch is nine months past the newest tag, and the manifest still says 1.7.0

Three tagged releases exist and the newest is from December 2025. v1.7.0 is dated 2025-12-08, v1.6.0 is dated 2025-09-03 and v1.5.0 is dated 2025-05-29, two gaps of about three months each, then nothing for the nine months since. The last push to the default branch is 2026-09-25, and the version field has not moved:

json
"version": "1.7.0",

So two different things answer to the same name. Installing `@met4citizen/talkinghead` gives you the December 2025 snapshot, while the documentation on the branch describes modules, add-ons and example pages that the branch has been changing since. Anyone following the four files in `examples/` is working from the repository, not from the package, and the package is the one carrying a version number. Pin deliberately: there is no release between the one on npm and the code in front of you.

The audience is research prototypes, and the class is a library rather than a service

Read the use case list as a map of where this ends up. DialogLab, developed with researchers from UVA, Google, Northeastern, Google DeepMind and Google Research, is a toolkit for authoring, simulating and testing human-AI group conversations. A CHI 2025 paper from the MIT Media Lab and Harvard used the class to build interactive dating profiles around digital twins that a potential partner could interact with. Two University of Florida efforts appear, one on multiple virtual agents helping people join cancer clinical trials, and one published in Frontiers in Digital Health on embodied chatbots and mental well-being. Smaller projects fill the rest: a kid friendly chat, a game show with its own host, a live Twitch adventure, and a quantum physics walkthrough built on the CHSH game.

Two things follow. The class is a browser library embedded in pages rather than a service, so every one of these examples inherits the browser's limits, including the WebGPU note attached to the fastest in-browser build. And the name collides with an unrelated film and marketing phrase, which is why the manifest keyword list reaches past the name to lip-sync, 3D avatar and talking avatar while anyone searching the bare words lands on the other meaning.

Editorial conclusion

TalkingHead earns adoption when you already have, or can get, a Ready Player Me style avatar with a Mixamo rig and the ARKit and Oculus viseme shapes, because that requirement is the one thing the class cannot work around and nothing in the package supplies it. Three things are worth checking before you start. The version you install from npm is the December 2025 release, so read the documentation against the repository rather than the package. The default voice bills Google Cloud, so decide between that, the free HeadTTS module, and an engine that emits blend shapes directly. And accept that there is no automated test to run, since the manifest test script exits 1 by design. For offline lip sync inside a build you can inspect, this class fits. For a documented public API surface or CommonJS support, the manifest offers neither.

Frequently asked questions

What is TalkingHead, the 3D avatar library from met4citizen?

It is a browser JavaScript class that renders a full body 3D avatar with ThreeJS and WebGL and lip syncs it in real time, MIT licensed and authored by Mika Suominen. It also reads emojis and turns them into facial expressions. Rendering and lip sync are built in; speech comes from a TTS engine you choose.

What avatar format does TalkingHead need?

A full body GLB with a Mixamo-compatible rig and both ARKit and Oculus viseme blend shapes, animated with Mixamo FBX clips. The package manifest narrows this to Ready Player Me avatars, which ship with that rig and those viseme sets. Appendix A covers creating your own avatar.

Which languages does TalkingHead lip sync in?

English, German, French, Brazilian Portuguese, Finnish and Lithuanian are built in, one module per language. New languages need a new module, unless the TTS engine outputs viseme IDs or blend shapes directly, in which case the Microsoft Azure Speech SDK path extends support to 100+ languages with no module.

How do I install TalkingHead from npm?

The package is `@met4citizen/talkinghead`, an ES module whose entry point is `modules/talkinghead.mjs`, with Three.js as its only runtime dependency at `^0.180.0`. The published files are the README, the LICENSE and `modules/*`, so avatars, animations, poses, fonts and the example pages come from the repository rather than the package.

Does TalkingHead need an API key to speak?

In its default configuration, yes. Google Cloud TTS is the default engine and no dependency supplies a credential, so the key comes from your own code. The free alternative is the HeadTTS add-on module, an open-source English TTS with Kokoro neural voices, and engines such as the Microsoft Azure Speech SDK can emit blend shapes instead.

Official sources

  1. Issues
  2. License: MIT
  3. met4citizen/TalkingHead on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/met4citizen-talkinghead.svg)](https://hysenlabs.com/projects/met4citizen-talkinghead)