# ChatdollKit: Turning a Unity VRM Model into a Voice Chatbot

> ChatdollKit is an Apache-2.0 Unity SDK that wires an LLM, speech recognition and speech synthesis to a 3D model. It solves the plumbing problem of voice conversation in Unity, but it is not a drop-in chatbot for non-Unity projects.

**uezo/ChatdollKit** — ChatdollKit enables you to make your 3D model into a chatbot

- Repository: https://github.com/uezo/ChatdollKit
- Stars: 1,227 · Forks: 126
- Language: C#
- License: Apache-2.0
- Published: 2026-09-14 · Updated: 2026-09-14 · Language: en
- Canonical page: https://hysenlabs.com/projects/uezo-chatdollkit

## What ChatdollKit actually solves for a Unity developer

Building a talking 3D character in Unity means solving at least five separate problems: capturing microphone audio, detecting when the user has finished speaking, sending text to a language model, turning the reply back into audio, and then driving the model's mouth and body in time with that audio. Each of those has its own ecosystem of packages, and the glue between them is where projects stall. ChatdollKit's contribution is that glue. The README describes it as a "3D virtual assistant SDK that enables you to make your 3D model into a voice-enabled chatbot", and the repository layout backs that up: Scripts/, Prefabs/, Audio/, Textures/ and Demo/ sit side by side, which is the shape of a Unity package that ships both code and scene assets.

The intended user is a Unity developer who already has a character. The topics list on the repository includes vrm, which points at the VRM humanoid avatar format rather than Unity's native rigs, and the README mentions VRM runtime loading in the 0.8.6 notes. If your model is a VRM file and your project is Unity, you are the target audience. If you are building a chat interface in React, or a voice assistant on a Raspberry Pi, nothing here applies to you.

The scope is deliberately broad on the backend side. The feature list names ChatGPT, Anthropic Claude, Google Gemini Pro and Dify as supported LLMs, with function calling on ChatGPT and Gemini, and lists OpenAI, Azure, Google, VOICEVOX, AivisSpeech, Aivis Cloud API and Style-Bert-VITS2 among the speech components. That breadth is the selling point: you are not locked to one vendor's stack, and swapping providers is a configuration change rather than a rewrite.

## How the pieces fit together: LLM, STT, TTS and the model controller

The architecture diagram in the repository is named chatdollkit_architecture_overview.png, and the release history fills in what the diagram implies. Audio comes in through a speech listener. In v0.8.3 the project added AzureStreamSpeechListener for recognizing speech as it is spoken, and v0.8.16 added WebSocket-based streaming speech recognition that, per the release notes, "offloads VAD to the server and completes recognition during turn-end detection". That is the input path: microphone to text, with voice activity detection deciding where a turn ends.

The text then goes to the dialog layer. The README describes dialog control as managing dialog state (context), extracting intents, routing topics and supporting wakeword detection. In practice this means the toolkit keeps the conversation history, decides whether an utterance is addressed to the character at all, and forwards the prompt to whichever LLM you configured. Function calling support on ChatGPT and Gemini means the model can invoke your own code as part of a reply, which is how you would let the character look something up rather than improvise.

The reply comes back as text and is synthesized to audio by one of the TTS providers. That audio drives lip-sync, and separately the toolkit drives facial expressions, blinking and animations on the model. The 0.8.16 notes describe a refactor that extracted speech handling into a SpeechController and face expressions into a FaceController, both pulled out of the older ModelController. That split matters if you plan to customize: expression logic and speech logic now have separate homes, so changing how the mouth moves does not require touching how audio is scheduled.

Barge-in, also new in 0.8.16, changes the input path in a way worth understanding. The user can interrupt the character mid-sentence. That means the speech listener has to stay active while TTS is playing, and the system has to cancel or duck the ongoing synthesis. The 0.8.14 notes mention "automatic volume control when users interrupt during AI speech", which is the earlier, softer version of the same idea.

## Installing ChatdollKit and getting a first conversation running

The README does not contain a step-by-step installation section. It points readers at the live WebGL demo and at an iOS app called OshaberiAI built with the toolkit, and the repository ships a Demo/ directory containing Demo08.unity alongside controller assets such as AGIA.controller and AGIAFree.controller. The practical starting point is therefore the demo scene rather than a documented install command. Clone the repository and open the project in Unity, then open Demo/Demo08.unity from the Project window.

```bash
git clone https://github.com/uezo/ChatdollKit.git
```

The clone gives you the full repository, including Demo/, Prefabs/ and Scripts/. Because the README gives no package manifest or UPM install line, treat the clone as the supported route until the documentation says otherwise. If you prefer to vendor the code into an existing project, the top-level directories are what you would copy.

Once the demo scene is open, the configuration you need to supply is provider credentials. The README lists OpenAI, Azure and Google among the speech services and ChatGPT, Claude, Gemini Pro and Dify among the LLMs, so the exact keys depend on which combination you pick. The repository does not document a single canonical config file, so check the demo scene's inspector fields and the component scripts under Scripts/ for the fields that expect a key.

```bash
# Run the demo scene in the Unity Editor, then say "Hello" to start.
# The README describes the WebGL demo this way: say "Hello" to start conversation.
```

The README's own description of the live demo is that saying "Hello" starts the conversation, and that the character is multilingual, so you can ask her to switch languages mid-conversation. Expect the same behaviour from the demo scene. If you hear nothing, the failure is almost always a missing or rejected API key, because the pipeline has no offline fallback.

## Where ChatdollKit is the wrong tool

The most obvious limitation is that everything runs inside Unity. There is no headless server mode described in the README, no HTTP API for driving a character from another process, and no non-Unity client. If your product is a web app with a 2D avatar, or a Slack bot, ChatdollKit gives you nothing you could not get from calling an LLM and a TTS API directly.

A second constraint is dependency on external services. Every conversation path in the README runs through a remote LLM and a remote or local speech service. That means latency you do not control, cost per conversation, and a hard dependency on network availability. The v0.8.16 streaming STT work is explicitly aimed at reducing latency, which tells you latency was a real problem worth a release. The release notes claim a reduction of "several hundred milliseconds" from offloading VAD to the server; that is a vendor claim from the changelog, not an independent measurement, and it will vary with your network and provider.

There is also a migration cost that the README acknowledges rather than hides. Version 0.8.4 removed legacy components and the README directs anyone updating from 0.7.x to a migration section. If you find an older tutorial or sample online, check which version it targets before following it, because the component names have changed. The 0.8.16 refactor moved speech and face handling out of ModelController, so code written against 0.8.15 that reached into ModelController for those responsibilities will need adjusting.

Finally, the documentation is thinner than the feature list. The README is organized as a changelog with a feature summary, not as a reference manual. There is no documented rollback path, no version compatibility matrix, and no troubleshooting section. You will be reading source under Scripts/ to answer questions the README does not address.

## ChatdollKit compared with ChatVRM and AIAvatarKit

The related searches around this project surface two names worth separating. ChatVRM is the other frequently searched term, and it is a different kind of artifact: an application you run to get a talking VRM character in a browser. ChatdollKit is an SDK you build with, in Unity. The difference in approach matters for what you can change. With ChatVRM you are working inside someone else's app; with ChatdollKit you are writing the app and the toolkit supplies the conversation and animation layers. If you want a working demo in five minutes, the app-shaped project wins. If you need the character to do something specific in your own scene, the SDK is the only one of the two that gives you the seam.

AIAvatarKit appears in the same search list and in the 0.8.11 and 0.8.12 release notes, where the README describes an "AIAvatarKit Backend" that "offloads AI agent logic to the server". This is the same author's server-side counterpart, and the two are designed to be used together rather than as alternatives. The split is architectural: ChatdollKit keeps the character, the animation and the audio in the Unity client, while AIAvatarKit holds the agent logic, which the notes say lets you plug in frameworks like AutoGen. If your agent logic is getting heavy, or you want to share it across multiple clients, that division is the documented answer. If your logic is a prompt and a function call, keeping it in ChatdollKit is simpler.

ChatMemory is the third related name, and it addresses a gap rather than competing. The 0.8.10 notes describe long-term memory support, with components provided for ChatMemory and the option to integrate services like mem0 or Zep. ChatdollKit's own dialog state is per-conversation context; persistent memory across sessions is a separate component you add.

## Licence, maintenance and what an upgrade costs you

ChatdollKit is licensed under Apache-2.0, which permits commercial use, modification and distribution provided you keep the licence and notices intact and state significant changes. That is permissive enough for a commercial product. The licence covers the SDK, not your assets: the VRM model you load, the voice data you synthesize with, and the LLM provider's terms are separate agreements, and the repository says nothing about them. If you are shipping a character based on a licensed model or a cloned voice, that is where your legal exposure sits, not in the Apache-2.0 grant. This is not legal advice; it is a note about which parts of the stack the repository licence does and does not reach.

The repository is not archived. The last push was on 2026-09-10, and the most recent tagged release is v0.8.16 from 2026-02-14, following v0.8.15 on 2025-08-21 and v0.8.14 on 2025-08-12. So commits continue between releases, and the gap between the v0.8.15 and v0.8.16 tags is roughly six months. Plan for that cadence: a release every few months, with the component structure occasionally reorganized. The 0.8.4 removal of legacy components and the 0.8.16 controller refactor are both examples of changes that can touch your code, so pinning to a tag and reading the changelog before moving is the reasonable posture.

Upgrade cost concentrates in three places: renamed or relocated components, provider API changes, and platform-specific behaviour. The 0.8.13 notes mention removing OpenAI-specific parameters from the OpenAI-style endpoint so that Grok and Gemini work, which is exactly the kind of change that alters request payloads. The 0.8.14 and 0.8.15 notes cover platform-specific microphone and VAD work for Android, iOS, macOS and WebGL, so if you ship on multiple platforms, budget test time per platform per upgrade.

## Conclusion

Adopt ChatdollKit if your character already lives in Unity, you have a VRM model, and you want the LLM, speech-to-text, text-to-speech and animation layers wired together rather than assembled by hand. Do not adopt it if your avatar is a browser page, a mobile app without a Unity build, or a backend service with no 3D output; the whole toolkit assumes a Unity scene and a model to animate. Before committing, verify three things: that the v0.8.16 release notes' claims about WebSocket streaming STT and barge-in match the components actually present in Scripts/, that your chosen LLM and TTS providers are among those the README lists, and that you can accept the Apache-2.0 terms alongside whatever licence your VRM model and voice data carry. The migration section for 0.7.x is in the README, and it is worth reading before you start rather than after.

## FAQ

### What is ChatdollKit?

It is a 3D virtual assistant SDK for Unity that turns a 3D model into a voice-enabled chatbot, wiring together an LLM, speech-to-text, text-to-speech and model animation. It is licensed under Apache-2.0 and written in C#.

### Which LLMs and speech services does ChatdollKit support?

The README lists ChatGPT, Anthropic Claude, Google Gemini Pro and Dify among the LLMs, with function calling on ChatGPT and Gemini. For speech it names OpenAI, Azure, Google, VOICEVOX, AivisSpeech, Aivis Cloud API and Style-Bert-VITS2.

### Which platforms can ChatdollKit run on?

The README states compatibility with Windows, Mac, Linux, iOS, Android and other Unity-supported platforms, including VR, AR and WebGL. Platform-specific microphone and voice activity detection work appears in the 0.8.14 and 0.8.15 release notes.

### Does ChatdollKit support VRM models?

Yes. The repository topics include vrm, and the 0.8.6 release notes describe improved VRM runtime loading that allows switching 3D models at runtime. The demo scenes in Demo/ use controller assets such as AGIA.controller.

### Can users interrupt the character while it is speaking?

Yes. The v0.8.16 release notes list barge-in support, which lets users interrupt AI speech mid-sentence with their voice. The earlier 0.8.14 notes mention automatic volume control when users interrupt during AI speech.

### Does ChatdollKit remember past conversations?

The 0.8.10 release notes describe long-term memory support, with components provided for ChatMemory and the option to integrate services such as mem0 or Zep. The per-conversation dialog state and this persistent memory layer are separate.

## Sources

- [Issues](https://github.com/uezo/ChatdollKit/issues)
- [License: Apache-2.0](https://github.com/uezo/ChatdollKit/blob/master/LICENSE)
- [README](https://github.com/uezo/ChatdollKit/blob/master/README.md)
- [Releases](https://github.com/uezo/ChatdollKit/releases)
- [uezo/ChatdollKit on GitHub](https://github.com/uezo/ChatdollKit)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/uezo-chatdollkit
