# AIAvatarKit: a Python speech-to-speech pipeline for talking avatars

> AIAvatarKit wires VAD, STT, an LLM and TTS into one streaming pipeline so a Python service can speak, show expressions and call tools. It is Apache-2.0, needs Python 3.11+ and a VOICEVOX-compatible server, and the README claims under one second from end of speech to first audio.

**uezo/aiavatarkit** — 🥰 Building AI-based conversational avatars lightning fast ⚡️💬

- Repository: https://github.com/uezo/aiavatarkit
- Stars: 686 · Forks: 68
- Language: Python
- License: Apache-2.0
- Published: 2026-09-14 · Updated: 2026-09-14 · Language: en
- Canonical page: https://hysenlabs.com/projects/uezo-aiavatarkit

## What AIAvatarKit actually solves for Python teams

Most conversational avatars are assembled from four moving parts: something that decides the user stopped talking, something that transcribes, a model that answers, and something that speaks. Gluing those together is where projects stall, because each vendor ships its own streaming semantics and its own idea of when a turn ends. AIAvatarKit's answer is a single STSPipeline object that owns the whole chain and streams each stage into the next.

The target reader is a Python developer who already has an avatar front end, a VRChat character, a Raspberry Pi, or a phone number, and needs the conversation layer behind it. The README lists those cases directly: talking avatars inside a web or mobile app, interactive signage and virtual store staff, companion devices such as Raspberry Pi, M5Stack and StackChan, metaverse characters on VRChat, cluster and Vket Cloud, and inbound or outbound calls through Twilio or Asterisk. That is a wide net, and it is the honest description of the project: a framework, not an application.

The distinction matters when you evaluate it. AIAvatarKit does not give you a character model, an animation rig or a hosted endpoint. It gives you the pipeline, the adapters that connect channels to it, and a set of built-in provider implementations you can swap.

## The STSPipeline and the adapter split

The architecture diagram in the README shows one STSPipeline with four sequential stages: VAD, STT, LLM with tools and agent logic, and TTS. A branch leaves the LLM stage for face, animation and artifacts. An Adapter sits on the other side of the pipeline and wraps it for one channel: WebSocket, HTTP, telephony, messaging. The README states that an adapter owns only transport concerns, and that multiple adapters can attach to the same pipeline instance, which is what makes the claim about a conversation following the user across channels possible. Hang up the phone and open LINE, and the same pipeline instance is still there.

Each stage is replaceable. The shipped implementations are worth reading as a map of what the project expects you to plug in. For voice activity detection there are Silero VAD, a streaming Silero VAD, Azure Speech, Amazon Transcribe, Parapper and a plain volume threshold. Turn-end gates, which the README groups under semantic VAD, include Smart Turn, Namo Turn, a filler-only gate, an LLM-based gate, a session hold gate and a custom gate. Speech-to-text covers Azure Speech, Google Cloud Speech-to-Text, OpenAI, AmiVoice and any OpenAI-compatible endpoint. The LLM side covers OpenAI Chat Completions, Azure OpenAI, the OpenAI Responses API, Anthropic Claude, Google Gemini, xAI Grok and OpenRouter.

Two design choices stand out. The first is the split between VAD and turn-end gates. Detecting speech is not the same problem as detecting that a sentence has finished, and the project treats them as separate modules, which is why a filler-only gate can exist at all. The second is that tools load only when needed, per the feature list, with background execution or a reply from a template for slow ones. That is a real constraint on how you write tool handlers, not just a marketing line.

## Install AIAvatarKit and run the built-in app

The quick start needs Python 3.11 or newer, an OpenAI API key, and a reachable VOICEVOX-compatible server at its default URL. Install the package first.

```bash
pip install aiavatar
```

Then export your key and start the built-in default application. The `aiavatar` console script is declared in pyproject.toml as `aiavatar.cli.command:main`, so it is on your PATH after the install.

```bash
export OPENAI_API_KEY=sk-xxx
aiavatar
```

On first run the built-in application uses Namo Turn semantic VAD. If its optional dependencies are missing, the command asks whether to install them, printing `Additional dependencies are required to enable Semantic VAD. Install them now? [y/N]:`. Answering `y` installs `aiavatar[namo-turn]` into the current Python environment and continues startup. Answering `n` starts without the Namo Turn gate, and the README says Silero VAD and the filler gate remain enabled. That prompt is the first place the project's optional-dependency design becomes visible: `namo-turn` and `smart-turn` both pull onnxruntime, transformers and huggingface-hub.

Once it is up, the avatar is at http://127.0.0.1:8000/ and the Admin Panel is at http://127.0.0.1:8000/admin/. Launch VOICEVOX beforehand or the TTS stage has nothing to talk to.

## Writing your own server instead of using the CLI

The built-in application is a convenience. The README is explicit that you write your own script when you need fine-grained control over components, routes and application lifecycle. The example saves as `run.py` and builds a WebSocket adapter around the pipeline.

```python
import os

from fastapi import FastAPI
from fastapi.staticfiles import StaticFiles
from aiavatar.adapter.websocket.server import AIAvatarWebSocketServer
from aiavatar.admin import setup_admin_panel
from aiavatar.util import download_example

html_dir = download_example("websocket/html")

aiavatar_app = AIAvatarWebSocketServer(
    openai_api_key=os.environ["OPENAI_API_KEY"]
)

app = FastAPI()
app.include_router(aiavatar_app.get_websocket_router())
setup_admin_panel(app, adapter=aiavatar_app)
app.mount("/", StaticFiles(directory=html_dir, html=True), name="ui")
```

The comments in that example carry two constraints worth keeping. The `setup_admin_panel` call is optional. The `app.mount` catch-all must come after the routes above it, or the UI mount swallows the WebSocket and admin routes. Start it with the module form.

```bash
python -m uvicorn run:app
```

The same URLs apply: port 8000 for the avatar, `/admin/` for the panel. If you prefer configuration over code, the repository ships `.env.example`, which documents the `AIAVATAR_*` settings used when no application script is supplied. Copy it to `.env`, edit `OPENAI_API_KEY`, and run `aiavatar`; the command loads `.env` from the current working directory without replacing variables already set in the process environment. Command-line options override the matching environment variables, and component-specific OpenAI settings override the shared `OPENAI_*` values.

## Where AIAvatarKit is the wrong tool

The quick start assumes a VOICEVOX-compatible server at its default URL. That is a hard dependency of the default path, and it is the first thing to check if you are deploying to a container or a device that cannot run a second service alongside your Python process. The README does not present a cloud TTS substitute as part of the quick start; the built-in application is arranged around that local server.

The optional dependency groups are the second boundary. `namo-turn` and `smart-turn` each pull onnxruntime, transformers and huggingface-hub, and the semantic VAD path is the one the built-in application prefers. On a Raspberry Pi or an M5Stack, that is a meaningful amount of runtime to install and load. The README's own escape hatch is to answer `n` at the prompt and fall back to Silero VAD with the filler gate, which is a real downgrade in turn detection quality rather than an equivalent option.

There is also a version trap the README calls out directly: if the steps in a technical blog do not work, the blog may describe a version prior to v0.6, and you can try `pip install aiavatar==0.5.8` to match that environment. That is a candid warning, and it also tells you the API moved. Treat third-party tutorials as historical unless they name a version.

Finally, this is infrastructure, not a product. If you want a hosted avatar you configure in a browser, AIAvatarKit is the wrong layer. You will be writing Python, choosing providers, and running at least one extra process.

## How it differs from a full avatar platform

The closest comparison is a complete avatar platform that ships the character, the voice, the hosting and the editor as one product. Those platforms decide your STT, your LLM and your TTS for you. AIAvatarKit inverts that: you supply the providers, and the project supplies the streaming contract between them plus the adapters to reach channels.

The practical difference shows up when a requirement falls outside the platform's menu. If you need Azure Speech for STT because of a regional constraint, Anthropic Claude for the model, and a local VOICEVOX server for the voice, a hosted platform cannot express that combination at all. AIAvatarKit's component table exists precisely for that case, and the README notes that a small interface covers providers beyond the built-in ones. The cost is that you own the deployment, the latency budget and the failure modes of every provider you pick.

A narrower alternative is to assemble the same pieces yourself over FastAPI and WebSocket. That is more work than it sounds, because the interesting parts are not the provider calls but the streaming and the turn-taking: the README describes STT running speculatively and the avatar giving a spoken nod before the answer itself, plus guardrails running in parallel that can interrupt the avatar mid-sentence. Those behaviors are pipeline-level, and they are the reason to take a dependency here rather than write your own loop.

## Maintenance, licence and upgrade cost

The repository is not archived. The last push was on 2026-09-10, and the most recent release listed is v0.9.0 from 2026-08-29, following v0.8.19 in July and v0.8.18 in early July. The package version in pyproject.toml matches v0.9.0. That release cadence, with patch releases landing roughly every two to three weeks, is the practical signal: the surface you build against is still moving.

The licence is Apache-2.0, declared both in the repository metadata and in pyproject.toml as `license = { text = "Apache-2.0" }`. That is a permissive licence with an explicit patent grant, which matters if you are shipping a commercial avatar product. It does not settle the terms of the providers you connect to. OpenAI, Azure, Google, Anthropic and the rest have their own commercial conditions, and VOICEVOX has its own terms for voice usage. Read those separately; the Apache-2.0 grant covers this library's code, not the services it calls.

On upgrade cost, the README's note about pre-v0.6 tutorials is the concrete evidence that breaking changes have happened. The optional-dependency split into `local`, `local-audio`, `smart-turn` and `namo-turn` groups means an upgrade can also change what gets installed, and the interactive prompt at startup is a symptom of that. Pin the version in production, and read the release notes before moving, because the difference between v0.8.18 and v0.9.0 is not documented in the README.

## Conclusion

Adopt AIAvatarKit when your avatar needs a real Python service behind it: swappable VAD, STT, LLM and TTS modules, adapters for WebSocket, telephony and messaging, and an admin panel at /admin/. Do not adopt it if you want a hosted, no-code avatar or if you cannot run a VOICEVOX-compatible server at http://127.0.0.1:50021, because the quick start assumes one. Verify first that your Python is 3.11 or newer and that pip install aiavatar resolves the onnxruntime and silero-vad dependencies on your target hardware, since those are the parts most likely to fail on a small device.

## FAQ

### What is AIAvatarKit?

It is a Python framework, distributed as the aiavatar package under Apache-2.0, that builds AI-based conversational avatars. It wraps voice activity detection, speech-to-text, an LLM with tools, and text-to-speech into a single STSPipeline, with adapters for channels such as WebSocket, telephony and messaging.

### How do I install AIAvatarKit?

Install it with pip install aiavatar. The quick start requires Python 3.11 or newer, an OpenAI API key, and a reachable VOICEVOX-compatible server at its default URL; then export OPENAI_API_KEY and run the aiavatar command.

### Does AIAvatarKit need a VOICEVOX server?

The quick start assumes one. The README states that the minimal setup is OPENAI_API_KEY plus a reachable VOICEVOX-compatible server at http://127.0.0.1:50021, and the script example tells you not to forget to launch VOICEVOX beforehand.

### Which LLM and speech providers does AIAvatarKit support?

The component table lists OpenAI Chat Completions, Azure OpenAI, the OpenAI Responses API, Anthropic Claude, Google Gemini, xAI Grok and OpenRouter for the LLM stage, and Azure Speech, Google Cloud Speech-to-Text, OpenAI and AmiVoice, plus any OpenAI-compatible endpoint, for speech-to-text.

### Why do steps from a technical blog not work with AIAvatarKit?

The README says the blog may be based on a version prior to v0.6. It suggests trying pip install aiavatar==0.5.8 to match the environment described in the blog, noting that some features may be limited.

## Sources

- [Issues](https://github.com/uezo/aiavatarkit/issues)
- [License: Apache-2.0](https://github.com/uezo/aiavatarkit/blob/main/LICENSE)
- [README](https://github.com/uezo/aiavatarkit/blob/main/README.md)
- [Releases](https://github.com/uezo/aiavatarkit/releases)
- [uezo/aiavatarkit on GitHub](https://github.com/uezo/aiavatarkit)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/uezo-aiavatarkit
