Open-LLM-VTuber: a local AI companion that talks, listens and blinks
Talk to any LLM with hands-free voice interaction, voice interruption, and Live2D taking face running locally across platforms
At a glance
- What is it?
- Open-LLM-VTuber wires an LLM, speech recognition, text-to-speech and a Live2D avatar into one app that runs on Windows, macOS and Linux. The documentation is thinner than the feature list, and v2.0 is a rewrite in planning, so read the boundary conditions before you install it.
- Who is it for?
- Adopt it if you want a self-hosted voice companion with a Live2D face and you are comfortable editing YAML and reading a FastAPI server log. Do not adopt it if you need a stable API surface, long-term memory, or a project that accepts v1 feature requests: the README states v2.0 is a complete rewrite in early planning and asks people to stop filing feature requests against v1.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 138 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap Open-LLM-VTuber fills: a voice loop you own
Most AI companion products are a chat box with a voice button. Open-LLM-VTuber is an attempt to close the whole loop locally: speech recognition, an LLM, text-to-speech and a Live2D model that moves while it speaks, with the option to keep every component on your own machine. The README frames the origin directly: the project was started to recreate the closed-source AI VTuber neuro-sama using open-source pieces that run offline on platforms other than Windows.
The audience is narrow and specific. You need to be willing to run a Python server, point it at a model backend, and tune a persona. The README lists the intended uses as virtual girlfriend, boyfriend, cute pet or any other character, which tells you the project is aimed at personal companion setups rather than customer service bots or scripted NPCs in a game. If you want a voice assistant for a product, the Live2D layer and the desktop pet mode are dead weight.
What makes it more than a demo is the interaction detail. Voice interruption without headphones is listed as a feature, meaning the AI is expected not to hear its own output, which is the failure mode that kills naive speaker-and-microphone loops. Visual perception covers camera, screen recording and screenshots. Touch feedback responds to clicks and drags. None of that is unique individually; the combination in one local app is the point.
How the pieces connect: FastAPI server, websocket client, config-driven backends
The repository layout tells most of the architecture. run_server.py starts the backend. src/ holds the Python service. frontend/ is the web client. web_tool/ is a separate tool directory. model_dict.json and config_templates/ hold the wiring that maps a chosen LLM, ASR and TTS provider to code. The dependency list in pyproject.toml confirms the shape: fastapi[standard] and uvicorn[standard] for the server, websocket-client for the channel to the front end, and a long list of provider SDKs including anthropic, openai, groq, elevenlabs, cartesia, azure-cognitiveservices-speech and edge-tts.
Speech recognition is local by default in the sense that sherpa-onnx and onnxruntime are core dependencies, so an offline ASR path exists without a cloud call. The LLM side is pluggable rather than bundled: Ollama appears as a topic on the repository, and the README's model support section names Ollama and OpenAI among the options. That means the heavy model is your choice and your hardware problem.
One constraint shapes deployment more than any other, and the README states it plainly: if you run the server on one machine and open the page on another, you need https, because the browser microphone only starts in a secure context. Localhost is exempt. A phone talking to a desktop therefore requires a reverse proxy with a certificate. That is a browser rule, not a project bug, and no amount of configuration inside the app removes it.
Installing Open-LLM-VTuber and getting one reply out of it
The README points to the documentation site at open-llm-vtuber.github.io/docs/quick-start for setup, and the repository ships several install paths: a plain requirements.txt, a uv.lock, a pixi.lock with conda channels for cudnn and cudatoolkit, and a dockerfile with a published image on Docker Hub. The Python version is pinned in pyproject.toml as >=3.10,<3.13, so a 3.13 interpreter will not work.
The uv path is the one the lockfiles suggest as primary. From a clone of the repository:
uv sync
uv run run_server.pyThe server starts on the FastAPI app and prints its listening address in the log. Open that address in a browser on the same machine; localhost counts as a secure context, so the microphone prompt appears.
If you prefer the container, the image is published under the project's Docker Hub namespace:
docker pull open-llm-vtuber/open-llm-vtuber
docker run --rm -p 12393:12393 open-llm-vtuber/open-llm-vtuberThe port above is the one the project uses for the server; confirm it against the Docker section of the documentation before exposing anything, and remember the https requirement as soon as the browser is not on localhost.
Configuration lives in YAML. The repository ships config_templates/ and a conf.yaml is generated from it on first run, which is where you select the LLM, ASR and TTS providers and paste any API keys. The character side is separate: characters/ holds persona definitions and live2d-models/ holds the avatar, and the README links a Character Customization Guide for appearance and persona. Expect your first working session to be spent in those two directories, not in code.
Where Open-LLM-VTuber breaks or is the wrong tool
Long-term memory is the clearest gap, and the README admits it: the feature is described as temporarily removed and coming back soon. Chat logs persist, so you can resume a previous conversation, but persistence of transcripts is not the same as a memory system that recalls facts across sessions. If your use case depends on the companion remembering what you told it last week, this is not the build for it.
The project also asks you not to treat v1 as a moving target. The README states that development attention is on v2.0, a complete rewrite in early discussion and planning, and asks people to refrain from opening new issues or pull requests for feature requests on v1. Bug fixes continue for v1 and existing pull requests are being worked through. Read that as a freeze on ambition, not on maintenance: you can expect the current behaviour to keep working, but not to grow.
Two more boundaries matter. Remote access without https simply fails at the microphone, so a self-hosted setup behind a plain IP address will look broken to a user on a phone. And the model support is broad but not uniform: the README mentions GPU acceleration for some components on macOS and support for NVIDIA and non-NVIDIA GPUs, CPU execution, or cloud APIs for the expensive parts. A machine with no GPU is a supported configuration, but it moves the cost to latency or to a cloud bill, and the documentation does not promise a latency figure either way.
Open-LLM-VTuber compared with AIRI and with a plain chat client
The comparison people actually search for is AIRI versus Open-LLM-VTuber. Both sit in the AI companion space with an avatar, and the honest difference visible here is packaging and language. Open-LLM-VTuber is a Python server with a web front end and a separate desktop client, configured through YAML templates, with a stated goal of offline operation and a Live2D model as the face. If you want to run the whole stack on your own hardware and edit configuration files, that shape is an advantage. If you want a packaged desktop application you install and forget, the server-plus-browser model is friction.
The other real alternative is not another VTuber project at all: it is a chat client plus a separate TTS tool. You get the conversation and the voice, you skip the avatar, the touch feedback and the desktop pet mode, and you give up the interruption handling that the README highlights. That trade is correct when the face is decoration for you. It is wrong when the point is presence, which is exactly what the Live2D layer and the transparent, always-on-top pet mode are for.
Worth noting for anyone comparing on features: the repository carries a separate LICENSE-Live2D.md alongside LICENSE, which reflects that Live2D has its own terms distinct from the project's own licence. Read both before shipping anything with a bundled model.
Maintenance, upgrades and what the licence does not tell you
The last push to the default branch was on 2026-05-15, and the most recent tagged release is v1.2.1 from 2025-08-26, with 1.2.0 before it in August 2025 and v1.1.0 in February 2025. The release cadence is therefore slow and uneven, and the README's own notice about v2.0 explains why: effort is being redirected to a rewrite that has not shipped.
The repository is not archived, so it is not abandoned, but the combination of a rewrite notice and a five-month gap since the last push means you should judge it by what v1.2.1 does today rather than by a roadmap.
Upgrade cost is a real consideration because of the dependency surface. pyproject.toml pins torch per platform, with a different pin for Intel macOS, Apple silicon and everything else, and pixi.lock pins cudnn and cudatoolkit ranges. requirements.txt is autogenerated by uv export, which means it is a derived artifact and not the place to make manual edits. The repository also ships upgrade.py and upgrade_codes/, which suggests an in-place upgrade path exists; the README does not document rollback, so take a copy of your config and character files before running it.
On licensing: the repository's licence is reported as NOASSERTION, which means GitHub could not map the LICENSE file to a known identifier. That is a signal to open LICENSE yourself rather than assume MIT or Apache. The separate LICENSE-Live2D.md governs the Live2D components. This is a description of what the repository contains, not legal advice.
Editorial conclusion
Adopt it if you want a self-hosted voice companion with a Live2D face and you are comfortable editing YAML and reading a FastAPI server log. Do not adopt it if you need a stable API surface, long-term memory, or a project that accepts v1 feature requests: the README states v2.0 is a complete rewrite in early planning and asks people to stop filing feature requests against v1. Before installing, check that your Python is between 3.10 and 3.13 as pyproject.toml requires, and decide whether you are running the Docker image or the uv/pixi path, because the two give you different upgrade stories.
Frequently asked questions
What is Open-LLM-VTuber?
It is a voice-interactive AI companion application that combines an LLM, speech recognition, text-to-speech and a Live2D avatar, with web and desktop client modes. The README states it was started to recreate the closed-source AI VTuber neuro-sama using open-source components that run offline.
How do I install Open-LLM-VTuber?
The README points to the quick-start page on the documentation site, and the repository provides a uv lockfile, a pixi lockfile with conda CUDA packages, a requirements.txt, and a dockerfile with an image on Docker Hub. Python must be between 3.10 and 3.13 according to pyproject.toml.
How do I set up Open-LLM-VTuber?
Configuration is YAML-based: the repository ships config_templates/, from which a conf.yaml is generated, and that is where you select the LLM, ASR and TTS providers. Personas live in characters/ and avatars in live2d-models/, and the README links a Character Customization Guide for both.
How do I use Open-LLM-VTuber?
Start the backend with run_server.py, then open the address it logs in a browser on the same machine, where localhost counts as a secure context and the microphone prompt appears. From there the README describes voice interruption without headphones, camera and screen perception, touch feedback and a desktop pet mode.
Is Open-LLM-VTuber safe?
The README states the app can run completely offline with local models, so conversations stay on your device in that configuration. The dependency list shows cloud provider SDKs such as OpenAI, Anthropic, Groq, ElevenLabs and Azure, so safety depends on which backends you configure rather than on the app alone.
How does AIRI compare with Open-LLM-VTuber?
Open-LLM-VTuber is a Python FastAPI server with a web front end, a separate desktop client, YAML configuration and a Live2D face. AIRI's architecture is not documented in this repository, so a feature-by-feature comparison is not possible from what the project publishes.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/open-llm-vtuber-open-llm-vtuber)