Model or dataset
elevenyellow/handcrafted-persona-engine avatar
elevenyellow/handcrafted-persona-engine

Persona Engine: a Windows-only Live2D VTuber stack with an LLM brain

An AI-powered interactive avatar engine using Live2D, LLM, ASR, TTS, and RVC. Ideal for VTubing, streaming, and virtual assistant applications.

1,374 stars163 forksC#License varies

At a glance

What is it?
Persona Engine wires Whisper, an LLM, Kokoro TTS and RVC into a Live2D avatar on a CUDA-only Windows build. It is the right tool if you already own an NVIDIA GPU and want a streaming-ready character; it is the wrong tool on a laptop or a Linux box.
Who is it for?
Adopt Persona Engine if you stream from a Windows x64 machine with an NVIDIA CUDA GPU and want the ASR, LLM, TTS and lip-sync pipeline assembled for you. Do not adopt it for a CPU-only box, an AMD or Intel GPU, Linux, or any deployment where you cannot ship a 16 GB model download.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository last received commits 134 days ago.
What is it written in?
Mainly C#, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Persona Engine actually solves, and for whom

Building a talking Live2D character normally means stitching together four or five separate projects: a speech recognizer, a language model client, a text-to-speech engine, a voice-conversion model, and something that drives the avatar's mouth and expressions. Each has its own runtime, its own model files, and its own way of failing. Persona Engine's pitch is that this assembly is the product.

The README describes it as an "AI-driven voice, animation, and personality stack for your Live2D character." The intended user is a streamer or VTuber who wants a character that listens through a microphone, answers through an LLM guided by a personality file, speaks back with real-time TTS that can optionally be voice-cloned, and animates in sync. A second audience is anyone building a virtual assistant on top of the same loop.

The constraint is stated up front and it is not negotiable in this build: Windows x64 and an NVIDIA GPU with CUDA. The README says ASR, TTS and RVC all run on CUDA through ONNX Runtime, and that CPU, AMD and Intel are not supported. That single sentence decides most adoption questions before anything else is considered.

A bundled model named Aria is rigged for the engine's lip-sync and expression pipeline out of the box, so a first run does not require you to bring your own Live2D asset. The repository also ships a Live2D.md integration guide for replacing it.

The pipeline: Whisper in, LLM in the middle, Kokoro and RVC out

The data flow follows the description closely. Audio arrives from the microphone and is transcribed by Whisper, whose size depends on the install profile: Tiny, Small, or Large-v3 Turbo. The transcript goes to an LLM, which the README says is guided by a personality file. The reply is synthesized by Kokoro TTS, optionally passed through RVC for voice cloning, and used to drive the Live2D avatar's lip-sync and expressions.

Two output paths exist. The README describes a built-in transparent overlay for watching the character locally, and a Spout output for piping the avatar into OBS for streaming. Spout is a Windows texture-sharing mechanism, which is consistent with the Windows-only framing.

Lip-sync has three tiers across the profiles: VBridger in the lighter profiles, and VBridger plus Audio2Face in the heaviest one. The voice tier similarly escalates from Kokoro alone to Kokoro plus "Qwen3 expressive."

The design decision worth flagging is the LLM. The README states that Persona Engine "feels most natural with a fine-tuned LLM trained on the engine's communication format," and that standard OpenAI-compatible endpoints such as Groq, OpenAI and Ollama work but require care in personality.txt. That means the engine is not a thin wrapper over any model you like. There is a formatting contract between the LLM and the engine, and the fine-tuned model is distributed through Discord rather than through the repository or the releases page. If you plan to use a hosted API, budget time for prompt work that a user of the fine-tuned model would skip.

Installing Persona Engine and getting to a first pixel

The project ships as a prebuilt zip rather than as a source build. The README's getting-started section says to download PersonaEngine-<version>-win-x64.zip from the Releases page, extract it somewhere with at least 16 GB free, and double-click PersonaEngine.exe. On first launch the bootstrapper asks you to pick an install profile, then downloads, hash-verifies and installs the models and the NVIDIA runtime into a Resources/ folder next to the executable.

The repository lists three profiles. Try it out uses Whisper Tiny and Kokoro. Stream with it uses Whisper Small and Kokoro. Build with it uses Whisper Large-v3 Turbo, Kokoro plus Qwen3 expressive, and VBridger plus Audio2Face, with an approximate download of 16 GB.

If you pick the Build profile, the README warns that the UI defaults still point at the lighter models. You have to change three things manually: set the Voice panel mode to Expressive (Qwen3), pick the Accurate Whisper template in the Listening panel, and enable Audio2Face in the Avatar panel.

To redo the profile choice later, the README gives this command:

bash
PersonaEngine.exe --reinstall

Running it reopens the picker. The README documents several other flags, including --profile=try|stream|build to skip the picker, --repair to re-download anything that fails hash verification, --verify to re-hash installed assets and report mismatches without downloading, --offline to refuse all network access, --non-interactive to treat any prompt as fatal, and --skip-gpu-check to bypass the GPU capability gate, which the README marks as not recommended.

The upgrade path from a pre-installer build is a documented break. The asset directory layout changed when the in-app installer landed, so existing Resources/Models/ and Resources/Live2D/Avatars/ trees are ignored and the installer re-downloads into new locations on first launch. The README tells you to free roughly 16 GB before starting and to delete the old folders once the bootstrapper finishes. There is no documented rollback if the new install fails partway.

Where Persona Engine is the wrong tool

The hardware gate is the first and largest limitation. A CUDA GPU from NVIDIA is required for ASR, TTS and RVC, and the README states plainly that CPU, AMD and Intel are not supported. There is a --skip-gpu-check flag, but the README describes it as not recommended, which reads as a diagnostic escape hatch rather than a supported mode. If your target machine is a Mac, a Linux server, or a Windows laptop with integrated graphics, this project does not run there.

Disk is the second constraint. The Build profile is approximately 16 GB of downloads, and the README repeatedly asks you to have at least that much free before extraction. That is before any Live2D models you bring yourself.

Quality is the third. The README is explicit that the engine depends on a personality file and that a fine-tuned LLM produces the most natural results. A generic model will answer, but the README frames the fine-tuned variant as the intended pairing and points to Discord to obtain it. Nothing in the repository indicates a hosted service that removes this setup burden.

The fourth issue is maintenance. The last push to the repository was on 2026-05-20, and the most recent release listed is v3.0.2 from 2026-04-23. That is a real gap of several months, and the README does not describe a support commitment or a release cadence. Treat the project as something you run at a known-good version rather than something you update on a schedule.

Finally, the licence. The repository metadata does not identify a licence, and the README's badge row includes a license badge but the text does not state terms. Before shipping anything commercial on top of the bundled Aria model or the downloaded weights, check the actual files in the repository rather than assuming.

How it compares to assembling the parts yourself

The closest alternative in practice is not another all-in-one engine but the do-it-yourself route: run Whisper through a local server, point an LLM client at Ollama or a hosted API, drive a TTS engine, and connect the result to a Live2D renderer with something like VBridger. That path works on Linux and on AMD hardware, because each component can be chosen for the platform. It also means you own every integration bug, every model download, and every version mismatch.

Persona Engine's trade is the opposite: it collapses the integration into one executable with a profile picker and a hash-verifying installer, and in exchange it fixes the platform to Windows x64 with an NVIDIA GPU and fixes the model choices to the ones the profiles name. The Konami-style comparison that search results imply, treating this as a game engine, is a naming collision rather than a real alternative; Persona Engine is an avatar stack, not a game runtime.

A second real alternative is a hosted interactive-avatar service that runs the ASR, LLM and TTS on someone else's GPUs and streams video back to you. That removes the CUDA requirement and the 16 GB download entirely, at the cost of per-minute billing, network latency in the conversation loop, and sending your microphone audio off your machine. Persona Engine's local pipeline keeps audio and the LLM conversation on your own hardware, which is the reason to accept the Windows and CUDA constraints in the first place.

Configuration, licensing and the cost of keeping it running

The repository's top level carries CONFIGURATION.md, INSTALLATION.md and Live2D.md alongside the README, which suggests the documented surface is wider than the README alone. The README points at INSTALLATION.md for the full profile walkthrough and at Live2D.md for bringing your own model. The personality file format is the piece to read first, since the README treats personality.txt as the main lever on conversational quality and ships personality_example.txt as a template.

Upgrade cost is mostly bandwidth and disk. The installer re-downloads and hash-verifies assets, and the documented migration from pre-installer builds wipes the old directory layout in favour of new locations. There is a --verify flag to re-hash installed assets and report mismatches without downloading, and a --repair flag to re-fetch anything that fails verification. Those two flags are the practical tools for keeping an install healthy when the network or a mirror misbehaves.

On licensing, the repository metadata does not state a licence and the README does not spell out terms for the bundled Aria model, the Whisper, Kokoro, Qwen3 or RVC weights, or the NVIDIA runtime the installer pulls down. Each of those components may carry its own terms. This is not legal advice, but the concrete step is to read the licence files that ship with the repository and with each downloaded asset before any commercial stream or product depends on them.

Editorial conclusion

Adopt Persona Engine if you stream from a Windows x64 machine with an NVIDIA CUDA GPU and want the ASR, LLM, TTS and lip-sync pipeline assembled for you. Do not adopt it for a CPU-only box, an AMD or Intel GPU, Linux, or any deployment where you cannot ship a 16 GB model download. Before committing, run PersonaEngine.exe --verify against a finished install and read CONFIGURATION.md for the personality.txt format, because the README treats that file as the main quality lever and does not document rollback for a failed upgrade.

Frequently asked questions

Does Persona Engine run on Linux or on an AMD GPU?

No. The README states that ASR, TTS and RVC all run on CUDA through ONNX Runtime and that CPU, AMD and Intel are not supported, with Windows x64 and an NVIDIA GPU as the requirement. A --skip-gpu-check flag exists but the README marks it as not recommended.

How much disk space does Persona Engine need?

The README asks for at least 16 GB free before extraction, and the Build with it profile is listed at approximately 16 GB of downloads. The lighter Try it out and Stream with it profiles download less, though the README does not give exact figures for them.

How do I reinstall or repair a Persona Engine install?

Run PersonaEngine.exe --reinstall to reopen the profile picker. To fix a failed asset download, --repair re-downloads anything that fails hash verification, and --verify re-hashes installed assets and reports mismatches without downloading anything.

Do I have to use the fine-tuned LLM with Persona Engine?

No, but the README says the engine feels most natural with a fine-tuned LLM trained on its communication format. Standard OpenAI-compatible models such as Groq, OpenAI and Ollama work, and the README advises putting care into personality.txt when using them.

Why does the Build profile still sound like the light models?

The README warns that the Build profile downloads the larger models but the UI defaults keep the lighter ones active. You have to set the Voice panel mode to Expressive (Qwen3), pick the Accurate Whisper template in the Listening panel, and enable Audio2Face in the Avatar panel.

Official sources

  1. elevenyellow/handcrafted-persona-engine on GitHub
  2. Issues
  3. README
  4. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/elevenyellow-handcrafted-persona-engine.svg)](https://hysenlabs.com/projects/elevenyellow-handcrafted-persona-engine)