Open-source project
xikhar/persona avatar
xikhar/persona

Persona: A VRM Desktop Avatar That Reacts to Voice Output for AI Agent Workflows

Bringing real-time voice to life.

955 stars96 forksTypeScriptMIT

At a glance

What is it?
Persona is a cross-platform Electron application that places an animated VRM character on the desktop and drives its speaking animation from the audio output of a selected process, without capturing microphone input, recording audio, or sending data over the network. It connects to AI agents through MCP, allowing those agents to trigger custom animations on the character.
Who is it for?
Persona suits developers building voice-based AI agent pipelines who want a visual character presence that animates in sync with AI audio output, and who are willing to work with a 0.1.0-beta.0 release. It is not suited to streaming or face-tracking workflows; VTube Studio, which drives VRM avatars from webcam face tracking, is the better fit for those.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 29 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

A Character Presence for Desktop Voice Conversations

Persona addresses a specific gap in AI voice agent interfaces: when a text-to-speech model or AI assistant plays audio, there is no visual representation of the speaking entity. Persona sits as a floating VRM avatar on the desktop, animates its speaking motion when audio output is detected, and stays out of the way otherwise through click-through mode. The target user is a developer running a local AI voice pipeline who wants the agent to have a visual presence without building a custom UI.

The application does not interact with the content of the audio. The README states explicitly that Persona does not capture the microphone, save audio, produce speech, transcribe content, or send audio over the network. It is a visual layer only, driven by the presence and level of audio output from a selected process.

How Persona Captures Voice Output on Each Platform

The audio detection mechanism differs by platform, and each has specific prerequisites. On Linux, Persona uses PipeWire playback-stream capture through pw-dump and pw-record, which must be present on the PATH. On Windows, it uses WASAPI process-loopback capture, which requires Windows 10 build 20348 or later. On macOS 14.2 and later, it uses Core Audio process tap and asks once for System Audio Recording permission.

In every case, the listener is scoped to a single automatically detected or user-selected playback process, not to the whole system audio. This means that when the target process is speaking, the avatar animates; when other audio plays on the system, the avatar does not react. The platform-specific capture method is the main reason the macOS minimum is 14.2: the Core Audio process tap API was not available in earlier macOS releases.

The distribution format also varies: AppImage and DEB on Linux, an NSIS installer on Windows, and DMG or ZIP for arm64 and x64 on macOS.

Installing and Running Persona from Source

Persona requires Node.js 24 or newer, npm, and a desktop session with hardware-accelerated graphics. To run it from source, copy the example catalog files and install dependencies:

bash
cp public/assets/library.json.example public/assets/library.json
cp public/assets/manifest.json.example public/assets/manifest.json

Both example files are described in the README as directly usable and as documentation for the catalog format. Their media is test-only and not intended for distribution.

bash
npm install
npm run demo

The npm run demo command builds the renderer and launches Persona with automatic voice-output detection. For a launch in the background:

bash
npm start -- --background

Persona is at version 0.1.0-beta.0 and is marked private in package.json, so it is not published to npm. The full build chain uses Vite for the renderer and TypeScript compilation for the Electron runtime, as visible in the package.json scripts.

VRM Models, VRMA Animations, and the Catalog System

Persona uses the VRM and VRMA formats for character models and animations. VRM is an open format for 3D humanoid avatars used in virtual reality and avatar applications; VRMA is the companion animation format. The application ships with a default character model and default idle and speaking motion clips, so it works out of the box without importing custom files.

Custom models are managed through Settings, where the character library shows installed models, allows preview of each model with its animations, and lets the user choose the packaged default. Packaged VRM files go under public/assets/models/ and packaged VRMA files under public/assets/animations/ in the source tree. A catalog can declare multiple packaged models, and when default_model_id is null, Persona picks the first model record automatically.

The animation system provides Idle and Speaking action slots that begin without assigned clips. Uploading a VRMA file to one of these slots gives it a numbered name such as idle1 or speaking1. Custom actions can also be defined with a name, description, and trigger scenario. Persona adds that metadata to its MCP animation tool so a connected AI agent can understand what the action represents and choose when to trigger it.

Connecting Persona to Codex and Other MCP Clients

Persona exposes a local MCP server on port 47831 that lets any MCP-compatible agent control the character's visual state. The README shows the registration command for Codex:

bash
codex mcp add persona --url http://127.0.0.1:47831/mcp

Once registered, a Codex session can ask Persona to play an installed animation, show or hide its window, and query whether the character and voice listener are active. The README notes that any compatible MCP client can use the same animation tools, and points to docs/INTEGRATIONS.md for broader integration documentation.

The README also describes an external-events mode for pipelines that already compute output audio levels themselves, letting them drive the animation without going through the platform audio capture layer. This is described as an option for advanced users with local model stacks.

Click-Through Mode and Platform Limitations

Click-through lets the avatar float over the desktop while mouse events pass through to the windows behind it. The behavior differs by platform in a meaningful way. On Windows and macOS, only the transparent area around the character passes clicks through; the character mesh itself remains interactive for orbit, zoom, and Alt+drag. On Linux, the whole window passes clicks through at once, which makes the character unresponsive to mouse interaction while click-through is on. The README describes the Linux behavior as experimental and notes it depends on X11 input shapes; no Wayland compositor has been verified.

VTube Studio is a widely used VRM avatar application that takes the opposite approach to driving animation: it uses webcam-based face tracking to animate facial expressions and head movement in real time, and is designed primarily for live-streaming. Persona does not do face tracking at all; it reacts to audio output levels rather than camera input. The two tools have different use cases, and Persona's design is unsuitable for streaming purposes where facial expression capture matters.

Persona also requires a desktop session with hardware-accelerated graphics. Headless or purely server-side deployments will not run the application.

Beta Status, Maintenance, and MIT License

Persona is at version 0.1.0-beta.0, indicating an early stage of development. The last push was on 2026-09-02, less than four weeks before 2026-09-28, and the repository is not archived. The repository includes CHANGELOG.md, CONTRIBUTING.md, CODE_OF_CONDUCT.md, and SECURITY.md, which are signs of an intentional open-source project structure even at this early version.

The license is MIT, as declared in both package.json and the LICENSE file at the repository root. MIT is permissive and allows use, modification, and redistribution with attribution. The beta designation means API and behavior changes between releases should be expected, and there is no documented migration path for the pre-1.0 period.

Editorial conclusion

Persona suits developers building voice-based AI agent pipelines who want a visual character presence that animates in sync with AI audio output, and who are willing to work with a 0.1.0-beta.0 release. It is not suited to streaming or face-tracking workflows; VTube Studio, which drives VRM avatars from webcam face tracking, is the better fit for those. Before adopting, confirm that your platform meets the audio capture prerequisites: PipeWire with pw-dump and pw-record on Linux, Windows 10 build 20348 or newer for WASAPI process-loopback, and macOS 14.2 or later for Core Audio process tap.

Frequently asked questions

Does Persona record or transmit audio?

No. The README states explicitly that Persona does not capture the microphone, save audio, produce speech, transcribe content, or send audio over the network. It listens to the output level of a selected process to drive the speaking animation and does nothing else with the audio.

Can Persona work with local AI voice models?

Yes. The README describes selecting a running local app or Linux playback stream from Settings to use as the voice source. Advanced users can also supply a cross-platform process pattern. Any MCP client, including those connected to local models, can use the animation tools through the local MCP server on port 47831.

What 3D model format does Persona use?

Persona uses VRM for character models and VRMA for animation clips. The application ships with a default model and default idle and speaking animations. Custom VRM and VRMA files can be added through the Settings interface without modifying the installed application.

Official sources

  1. Issues
  2. License: MIT
  3. README
  4. xikhar/persona on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/xikhar-persona.svg)](https://hysenlabs.com/projects/xikhar-persona)