Model or dataset
fluxions-ai/vui avatar
fluxions-ai/vui

fluxions-ai/vui: a self-hosted real-time voice assistant built around Vui Nano TTS

Real-time voice assistant — WebRTC streaming, faster-whisper ASR, local LLM, Vui Nano (300M) TTS. OpenAI Realtime API compatible. Voice cloning, barge-in, ~9× realtime on a 4090. Apache 2.0.

769 stars84 forksPythonNOASSERTION

At a glance

What is it?
Vui is Fluxions AI's open core: a single Python server that runs WebRTC audio in, faster-whisper or Moonshine ASR, a local LLM, and a 300M streaming TTS model back out. It is OpenAI Realtime API compatible, and its streaming server is designed for Linux with an NVIDIA GPU.
Who is it for?
Adopt Vui if you want a voice loop you can run on your own hardware, you already have Ollama and an NVIDIA GPU on Linux, and you accept that the streaming server is the open core while turn-taking improvements live in the paid API. Do not adopt it if you need a documented production support contract, a Windows-native path, or Apple Silicon streaming today, since the README describes the MLX streaming glue as WIP.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 19 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem Vui solves, and who it is actually for

Most speech stacks are assembled from parts that were never designed to run together. You pick an ASR model, wire it to an LLM client, then bolt on a TTS engine, and the seams show up as latency: the user finishes a sentence, the transcript waits for a final segment, the LLM waits for the full transcript, and the TTS waits for the whole reply. Vui's pitch is that the whole loop ships as one server with the seams already closed. The README describes a pipeline where the transcript streams, the LLM is prefilled speculatively while you are still speaking, and the TTS emits sentence-level chunks with backpressure.

The intended audience is developers who want voice in an application without sending audio to a hosted API. The repository targets Linux with an NVIDIA GPU for the streaming server, and the install one-liner clones into ~/vui and launches the stack on localhost:8080. A second audience exists for the model alone: the PyPI distribution is named vui-tts, and the pyproject.toml notes that `pip install vui-tts` gives the inference engine only, which is what integrations such as Pipecat and LiveKit depend on. That split matters. If you only want a TTS model with voice cloning, you do not need the assistant server at all.

How the ASR, LLM and TTS loop is wired together

The architecture visible in the README is a WebRTC plus WebSocket pipeline. Audio enters over WebRTC, VAD decides when a turn has ended, and the ASR stage runs either faster-whisper on GPU or Moonshine on CPU via ONNX. The LLM stage is pluggable: Ollama, vLLM, or any OpenAI-compatible endpoint. The TTS stage is Vui Nano, described as a Llama-style decoder with an RQ-Transformer head over the Qwen3-TTS-12Hz codec, running bf16 inference with CUDA graphs.

Two design choices stand out. The first is barge-in: the README states that if you start talking mid-reply, the model cancels and listens. That is a turn-taking decision, not a model capability, and it is the kind of thing that is easy to claim and harder to get right when the TTS is mid-sentence. The second is the thoughts stream, a parallel LLM route that maps voice intent to roughly fifteen tools (memory operations, task control, timers, web search, delegation) without a wake-word grammar. The README says it is pluggable for your own local tools. That is a more interesting claim than the feature list suggests, because routing intent without a fixed grammar usually means the routing prompt is doing real work and will need tuning.

State persists in two places worth knowing about. Memories are written to ~/.vui/memories.json, so the assistant accumulates facts about a user across sessions. Model selection is hot-swappable from the UI, which means the LLM and ASR backends can change without a restart.

Installing Vui with docker compose and getting a first reply

The README offers a one-liner installer that clones into ~/vui, auto-detects Docker versus native, installs uv, Ollama, ffmpeg and the Claude Code CLI, pulls the model, and launches on port 8080. Flags such as --docker, --native, --no-claude, --upgrade, --model <name> and --dry-run forward to install.sh. If you would rather control the steps, the compose path is the documented recommendation and assumes Ollama on the host.

The prerequisite is the NVIDIA Container Toolkit, because the container needs to see the GPU. The README gives these commands for Debian and Ubuntu:

bash
sudo apt install -y nvidia-container-toolkit
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker

Before going further, confirm the GPU is visible to Docker. The README suggests this check, and it is worth running because a silent failure here produces a container that starts and then cannot load the model:

bash
docker run --rm --gpus all nvidia/cuda:13.0.0-base-ubuntu22.04 nvidia-smi

With the toolkit working, pull the LLM on the host and bring up the stack. The compose file's own header repeats these two commands, so they are the canonical first run:

bash
ollama pull qwen3.5:4b
docker compose up -d

Open http://localhost:8080, allow microphone access, and start talking. The README states that the Vui checkpoint and the Qwen codec download automatically from Hugging Face on first run and persist in a named volume, so the first start is slower than later ones. If you do not have Ollama on the host, the bundled service is gated behind a profile and is off by default:

bash
docker compose --profile ollama up -d
docker compose exec ollama ollama pull qwen3.5:4b

The Vui container talks to whichever Ollama is listening on localhost:11434, so the two paths are mutually exclusive in practice. One caveat from the compose file: the claude-task service starts by default, and without auth (a ~/.claude login or ANTHROPIC_API_KEY) it stays up but task requests fail at runtime. The README is explicit that the voice loop is unaffected in that case.

Where Vui gets in the way: the open-core line and unpublished weights

The README states plainly that this repository is the open core and that the production API ships ongoing model updates and a more advanced turn-taking system. That is a real boundary, not marketing noise. If your evaluation depends on turn-taking quality, you are evaluating the weaker of the two systems, and the better one is not in this repository.

The compose file carries a sharper constraint. Its vui-stream service overrides the default command to load a local checkpoint instead of fetching vui-nano.safetensors from huggingface.co/fluxions/vui, with the comment that the file is not yet published. The same comment says to drop the override once the HF repo has the v2 weights. So a fresh clone may not behave the way the feature list implies without that local checkpoint, and the README's claim that weights download automatically on first run sits in tension with the compose comment. That is the first thing to check before you plan a deployment around it.

Platform support is uneven in a way that is easy to miss. The streaming server is designed for Linux plus an NVIDIA GPU. Apple Silicon is supported through the Engine Python API, which auto-dispatches to an MLX backend with quantized vui-190k weights, and the README gives a figure of roughly 1.5 to 2.7 times real time on an M4 for that path, with demo.py and demo.py --render working end to end. The streaming-server MLX glue is described as WIP. If your team is on Macs, you get the model, not the assistant. Windows is not addressed at all in the README.

How Vui differs from Pipecat and LiveKit Agents

The obvious comparison is with voice-agent frameworks such as Pipecat and LiveKit Agents. Those are orchestration layers: you bring an ASR provider, an LLM provider and a TTS provider, and the framework handles transport and turn-taking. Vui inverts that. It ships a specific model, Vui Nano, and the server exists to serve it, with the ASR and LLM left pluggable. The pyproject.toml even notes that Pipecat and LiveKit integrations depend on the vui-tts inference package, so the relationship is closer to component than competitor.

The practical difference is where you spend your effort. With an orchestration framework, the work is in choosing and paying for providers, and the ceiling is set by whichever hosted TTS you pick. With Vui, the work is in hardware and local model management, and the ceiling is set by a 300M-parameter model running on your own GPU. The README gives a streaming TTS figure of roughly 9 times real time on a 4090 with bf16 and CUDA graphs. That number is the project's own claim and is not independently reproduced here, but it is the kind of figure that decides whether a single GPU can serve a given number of concurrent conversations.

The other difference is the API surface. Vui exposes a ws:// endpoint at /v1/realtime that the README describes as a drop-in for clients written against OpenAI's Realtime spec, documented in docs/realtime-api.md. There is also a one-shot REST endpoint, POST /v1/voice-note, that runs the whole ASR to LLM to TTS pipeline in a single HTTP call with audio in and JSON out. If your client code already speaks the OpenAI Realtime protocol, that compatibility is worth more than any feature comparison.

Licence, upgrades and what maintenance looks like

The repository is not archived, and the last push was on 2026-09-02. The two releases listed are v0.1.0 (Vui 100M) and v1.0.0 (Vui Nano plus streaming voice assistant), both dated 2026-05-14. The pyproject.toml declares version 1.1.4 and a license of Apache-2.0, and the README's badge row points to a Hugging Face model page and a Discord. Note that the repository metadata reports the licence as NOASSERTION while pyproject.toml states Apache-2.0; if licence terms matter to your legal review, read the LICENSE file rather than trusting either field. Nothing here is legal advice.

Upgrade cost is shaped by two things. First, the installer has an --upgrade flag, and the compose path rebuilds from docker/Dockerfile.stream, so staying current is mostly a pull and rebuild. Second, the model weights live in a named volume and, per the compose comments, may need to be swapped manually while the v2 weights are unpublished. That is the upgrade step most likely to surprise you. The pyproject.toml also documents a deliberate dependency constraint: torch is a range rather than a pin because torch 2.11 and 2.12 declare setuptools<82, which collides with frameworks that pin newer setuptools, while the repository's own lock stays on release-tested torch 2.11 through constraint-dependencies. If you embed vui-tts inside another framework, that constraint is the thing to read before you file a dependency-conflict bug.

What to check before you commit to Vui

The strongest reason to pick Vui is the OpenAI Realtime API compatibility combined with local execution. If you have a client written against that spec and you want the audio to stay on your hardware, pointing it at ws://.../v1/realtime is a smaller change than rebuilding around a different protocol. The one-shot POST /v1/voice-note endpoint covers the simpler case where you just want audio in and JSON out.

The strongest reason to look elsewhere is operational. This is a single-server design aimed at one machine with one GPU, and the README does not document horizontal scaling, failover, or rollback of a model upgrade. The README is also silent on authentication for the streaming server, so exposing port 8080 beyond localhost is something you would have to reason about yourself; the mobile section instead points at cloudflared and Tailscale for phone access with mic over HTTPS, which suggests the intended posture is a private tunnel rather than a public endpoint.

Between those two poles, the deciding question is whether you are comfortable running inference hardware. Vui rewards teams that already have Ollama running and a 4090-class GPU idle. It frustrates teams that want a managed endpoint with a support contract, because the README directs those users to the production API at fluxions.ai instead. That is not a defect in the open repository; it is the stated shape of the project.

Editorial conclusion

Adopt Vui if you want a voice loop you can run on your own hardware, you already have Ollama and an NVIDIA GPU on Linux, and you accept that the streaming server is the open core while turn-taking improvements live in the paid API. Do not adopt it if you need a documented production support contract, a Windows-native path, or Apple Silicon streaming today, since the README describes the MLX streaming glue as WIP. Before committing, verify three things on your own machine: that `docker run --rm --gpus all nvidia/cuda:13.0.0-base-ubuntu22.04 nvidia-smi` sees your GPU, that your Ollama instance answers on localhost:11434, and that the Hugging Face checkpoint download completes, because the compose file notes the v2 weights are not yet published there.

Frequently asked questions

What is fluxions-ai/vui?

It is a real-time voice assistant that runs as a single Python server: WebRTC audio in, faster-whisper or Moonshine ASR, a local LLM, and streaming TTS from Vui Nano, a 300M speech transformer based on Qwen3 TTS. The README also describes it as OpenAI Realtime API compatible.

Is Vui free or paid?

The repository is the open core and the pyproject.toml declares Apache-2.0, so the code and the Vui Nano model are available at no cost. The README separately points to a production API at fluxions.ai that ships ongoing model updates and a more advanced turn-taking system.

How do I install fluxions-ai/vui?

The README gives a one-liner installer, curl -fsSL https://install.fluxions.ai | bash, which clones into ~/vui and launches the stack on http://localhost:8080. The recommended alternative is docker compose: pull qwen3.5:4b with Ollama on the host, then run docker compose up -d.

Does Vui work on Apple Silicon or Windows?

Apple Silicon is supported through the Engine Python API, which auto-dispatches to an MLX backend, and the README states demo.py works end to end there while the streaming-server MLX glue is WIP. The README describes the docker compose stack as designed for Linux plus an NVIDIA GPU, and it does not address Windows.

Does Vui support voice cloning?

Yes. The README lists voice cloning from an uploaded audio sample and states that four fine-tuned presets ship: maeve, abraham, rhian and harry.

Official sources

  1. fluxions-ai/vui on GitHub
  2. Issues
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/fluxions-ai-vui.svg)](https://hysenlabs.com/projects/fluxions-ai-vui)