Model or dataset
kyutai-labs/unmute avatar
kyutai-labs/unmute

Unmute: Adding Voice to Any Text LLM with Kyutai Speech Models

Make text LLMs listen and speak

1,518 stars249 forksPythonMIT

At a glance

What is it?
kyutai-labs/unmute is an open-source system that wraps any text LLM in Kyutai's speech-to-text and text-to-speech models, creating a low-latency voice assistant you can self-host. It is designed for engineers who want full control over the model stack rather than a managed real-time voice API.
Who is it for?
Unmute is for engineers who need a self-hosted voice LLM pipeline with control over the underlying speech models and latency budget. It is not a viable option for macOS users, developers without CUDA hardware, or teams that need the system running on ARM servers.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 20 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 22, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The Problem Unmute Solves: Voice for Any Text LLM

Text LLMs have no native ability to process spoken audio or generate speech. Building a voice interface around one typically requires integrating three separate systems: a speech-to-text model, the LLM itself, and a text-to-speech model. Getting those three to work together in real time, at low latency and over a websocket, involves non-trivial orchestration work.

Unmute packages that orchestration into a deployable system. The README describes it as a system that allows text LLMs to listen and speak by wrapping them in Kyutai's speech-to-text and TTS models. Both the STT and TTS components are optimized for low latency. The system works with any text LLM the user configures, including self-hosted models served via VLLM or cloud-hosted ones via OpenRouter.

The project is maintained by Kyutai Labs, the same team that built the underlying speech models. A pre-print about those models is cited in the README.

The Three-Component Pipeline

Unmute's architecture connects five services: a frontend, a backend, a speech-to-text server, an LLM server, and a text-to-speech server.

The user opens the Unmute web interface served by the frontend. Clicking connect establishes a websocket connection to the backend. The backend relays audio from the user's microphone to the STT server, which transcribes it in real time. Once the STT detects that the user has finished speaking, the backend sends the transcript to the LLM server and streams the response tokens back. As the LLM streams tokens, the backend feeds them to the TTS server, which generates audio and forwards it to the user's browser.

All three model services communicate with the backend over websockets, using the environment variables `KYUTAI_STT_URL`, `KYUTAI_TTS_URL`, and `KYUTAI_LLM_URL`. In the Docker Compose deployment these resolve to internal container hostnames: `ws://stt:8080`, `ws://tts:8080`, and `http://llm:8000`.

The default local LLM is Gemma 3 1B hosted on Hugging Face. In production at unmute.sh, Kyutai runs GPT OSS 120B via OpenRouter. The README notes this explicitly, giving users a reference point for what scale of model the hosted version uses.

Hardware and OS Requirements Before You Start

The hardware requirements are firm. Unmute requires a GPU with CUDA support and at least 16 GB of VRAM. The architecture must be x86_64; aarch64 is not supported, and the README states that no aarch64 support is planned. Running without a GPU is not a documented option.

The supported operating systems are Linux and Windows with WSL. The README explicitly states that running on macOS is not supported, and links to an open issue. This is a hard boundary, not a warning about degraded performance.

For Windows users, the README links to WSL installation instructions. The WSL path works because the NVIDIA Container Toolkit can access the GPU through WSL2.

The first setup check is verifying that the Nvidia Container Toolkit is installed and that Docker can access the GPU:

bash
sudo docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smi

If that command prints GPU information, the prerequisite is met. If it fails, the Unmute Docker Compose will not start the STT or TTS services.

Getting Started with Docker Compose

Docker Compose is the recommended deployment method. The README describes it as very easy compared to the Dockerless and Docker Swarm options.

The default LLM is Gemma 3 1B, which requires accepting a license on Hugging Face. Before starting the compose stack, create a Hugging Face account, accept the model terms, generate a read-access token, and export it:

bash
echo $HUGGING_FACE_HUB_TOKEN  # This should print hf_...something...

docker compose up --build

The compose file starts all services, including Traefik as a reverse proxy on port 80. The frontend and backend are built from the repository source. If you run into memory issues with the 16 GB VRAM limit, the README advises opening `docker-compose.yml` and checking for `NOTE:` comments that indicate adjustable settings.

Once the stack is up, the interface is accessible at `http://localhost:80`. The pre-print linked in the README describes the speech model architectures in more detail for users who want to understand the latency profile.

Multi-GPU Deployment for Lower TTS Latency

Running all three model services on a single GPU produces a TTS latency of approximately 750ms on an L40S class GPU. The README gives a concrete comparison: on the unmute.sh production system, which runs STT, TTS, and LLM on separate GPUs, TTS latency drops to approximately 450ms on the same GPU class.

To assign each service to its own GPU, the Docker Compose configuration needs a `deploy.resources.reservations.devices` block under the `stt`, `tts`, and `llm` services:

yaml
  stt:
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: 1
              capabilities: [gpu]

The README applies this pattern to all three services. The result is that each model runs on a separate physical GPU with no memory contention between them.

The Docker Swarm deployment path, documented in `SWARM.md`, is how the unmute.sh production site scales to larger setups. The README is explicit that Kyutai does not offer support for debugging Swarm deployments, since multi-node troubleshooting is difficult to assist with remotely.

What Unmute Cannot Do

Several deployment scenarios are outside the documented scope. Unmute does not run on CPU only; the STT and TTS models require CUDA. It does not run natively on macOS or on ARM hardware, two popular developer platforms. Windows without WSL is also not supported.

The HTTPS support for the Docker Compose path is not configured in the default compose file. The README notes that the default setup is HTTP only, and recommends the Docker Swarm path or asks users to modify the file themselves for production HTTPS. This is a meaningful gap for anyone deploying to a server accessible over the internet.

The default LLM is Gemma 3 1B, which is a small model. Swapping it for a larger model requires editing the compose file and potentially requiring more VRAM. The README does not document the exact memory requirements for larger LLMs in the context of running them alongside the STT and TTS services on a single GPU.

The project has no GitHub releases. All deployments pull from main or from a built Docker image. There is no stable release channel or version pinning for the Docker image tag beyond `latest`.

Comparison with Managed Real-Time Voice APIs

The direct alternative is a managed real-time voice API, such as OpenAI's real-time API, which provides STT, LLM, and TTS as a single hosted service. The fundamental difference is infrastructure ownership.

With a managed API, you pay per token or per minute of audio, and the provider handles model hosting, scaling, and latency optimization. You do not need a GPU. The trade-off is that you are tied to the provider's model choices and pricing.

With Unmute, you own the infrastructure and choose the LLM independently. You can run any OpenRouter-hosted model or a locally served VLLM instance. The STT and TTS are Kyutai's own models, optimized for low latency. The cost is a CUDA GPU with 16 GB of VRAM and the engineering effort to maintain the deployment.

For teams working with sensitive audio data who cannot send it to a cloud API, or for researchers who need to experiment with specific LLMs not available through managed voice services, Unmute provides the only documented path to a fully self-hosted voice LLM pipeline in this form.

Editorial conclusion

Unmute is for engineers who need a self-hosted voice LLM pipeline with control over the underlying speech models and latency budget. It is not a viable option for macOS users, developers without CUDA hardware, or teams that need the system running on ARM servers. Before deploying it, confirm that your GPU has at least 16 GB of VRAM and that the Nvidia Container Toolkit is installed and working, since the Docker Compose setup will not start without it. If you want lower TTS latency, run the STT, TTS, and LLM services on three separate GPUs, as the README documents this reduces TTS latency from approximately 750ms to approximately 450ms on comparable hardware.

Frequently asked questions

Does Unmute work on macOS?

No. The README explicitly states that running on macOS is not supported, and links to an open issue on the subject. Unmute requires Linux or Windows with WSL, a CUDA-capable GPU, and an x86_64 architecture.

Can Unmute use any LLM, or only specific models?

Unmute works with any text LLM accessible via an OpenAI-compatible API. The default local setup uses Gemma 3 1B from Hugging Face, and the README notes that the production unmute.sh uses GPT OSS 120B via OpenRouter. You can also serve your own model using VLLM and point Unmute to it with the KYUTAI_LLM_URL environment variable.

What GPU is required to run Unmute?

The README specifies a GPU with CUDA support and at least 16 GB of VRAM. The default configuration with Gemma 3 1B fits within that limit. If you run into memory issues with the default model, the README advises checking NOTE comments in docker-compose.yml for places to adjust the configuration.

Official sources

  1. Issues
  2. kyutai-labs/unmute on GitHub
  3. License: MIT
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/kyutai-labs-unmute.svg)](https://hysenlabs.com/projects/kyutai-labs-unmute)