# Voicebox: seven TTS engines, a 4 CPU cap in Docker, and one engine that reads [laugh] aloud

> jamiepine/voicebox is a local-first voice studio: a Tauri desktop shell over a Python backend that clones voices, runs seven TTS engines, dictates into any text field and hands agents a voicebox.speak MCP tool. The friction is in the packaging rather than the model list. Linux has no prebuilt binary, the container build skips type checking on purpose, and the compose file caps the service at 4 CPUs and 8 GB of memory.

**jamiepine/voicebox** — Voicebox is a local AI voice studio for recording, cloning voices, dictation, and speech generation.

- Repository: https://github.com/jamiepine/voicebox
- Website: https://voicebox.sh
- Stars: 55,764 · Forks: 6,947
- Language: TypeScript
- License: MIT
- Published: 2026-08-08 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/jamiepine-voicebox

## Only Chatterbox Turbo reads [laugh] as a sound, the other four read it as text

The engine list is the feature list. Seven engines are switchable per generation and they are not interchangeable. Qwen3-TTS comes in 0.6B and 1.7B sizes across 10 languages and takes delivery instructions in plain language, such as speak slowly. Qwen CustomVoice covers the same 10 languages with 9 curated preset voices and no reference audio at all. LuxTTS is English only and small, about 1GB of VRAM, with 48kHz output and 150x realtime on CPU. TADA is offered at 1B and 3B for 700 seconds or more of coherent audio. Kokoro is an 82M model with 50 preset voices and 8 languages.

The catch is the emotion tags. Chatterbox Turbo, the fast 350M English model, is the only engine that interprets [laugh], [sigh], [gasp] and the rest. Qwen3-TTS, LuxTTS, Chatterbox Multilingual and HumeAI TADA read those brackets literally and speak them. Type / in the text field to open the tag inserter, choose an engine, and the same script either performs or reads out a stage direction.

## A Tauri shell over a Python server, with the JavaScript side almost dependency-free

The runtime is two languages, and the split is visible in the package manifest. The root package.json is a bun workspace with four members, app, tauri, web and landing, and the desktop shell is Tauri in Rust rather than Electron. Its JavaScript dependencies amount to loaders.css and react-loaders, with biome, tailwindcss, typescript and @types/node as dev tooling, and bun is pinned as the package manager at 1.3.8 with an engine requirement of 1.0.0 or later.

The work happens in backend/, which is Python. The dev script starts it with uvicorn:

```bash
uvicorn backend.main:app --reload --port 17493
```

and requirements.txt names what that server needs: uvicorn, fastapi, sqlalchemy, torch, torchvision, soundfile, librosa, python-multipart and huggingface_hub. The app also exposes a REST API and a built-in MCP server, and the agent side is one tool call, voicebox.speak, which any MCP-aware client can invoke. A .mcp.json file sits in the repository root, so the configuration for talking to it is part of the source rather than something you write yourself. Voice personalities attach a free-form persona to a voice profile, and Compose, Rewrite and Respond run through a bundled local LLM, with the same modes reachable over MCP.

## macOS and Windows get a binary, Linux gets a compiler

The download table is four rows and one of them is a command. macOS on Apple Silicon takes a DMG, macOS on Intel takes a different DMG, Windows takes an MSI, and Docker is started with:

```bash
docker compose up
```

All prebuilt binaries are linked from the releases page. Linux is the gap, stated plainly: pre-built binaries are not yet available, and the project points to voicebox.sh/linux-install for build-from-source instructions. That is the single largest difference between platforms in this project, and it lands on the platform where running from source is already the normal way to work.

The Docker route has a second form for AMD cards, an overlay compose file applied on top of the base one:

```bash
docker compose -f docker-compose.yml -f docker-compose.rocm.yml up --build
```

The base file is the CPU build, which is the default, and the overlay is what switches PyTorch over to ROCm wheels. That overlay carries a version default of 6.3, aimed at RDNA1, RDNA2 and RDNA3, with a note that RDNA4 needs ROCm_VERSION set to 7.2 instead. Get that wrong and the container builds cleanly and runs on the CPU.

## The compose file moves the host port to 17600 so the desktop app can keep 17493

The port mapping is the most instructive line in docker-compose.yml, because the comment explains the collision it solves: the host side moved to 17600 so the dev or installed Voicebox can keep 17493, while the container still listens on its native port internally. So the desktop app talks to 17493 and the container is reached on 17600, and anything you write against one port has to be told which.

The rest of the service definition is where a self-hosted install lives or dies. Three volumes are declared: the host directory ./output is bind-mounted to /app/data/generations for generated audio, a named volume voicebox-data sits at /app/data for profiles, database and cache, and a huggingface-cache volume sits at the container's cache path so model weights are not downloaded again on every rebuild. Two environment variables are set, LOG_LEVEL and NUMBA_CACHE_DIR.

And then the limits:

```yaml
    deploy:
      resources:
        limits:
          cpus: '4'
          memory: 8G
```

Four CPUs and 8 GB is a hard ceiling for the CPU build. The LuxTTS and Kokoro entries in the engine table advertise fast CPU inference, and this is the budget they get. A generation that would run on a workstation can fail here, and nothing in the file tells you which model to blame.

## The container build runs vite directly because tsc has pre-existing errors

The Dockerfile is a three-stage build and stage one contains an admission. The frontend is built from oven/bun:1, package.json is normalized with sed to strip carriage returns, the tauri and landing workspaces are deleted from the manifest, and a trailing comma is patched, all because the build needs a package.json that only contains the web workspace. Then it installs and builds:

```dockerfile
RUN cd web && bunx --bun vite build
```

The comment above that line explains what is missing: tsc is skipped because upstream has pre-existing type errors. So the image you run from Docker Hub was built without type checking, while the repository's own check script runs bunx tsc over the app and web projects. Two different definitions of correct are in play, and the container is on the looser one.

The Python stage is python:3.11-slim, and it copies backend/requirements.txt rather than the requirements.txt at the root, so the two files are not the same list. There is also an ARG for the PyTorch variant defaulting to cpu, and a separate ROCm version argument, both of which have to be passed at build time for a GPU image.

## requirements.txt pins nothing, and the model weights arrive at first run

The Python dependency list is nine names with no version constraints: uvicorn, fastapi, sqlalchemy, torch, torchvision, soundfile, librosa, python-multipart and huggingface_hub. Nothing is held still. A working environment today is a coincidence of what the resolver picked, and there is no lock file for the Python side to make it reproducible. The JavaScript side is the opposite, with biome pinned to 2.3.12 and a bun.lock in the repository.

Two of those names decide what happens on a fresh machine. huggingface_hub is how the model weights arrive, and the compose file reserves a named volume at the container's huggingface cache path specifically so the weights survive a rebuild. NUMBA_CACHE_DIR is pointed at /tmp for the same kind of reason, keeping compiled artifacts out of the mounted volume. Neither is a network call your app makes at inference time, but both mean the first start of a fresh install is a download, and the app's promise that models and voice data never leave your machine holds after that fetch rather than before it.

For a local install the same applies. The first generation pulls weights, and the troubleshooting guide the project links from the download section exists for exactly that class of first-run failure, along with GPU detection problems.

## Cloning a voice and saying who consented are two different problems

The privacy claim is specific: models, voice data and captures never leave your machine, and the app is described as a free and open-source alternative to ElevenLabs and WisprFlow, the two cloud products that sit on opposite halves of the same loop, one on output and one on input. Bridging them is the bundled local LLM, used for text refinement and for per-profile personas rather than for speech.

What the app cannot do is decide whose voice you are allowed to clone. Zero-shot cloning takes a reference sample of a few seconds, and the 50-plus preset voices come from Kokoro and Qwen CustomVoice, but nothing in the feature list asks about consent. The repository ships a RESPONSIBLE_USE.md alongside SECURITY.md and CONTRIBUTING.md, which is the project acknowledging that the capability has an ethical edge, and it is the file to read before you point the tool at a recording of a colleague.

The agent route raises the stakes slightly. One MCP tool call is enough for Claude Code, Cursor or Cline to speak to you in a cloned voice, which means the voice profile is now reachable by whatever agent is running on your machine rather than only by the app you installed it in.

## v0.5.0 shipped in April and the branch has moved since

Release cadence and commit activity have drifted apart. The tagged releases are v0.4.4 on 2026-04-21, v0.4.5 on 2026-04-22 and v0.5.0 on 2026-04-25, and package.json still reads 0.5.0, with a .bumpversion.cfg at the root for the version bump. The last push to main was on 2026-08-09, four months later, and the repository is not archived. So main is ahead of every published binary, and a Linux user following the build-from-source route is compiling something no release has shipped.

The rest of the tooling is small and conventional: biome for lint, format and check, a justfile beside the shell scripts, scripts/ for the build, release preparation, API generation and icon updates, and separate web and landing builds. The generate:api and generate:keys scripts cover the two artefacts that would otherwise be hand-made, the API surface and the Tauri signing key.

Licensing is MIT, which covers the code and the preset voices' surrounding tooling but says nothing about the reference audio you bring in. That is the boundary worth holding on to: the licence is permissive about the software, and the responsibility for the voice is yours.

## Conclusion

Voicebox suits a developer on macOS or Windows who wants speech generation, dictation and an agent voice in one local process, and who is willing to pick an engine per job rather than expect one model to do everything. Do not expect the prebuilt path to cover your platform: Linux has no binary yet, and the ROCm overlay defaults to ROCm 6.3, which is the wrong default on RDNA4. Before you build anything on it, check the newest tag against your checkout, since v0.5.0 was published on 2026-04-25 and the last push was on 2026-08-09, and read RESPONSIBLE_USE.md before cloning a voice that is not yours.

## FAQ

### Is voicebox free?

Yes, the project is MIT licensed and describes itself as a free and open-source alternative to ElevenLabs and WisprFlow in one app. The app itself carries no licence fee, while the model weights are downloaded from Hugging Face on first use rather than bundled.

### what is voicebox ai

Voicebox is a local-first AI voice studio that clones a voice from a few seconds of audio, generates speech through seven switchable TTS engines across 23 languages, and dictates into any text field with a global hotkey. It is built as a Tauri application in Rust over a Python backend, and it exposes both a REST API and a built-in MCP server.

### how to install voicebox

macOS on Apple Silicon and macOS on Intel take a DMG, Windows takes an MSI, and Docker is started with docker compose up. Prebuilt binaries for every platform are on the releases page, except Linux, where the project states that pre-built binaries are not yet available and points to build-from-source instructions at voicebox.sh/linux-install.

### how to set up voicebox

For a container install, the base compose file is the CPU build and the AMD overlay is applied on top with docker compose -f docker-compose.yml -f docker-compose.rocm.yml up --build. The service publishes the host port 17600 to the container's 17493, mounts an output directory for generated audio, and stores profiles, database and cache in a named volume that survives restarts.

### what is voicebox.sh

voicebox.sh is the project's site, and it is where the platform downloads live, including separate DMG builds for Apple Silicon and Intel and an MSI for Windows. It also hosts the documentation at docs.voicebox.sh and the Linux build-from-source instructions, while the source itself is on GitHub under an MIT licence.

## Sources

- [Official documentation](https://voicebox.sh)
- [Official README](https://github.com/jamiepine/voicebox#readme)
- [Project repository](https://github.com/jamiepine/voicebox)
- [Release notes](https://github.com/jamiepine/voicebox/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/jamiepine-voicebox
