# Chatterbox-TTS-Server installs its engine from a fork at master, with dependencies switched off

> devnen/Chatterbox-TTS-Server puts three Chatterbox models behind an OpenAI-compatible API with a web UI. Its requirements file refuses to install the engine normally, its Dockerfile for every GPU vendor shares a CUDA base image, and its compose file ships a literal token placeholder.

**devnen/Chatterbox-TTS-Server** — Self-host the powerful Chatterbox TTS model. This server offers a user-friendly Web UI, flexible API endpoints (incl. OpenAI compatible), predefined voices, voice cloning, and large audiobook-scale text processing. Runs accelerated on NVIDIA (CUDA), AMD (ROCm), and CPU.

- Repository: https://github.com/devnen/Chatterbox-TTS-Server
- Website: https://colab.research.google.com/github/devnen/Chatterbox-TTS-Server/blob/main/Chatterbox_TTS_Colab_Demo.ipynb
- Stars: 1,461 · Forks: 360
- Language: Python
- License: MIT
- Published: 2026-09-16 · Updated: 2026-09-16 · Language: en
- Canonical page: https://hysenlabs.com/projects/devnen-chatterbox-tts-server

## The engine is installed from a fork with dependencies switched off

The requirements file describes itself as being for CPU-only installation and then says, in capitals, that chatterbox-tts is not included in it. The reason given is dependency resolver conflicts from ONNX source builds and torch version mismatches, and the instruction is to install the engine separately with --no-deps. The command it prints is:

```bash
pip install --no-deps git+https://github.com/devnen/chatterbox-v2.git@master s3tokenizer==0.3.0 onnx==1.16.0
```

Three things are in that line. The engine is not the upstream package but a fork under the project author's own account. It is referenced at a branch named master rather than at a tag, so there is no version to pin and a reinstall resolves to whatever that branch points at. And --no-deps means pip is told not to look at what the engine requires, which is why the file then lists the engine's dependencies by hand. The launcher handles this automatically, so a manual install has to follow the printed order.

## A protobuf ceiling is deliberately overridden to make the install work

The comments in the requirements file describe a dependency conflict in plain terms. onnx requires protobuf at 3.20.2 or newer, while descript-audiotools requires protobuf below 3.20, so the two cannot both be satisfied. The chosen resolution is to install both with dependency resolution disabled and then force protobuf to 4.25.0 or newer afterwards, which is what the Dockerfile does in its final install step. That means the ceiling declared by descript-audiotools is not met by the running environment, and nothing in the file says what breaks if that package does check its own requirement. The same pattern appears elsewhere: fastapi is capped below 0.116.0 with a note that starlette below 1.0 is needed for template response compatibility, so the upper bounds in this project are the result of running into a specific break.

## Seven images for seven GPU stacks, none of them for Apple

The tree carries seven Dockerfiles and seven compose files: a default one, a CPU one, two NVIDIA ones for CUDA 12.8 and CUDA 13.0, two AMD ones for ROCm and for RDNA4, and one for Strix Halo. The release notes add the mapping. The CUDA 13.0 file covers DGX Spark and sm_121 hardware with PyTorch 2.10, while RTX 30, 40 and 50 cards stay on CUDA 12.1 or 12.8. The Strix Halo file carries ROCm 7.2 with a graphics override environment variable. Against that, the README claims acceleration on Apple Silicon through MPS with a CPU fallback, and there is no compose file or Dockerfile named for Apple anywhere in the tree, so on a Mac the documented path is not the container path.

## The default image is a CUDA runtime even when you ask for CPU

The main Dockerfile starts from a CUDA 12.8.1 runtime image on Ubuntu 22.04, then takes a build argument that can be nvidia or cpu and defaults to nvidia. Inside, the base requirements are installed first, which pull CPU-only torch from the PyTorch CPU index, and the CUDA torch set is layered on top only when the argument says nvidia. So the CPU deployment path is the NVIDIA base image with the GPU torch deliberately not installed, and it carries the CUDA runtime libraries in the image either way. The file also installs ffmpeg, libsndfile, build-essential and git, sets the Hugging Face home to a cache directory, creates the working directories with a comment noting a syntax error that was fixed, and starts the server on port 8004.

## The compose file ships a literal token placeholder and a GPU fallback

The default compose file is where most people will start, and it has two details worth reading before running it. The environment block sets the Hugging Face token to the literal text YOUR_TOKEN_HERE, which is a placeholder rather than an empty value, so a copy that is left unchanged will send that string as a credential. It also sets the fast transfer flag, makes all NVIDIA devices visible, and declares compute and utility driver capabilities. GPU access is requested the modern way, through a deploy block reserving one device with the nvidia driver, and a comment explains the alternative for older setups: if a device injection error appears, comment out the deploy section and uncomment the legacy runtime line instead. Four host directories are mounted for configuration, voices, reference audio, outputs and logs, with the model cache on a named volume so it survives rebuilds.

## A path traversal fix and a chunker that mistook dashes for bullets

Two defects in the release notes are worth reading as design information. The first is labelled as a security fix: a path traversal weakness in the voice file parameters of the speech endpoints, now returning HTTP 400 on an attempt. That matters because voice cloning takes a file path as input, and the endpoints at issue are the OpenAI-compatible ones, so the fix lands on the path most integrations will use. The second is less obvious. Stray dashes in narrative text were being treated as bullet items by the chunker, which swallowed the rest of the paragraph. For a server whose main job is long-form audiobook text, that is a silent correctness bug rather than a crash, and it is filed as issue 144 in the same release that adds streaming.

## Streaming and bfloat16 are both opt-in, with the old behaviour as default

The two performance features in this release deliberately do not change what an existing caller gets. The speech endpoint takes an opt-in stream parameter that returns a streaming response flushing WAV bytes per chunk with 20 millisecond crossfades, and the default behaviour is stated as unchanged. Inference precision works the same way: an environment variable set to on or auto converts the T3 model to bfloat16 and runs under autocast, which the notes put at about 40 percent throughput on capable hardware, with the default off to preserve existing behaviour on upgrade. A third feature has no toggle. The voice conditioning cache means a repeated request against the same reference voice skips re-encoding it, which is the latency win that matters most for batch work through the OpenAI endpoint.

## Portable Mode bundles its own Python 3.10, and only on Windows

The Windows launcher offers a portable mode that is selected by default at first run. It copies the entire project folder into one self-contained directory, including Python and every dependency, so the recipient can unzip it on a machine with no Python at all and start it by double-clicking the batch file. Two details are specified precisely. The embedded runtime is always Python 3.10 and that is the only fully supported version, even though the launcher will run against any system Python 3.10 or newer, which is why portable mode is offered as the way to avoid dependency trouble on a 3.11 system. A flag skips the prompt in either direction. Linux and macOS get standard virtual environments with Python 3.10 installed separately, and portable mode is not offered there.

## Conclusion

This suits someone who wants the Chatterbox family behind an OpenAI-shaped endpoint with a browser UI, and who can accept a moving dependency on a fork rather than a released engine version. Three things to know before you start. The engine is installed with dependency resolution disabled from a branch called master, so a reinstall can pick up different code without any version changing on your side, and one protobuf ceiling is deliberately overridden to make it install at all. The Dockerfile used for the CPU path is still a CUDA base image, so a CPU deployment carries a GPU runtime it will not use. And the shipped compose file sets a Hugging Face token to the literal placeholder string, which fails confusingly rather than obviously if you leave it in place.

## FAQ

### What is devnen/Chatterbox-TTS-Server?

It is a self-hosted server that exposes Resemble AI's Chatterbox model family behind an OpenAI-compatible API with a web UI, and it runs on NVIDIA through CUDA, AMD through ROCm, or CPU. The project is MIT licensed, its web interface and architecture are based on the author's Dia-TTS-Server, and the repository's declared homepage is a Google Colab notebook rather than a site.

### Which Chatterbox models does the server support?

Three, hot-swappable from a dropdown in the web UI: the original 0.5B model, the Multilingual model with 23 languages including Arabic, Chinese, Japanese, Korean, Hindi and Turkish, and Turbo, a 350M-parameter model that distils the mel decoder from ten diffusion steps to one and supports paralinguistic tags such as laugh, cough and chuckle.

### Where does the Chatterbox engine come from?

From a git fork under the project author's account, referenced at a branch named master and installed with pip --no-deps together with s3tokenizer 0.3.0 and onnx 1.16.0, after which protobuf 4.25.0 or newer is forced. The requirements file states that the engine is deliberately not a normal entry because of ONNX source builds and torch version conflicts.

### What did Chatterbox TTS Server v2.0.0 add?

Support for the whole model family, CUDA 13.0 and Strix Halo hardware, an opt-in streaming speech endpoint with 20 millisecond crossfades, a voice conditioning cache, opt-in bfloat16 inference, optional HTTPS through config.yaml, a path traversal fix on the voice file parameters that now returns HTTP 400, endpoints for unloading GPU memory and listing voices, a dynamic language selector, and a chunker fix for issue 144.

### How do I run Chatterbox TTS Server in Docker on an NVIDIA GPU?

The default compose file builds from the main Dockerfile with a build argument that can be nvidia or cpu and defaults to nvidia, publishes port 8004, and requests the GPU through a deploy block reserving one device. If device injection fails, the file's own comment says to comment out the deploy section and uncomment the legacy nvidia runtime line. The base image is a CUDA runtime in both cases, including the CPU path.

## Sources

- [devnen/Chatterbox-TTS-Server on GitHub](https://github.com/devnen/Chatterbox-TTS-Server)
- [License: MIT](https://github.com/devnen/Chatterbox-TTS-Server/blob/main/LICENSE)
- [Project website](https://colab.research.google.com/github/devnen/Chatterbox-TTS-Server/blob/main/Chatterbox_TTS_Colab_Demo.ipynb)
- [README](https://github.com/devnen/Chatterbox-TTS-Server/blob/main/README.md)
- [Releases](https://github.com/devnen/Chatterbox-TTS-Server/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/devnen-chatterbox-tts-server
