# rsxdalv/TTS-WebUI: One Gradio and React Front End for Two Dozen Speech and Audio Models

> TTS-WebUI bundles Bark, XTTSv2, RVC, MusicGen and about twenty more models behind a single interface, with Docker images, a pip package and an OpenAI-compatible API. Here is how it is put together, where it breaks, and who should pick it over a single-model tool.

**rsxdalv/TTS-WebUI** — A single Gradio + React WebUI with extensions for ACE-Step, OmniVoice, Kimi Audio, Piper TTS, GPT-SoVITS, CosyVoice, XTTSv2, DIA, Kokoro, OpenVoice, ParlerTTS, Stable Audio, MMS, StyleTTS2, MAGNet, AudioGen, MusicGen, Tortoise, RVC, Vocos, Demucs, SeamlessM4T, and Bark!

- Repository: https://github.com/rsxdalv/TTS-WebUI
- Website: http://TTSWebUI.com/
- Stars: 3,278 · Forks: 333
- Language: TypeScript
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/rsxdalv-tts-webui

## What TTS-WebUI is for, and who ends up using it

TTS-WebUI is a single web interface that fronts a long list of speech and audio models. The README's supported-model table splits them into text-to-speech (Bark, Tortoise, XTTSv2, StyleTTS2, Parler TTS, OpenVoice, F5-TTS, MMS, SeamlessM4T and others), audio and music generation (MusicGen, MAGNeT, Stable Audio, ACE-Step), and conversion or cleanup tools (RVC, Demucs, Vocos, Whisper, Resemble Enhance, Audio Separator).

The project targets people who want to compare models without building a separate environment for each one. Voice cloning, speech synthesis and stem separation all live behind the same tabs, and the same server exposes an OpenAI-compatible API. That combination is the reason the project exists: instead of five repositories with five installers, you get one process and one set of output folders.

The cost is weight. A single-model tool such as Piper ships a small runtime and does one thing. TTS-WebUI pulls in torch, torchaudio, xformers on CUDA, gradio, fastapi and a set of separately versioned extension packages. If you only ever need one voice, that trade is bad. If you are evaluating several, it is the point.

## How the Gradio app, React UI and extension packages fit together

The repository separates the Python back end from the browser front end. tts_webui/ holds the Python package, react-ui/ holds the TypeScript front end (the primary language of the repository is TypeScript), and server.py is the entry point. The pyproject.toml declares a console script, tts-webui, mapped to tts_webui.cli:main, so the package is meant to be installed and launched rather than only cloned.

Models are not all vendored into the core. requirements.txt lists core extensions as separate distributions: tts_webui_extension.demucs, tts_webui_extension.mms, tts_webui_extension.seamless_m4t, tts_webui_extension.vocos, tts_webui_extension.openai_tts_api, tts_webui_extension.log_viewer and tts-webui-extension.vall_e_x. The Dockerfile installs two more from a dedicated index, tts-webui-extension.rvc and tts-webui-extension.styletts2, using --extra-index-url https://tts-webui.github.io/extensions-index/. That index is the distribution channel for model code that is too heavy or too licence-specific to sit in the core requirements.

The API side is a FastAPI and uvicorn service, described in requirements.txt as the same stack as the OpenAI TTS API. Storage is SQLite, and the docker-compose file notes that no separate database container is needed because the database files live under ./data/sqlite/. Ports are fixed by the compose file: 7770 for the Gradio interface, 3000 for the React UI, and 7778 for the OpenAI API.

## Installing TTS-WebUI with Docker

The compose file is the shortest path. It pins the image ghcr.io/rsxdalv/tts-webui:main, maps three ports, and mounts data, outputs and favorites as bind mounts so generated audio survives a container restart. The voices directory is present but commented out, so voice files are not persisted until you uncomment that line.

```yaml
services:
  tts-webui:
    image: "ghcr.io/rsxdalv/tts-webui:main"
    restart: unless-stopped
    ports:
      - ${TTS_PORT:-7770}:7770
      - ${UI_PORT:-3000}:3000
      - ${OPENAI_API_PORT:-7778}:7778
    volumes:
      - ./data:/app/tts-webui/data
      - ./outputs:/app/tts-webui/outputs
      - ./favorites:/app/tts-webui/favorites
```

Run docker compose up and the Gradio interface should answer on port 7770, the React UI on 3000, and the API on 7778. The same file reserves an NVIDIA device for the container. The commented NVIDIA_DISABLE_REQUIRE variable is worth knowing about: the compose comments say it disables the hard CUDA version check and may let some models run even when the CUDA version does not match, while warning that errors can still occur. That is a diagnostic switch, not a fix.

The image itself is built from nvidia/cuda:12.8.0-devel-ubuntu22.04 with Python 3.10, Node 22.9.0 and torch 2.11.0 with xformers 0.0.35 from the cu128 wheel index. The base image choice is the main constraint: this is a CUDA-first deployment, and the devel image is larger than a runtime-only image would be.

## Installing from pip and running a first synthesis

The package installs as TTS-WebUI and exposes the tts-webui command. The optional dependency groups in pyproject.toml select the torch build: cpu, cuda, mac, rocm and intel all pin torch 2.11.0, and only cuda adds xformers 0.0.35. Python 3.8 or later is declared, though the Docker image standardises on 3.10.

```bash
pip install TTS-WebUI[cuda]
tts-webui
```

After the server starts, the Gradio interface is the first real use. Pick a text-to-speech tab, paste a sentence, and generate. Outputs land in the outputs directory, which is why the compose file mounts it. If you want the API instead of the browser, the openai_tts_api extension is installed by default from requirements.txt and listens on port 7778, so an OpenAI-style client pointed at that port talks to the same models.

One practical note before you start: the RVC and StyleTTS2 extensions are not in the default pip path. The Dockerfile installs them with an explicit extra index, so a plain pip install will not give you those two voices.

```bash
pip install "tts-webui-extension.rvc>=0.0.6" --extra-index-url https://tts-webui.github.io/extensions-index/
pip install "tts-webui-extension.styletts2>=0.1.0" --extra-index-url https://tts-webui.github.io/extensions-index/
```

The README also links a Windows installer archive and a Google Colab notebook, which is the option to consider if you do not have a local GPU.

## Where TTS-WebUI gets in your way

The dependency surface is the first real limitation. requirements.txt pins gradio to 5.49.1, pydantic to 2.10.6 and pillow to 10.3.0, and the comments explain that the latter two exist to work around gradio and conda issues. Those pins are deliberate, but they mean installing TTS-WebUI into an environment that already has a different gradio or pydantic version is a conflict you have to resolve yourself. The same applies to torch 2.11.0, which is pinned in every optional dependency group.

Model coverage is uneven in a way the table does not show. Several entries carry an asterisk, and the README does not explain what the asterisk means. Some of those entries point at third-party forks rather than upstream projects, for example the AudioCraft Mac and AudioCraft Plus links. If you need a specific model, check whether it is a core extension, an index extension, or a fork before you plan around it.

The project is under active development, which cuts both ways. The last push was on 2026-09-07, and releases v1.5.0, v1.5.1 and v1.5.2 landed between 2026-04-30 and 2026-08-31. That pace means the main image tag moves. The README does not document rollback, so if you track ghcr.io/rsxdalv/tts-webui:main you should pin a digest or build your own image from a tagged commit.

Finally, this is a GPU-first project. The compose file reserves an NVIDIA device, the Dockerfile starts from a CUDA devel image, and the CPU extra exists but the README does not claim CPU parity for the heavier models. If your only machine is a laptop without CUDA, the Colab notebook is the more honest starting point.

## TTS-WebUI compared with Piper and other single-model tools

Piper is the clearest alternative for a narrow use case. Piper is a text-to-speech engine with a small runtime and no model menu; you install it, point it at a voice, and it synthesises. TTS-WebUI is the opposite shape: a host process that loads whichever model a tab asks for, plus conversion and separation tools that Piper does not attempt.

The difference shows up in operations. A Piper deployment is a binary and a voice file. A TTS-WebUI deployment is a container with CUDA, torch, gradio, FastAPI, SQLite and a set of extension packages, and its upgrade path is a moving image tag. If your requirement is a single stable voice on modest hardware, Piper wins on every axis that matters. If your requirement is to try XTTSv2 against StyleTTS2 against F5-TTS on the same text, TTS-WebUI is the only one of the two that can do it without a new environment each time.

The same logic applies against running a model's own repository. Upstream projects such as Tortoise or OpenVoice give you the model and its own inference script. TTS-WebUI gives you a shared interface, a shared output folder, and an OpenAI-compatible endpoint in front of all of them. What you give up is directness: when something breaks, you are debugging an integration layer as well as the model.

## Maintenance cost, licence and upgrade planning

TTS-WebUI is MIT licensed, and the LICENSE file sits at the repository root. The pyproject.toml classifiers list the MIT License as well. That covers the code in this repository, not the models it loads. The extension packages are separate distributions with their own licences, and the model weights they download carry their own terms. The README's model table links each model to its upstream project, which is where you check the terms that apply to weights and to generated audio. This is a description of what the repository states, not legal advice.

Upgrade cost concentrates in three places. Torch is pinned at 2.11.0 across every extra, so moving to a newer torch means waiting for the project to move first. Gradio is pinned at 5.49.1, and the pydantic and pillow pins exist to keep that combination working. The extension packages are version-ranged, for example tts_webui_extension.demucs>=0.0.4 and tts-webui-extension.rvc>=0.0.6, so they can advance independently of the core and break it.

Pinning the container image is the cheapest insurance. The compose file uses the main tag, which is not a release tag. Releases are cut as v1.5.0 through v1.5.2, so a deployment that matters should reference a digest rather than main. The README does not describe a downgrade procedure, and the compose file has no rollback section, so the safe assumption is that rolling back means redeploying a previously pinned image and keeping the ./data bind mount intact.

## Conclusion

Adopt TTS-WebUI if you want many speech and audio models behind one interface and one OpenAI-compatible endpoint, and you are willing to run a CUDA 12.8 container or a Python 3.10 environment to get it. Do not adopt it if you need a small dependency footprint, a documented rollback path, or a single model you already run well on its own. Before committing, verify three things in your own environment: that the ghcr.io/rsxdalv/tts-webui:main image starts with your NVIDIA driver, that the extensions you need resolve from the extensions index, and that port 7778 answers where your client expects the API. The README does not document rollback, so pin the image tag yourself rather than tracking main.

## FAQ

### What is TTS-WebUI used for?

It is a single Gradio and React web interface for a long list of speech and audio models, covering text-to-speech, music and audio generation, and conversion tools such as RVC and Demucs. It also exposes an OpenAI-compatible API on port 7778.

### Is TTS-WebUI AI generated?

The interface and the code are a normal software project under the MIT licence. The audio it produces is generated by the machine learning models it loads, such as Bark, XTTSv2 or MusicGen, each of which is a separate upstream project.

### How do I use TTS-WebUI?

Deploy it with the provided docker-compose.yml, which maps ports 7770 for Gradio, 3000 for the React UI and 7778 for the API, or install the package and run the tts-webui command. Then pick a model tab in the interface, enter text and generate.

### Is TTS WebUI safe to run?

The repository is MIT licensed and not archived, and the container runs from the published ghcr.io/rsxdalv/tts-webui image with data, outputs and favorites mounted from the host. The README does not make a security claim, so the relevant checks are your own: which extensions you install and where the model weights come from.

## Sources

- [License: MIT](https://github.com/rsxdalv/TTS-WebUI/blob/main/LICENSE)
- [Project website](http://TTSWebUI.com/)
- [README](https://github.com/rsxdalv/TTS-WebUI/blob/main/README.md)
- [Releases](https://github.com/rsxdalv/TTS-WebUI/releases)
- [rsxdalv/TTS-WebUI on GitHub](https://github.com/rsxdalv/TTS-WebUI)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/rsxdalv-tts-webui
