CLI tool
jianchang512/ChatTTS-ui avatar
jianchang512/ChatTTS-ui

ChatTTS-ui: a local web interface and HTTP API for ChatTTS

一个简单的本地网页界面,使用ChatTTS将文字合成为语音,同时支持对外提供API接口。A simple native web interface that uses ChatTTS to synthesize text into speech, along with support for external API interfaces.

7,660 stars917 forksPythonNOASSERTION

At a glance

What is it?
ChatTTS-ui wraps the ChatTTS speech model in a Flask page and a POST /tts endpoint, with prebuilt Windows packages and Docker files. It is convenient, but model download, GPU detection and voice stability are the parts that decide whether it fits your setup.
Who is it for?
Adopt ChatTTS-ui if you want a local ChatTTS endpoint on your own machine and can accept a Python 3.9 to 3.11 environment, a first-run model download and a voice seed that is not perfectly reproducible. Do not adopt it if you need a permissively licensed, stable production voice service or a supported macOS GPU path; the repository carries no recognised licence file and the README documents CPU-only behaviour on Mac.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 107 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap ChatTTS-ui fills between the ChatTTS model and a usable endpoint

The upstream ChatTTS project is a model. Running it means writing Python that loads weights, prepares text and writes a wav file. ChatTTS-ui is the layer above that: a Flask application that serves a page for typing text and a POST endpoint for programs. The README describes it as "a simple native web interface that uses ChatTTS to synthesize text into speech, along with support for external API interfaces." The target user is someone who wants speech output on their own hardware without building a service around the model first.

The scope is deliberately narrow. There is no user management, no queue, no storage layer beyond wav files written under static/wavs. The repository layout confirms this: app.py, templates/, static/, a speaker/ folder for voice data, and Dockerfiles for CPU and GPU. If you need multi-tenant scheduling or a hosted SLA, this is the wrong shape. If you need a local /tts endpoint that returns a wav URL, it is close to the shortest path available.

How the request flows from text to a wav file

The API contract in the README is the clearest view of the architecture. A client posts form data to http://127.0.0.1:9966/tts with text required and the rest optional: voice, prompt, temperature (default 0.3), top_p (default 0.7), top_k (default 20), skip_refine (default 0) and custom_voice (default 0). The server returns JSON with code 0 and an audio_files array, where each element carries a filename (the absolute path to the wav) and a url under /static/wavs/. Failures return code 1 with a message string.

Two details matter more than they look. First, prompt accepts control tokens such as [oral_2][laugh_0][break_6], which is how the project exposes prosody, laughter and pauses without a separate settings UI. Second, custom_voice is a seed value and, when set above zero, it overrides voice entirely. The README states plainly that the same seed on different devices produces different audio, and even on one device the timbre can shift, especially pitch. Treat voice selection as approximate, not as a reproducible asset identifier.

The service itself is served by waitress rather than the Flask development server, which is why the default binding is a host and port pair in .env rather than a debug flag.

Installing ChatTTS-ui and making a first synthesis request

The fastest path on Windows is the prebuilt archive. The README says to download the compressed package from Releases, extract it and double-click app.exe. Two caveats come with it: some security software may flag the executable, and GPU acceleration only turns on with an Nvidia card above 4G of VRAM and CUDA 12.8 or newer installed.

Source deployment follows the same shape on Linux, macOS and Windows. On Linux the README uses Python 3.9 to 3.11, an empty directory, a virtual environment and requirements.txt, then torch installed separately. The commands below are the Linux sequence as written in the README, including the CPU-only torch install.

bash
cd /data/chattts && git clone https://github.com/jianchang512/chatTTS-ui .
python3 -m venv venv
source ./venv/bin/activate
pip3 install -r requirements.txt
pip3 install torch==2.7.1 torchaudio==2.7.1
python3 app.py

For CUDA acceleration the README substitutes a PyTorch index URL and adds two Nvidia packages, then requires CUDA 12.8 or newer ToolKit installed separately.

bash
pip install torch==2.7.1 torchaudio==2.7.1 --index-url https://download.pytorch.org/whl/cu128
pip install nvidia-cublas-cu11 nvidia-cudnn-cu11

On first start the application downloads the model. It checks whether https://huggingface.co is reachable and falls back to modelscope.cn. The README warns that the modelscope path must not go through a proxy, and that source deployments need working access to one of the two hosts or the download fails. When the server is up it opens a browser at http://127.0.0.1:9966 by default.

Once running, a first real request is a POST to the tts endpoint. The README's own example sends form fields and prints the parsed JSON.

python
import requests

res = requests.post('http://127.0.0.1:9966/tts', data={
  "text": "若不懂无需填写",
  "prompt": "",
  "voice": "3333",
  "temperature": 0.3,
  "top_p": 0.7,
  "top_k": 20,
  "skip_refine": 0,
  "custom_voice": 0
})
print(res.json())

A successful response looks like code 0 with an audio_files list containing a filename and a url under /static/wavs/. If you get code 1, the message field carries the reason. The container route is different but shorter: clone the repository, then bring up docker-compose.gpu.yaml or docker-compose.cpu.yaml and follow the logs.

bash
git clone https://github.com/jianchang512/ChatTTS-ui.git chat-tts-ui
cd chat-tts-ui
docker compose -f docker-compose.cpu.yaml up -d
docker compose logs -f --no-log-prefix

The README notes the container binds to 0.0.0.0:9966, so the page is reachable at the host IP and port rather than only on loopback.

Configuration in .env and the GPU detection trap

Almost all operational tuning lives in one file. The README shows three keys: WEB_ADDRESS, which defaults to 127.0.0.1:9966 and can be changed to something like 192.168.0.10:9966 for LAN access; compile, set to false; and device, set to default. The device key is the interesting one. Its default means the application prefers cuda when VRAM exceeds 4G, otherwise mps, and it can be forced to cpu, mps or cuda.

That automatic choice is where deployments go wrong. The README documents a specific failure: on Windows or Linux, with an Nvidia card above 4G of VRAM, a source deployment may still run on CPU. The suggested fix is to uninstall torch and torchaudio and reinstall the CUDA build from the PyTorch index, with CUDA 12.8 or newer already present. Note the version mismatch risk here: the README's CUDA install command pins torch 2.7.1, while pyproject.toml lists torch ^2.3.0 against the cu118 index. The README is the newer instruction, but the two files disagree, and that is worth checking before you script an install.

A second constraint is that GPU memory below 4G forces CPU regardless of what you set. On a shared machine this can turn a fast synthesis into a slow one without any error message.

Voices are approximate, and the project says so

Voice handling is the least deterministic part of ChatTTS-ui, and the README is unusually direct about it. Since version 0.92 the project accepts fixed voices in csv or pt format, placed in the speaker folder. pt files come from the ChatTTS_Speaker project's demo page on modelscope, and numeric voice values can be read from a separate listing page and pasted into the custom voice field.

The caveat follows immediately: the same seed value produces different audio on different devices, and on the same device the timbre can still vary, particularly pitch. That is a property of the underlying model, not a bug in the wrapper, but it changes how you should use the project. Do not treat a voice value as a brand asset that will sound identical across machines or across restarts. If your application needs a consistent narrator identity, you will be re-auditioning output, not just setting a constant.

The API reflects the same looseness. voice defaults to 2222 and the README lists 2222, 7869, 6653, 4099 and 5099 as options, but says any other value is accepted and will be used to pick a random timbre. custom_voice, when set above zero, takes precedence over voice.

Where ChatTTS-ui is the wrong tool

The licence is the first hard limit. The repository's LICENSE file is not a recognised open source licence according to the metadata, and the project is marked NOASSERTION. That is not a statement about legality, and it is not legal advice, but it means you cannot assume the permissions a standard licence would give you. If your organisation requires an approved licence identifier before adoption, this project does not provide one, and the upstream ChatTTS model carries its own terms as well.

Platform support is the second. The README documents Windows, Linux and macOS source deployment, but the CUDA path is Nvidia-only and the macOS instructions install plain torch with no acceleration branch. GPU acceleration in the prebuilt package is described only for Nvidia cards above 4G VRAM with CUDA 12.8 or newer. If you are on AMD or Apple silicon and expect GPU speed, the documentation does not promise it.

The third is operational. There is no documented authentication on the /tts endpoint, no rate limiting and no queue. Changing WEB_ADDRESS to a LAN address, as the README suggests for local network access, exposes synthesis to anyone who can reach that port. The project is a local tool; treating it as a public service requires work the README does not describe.

ChatTTS-ui against calling the ChatTTS model directly

The obvious alternative is the upstream 2noise/chattts repository, which the README links as the original project. The difference is not quality but surface area. Upstream gives you the model and its Python API; you decide how text is segmented, how the wav is stored, how requests arrive and how the process is supervised. ChatTTS-ui gives you those decisions already made: Flask plus waitress on a configurable host and port, a form page, a POST /tts contract with fixed parameter names, wav files under static/wavs, and Dockerfiles for CPU and GPU.

That trade runs both ways. Going direct means you can batch, cache and shape the pipeline around your own storage, and you avoid inheriting the wrapper's version pins. Using ChatTTS-ui means you get a working endpoint in one command and a browser page for auditioning voices, at the cost of accepting its parameter names, its file layout and its release cadence. The project also integrates with pyVideoTrans 1.82 and newer, where you point the ChatTTS setting at the request address and select ChatTTS in the main interface, which is a concrete reason to choose the wrapper if you already use that tool.

Maintenance, releases and upgrade cost

The last push to the default branch was on 2026-06-14, and the same date carries the v260614 Windows release. The two prior releases are much older, v1.0 from 2024-08-05 and v0.95 from 2024-06-03, so the release history is sparse rather than steady. The repository is not archived. Recent activity is real, but the gap between v1.0 and v260614 means you should read the release notes for a given version rather than assume continuous small updates.

Upgrading has two distinct costs. For the Docker route the README gives an explicit procedure: check out main, pull, bring the compose stack down, then rebuild with the gpu or cpu compose file and follow the logs. For source deployments there is no documented upgrade path at all; the README covers first install only, and does not describe how to migrate an existing venv or how to handle the speaker folder across versions. The dependency pins in requirements.txt and pyproject.toml are another moving part, and as noted they do not currently agree on the torch version.

On licensing, the practical implication is procedural rather than legal: because the repository does not carry an OSI-recognised licence identifier, any adoption review that keys off licence metadata will stop here. The upstream ChatTTS model adds its own terms on top, and the README does not summarise them.

Editorial conclusion

Adopt ChatTTS-ui if you want a local ChatTTS endpoint on your own machine and can accept a Python 3.9 to 3.11 environment, a first-run model download and a voice seed that is not perfectly reproducible. Do not adopt it if you need a permissively licensed, stable production voice service or a supported macOS GPU path; the repository carries no recognised licence file and the README documents CPU-only behaviour on Mac. Before committing, verify that the machine can reach huggingface.co or modelscope.cn, check the .env device and WEB_ADDRESS values, and confirm your GPU has more than 4G of VRAM if you expect CUDA acceleration.

Frequently asked questions

What is ChatTTS-ui and who is it for?

It is a local web interface and HTTP API that uses ChatTTS to turn text into speech, supporting mixed Chinese, English, numbers and symbols. The README positions it for people who want synthesis on their own machine plus a POST endpoint for other programs.

How do I install ChatTTS-ui on Windows?

The README offers a prebuilt package: download the archive from Releases, extract it and double-click app.exe. Source deployment requires Python 3.9 to 3.11, Git, a virtual environment and a separate torch install, with CUDA 12.8 or newer for GPU acceleration.

What are the request parameters for the ChatTTS-ui API?

POST to http://127.0.0.1:9966/tts with text required and optional voice, prompt, temperature, top_p, top_k, skip_refine and custom_voice. The response is JSON with code 0 and an audio_files array, or code 1 with an error message.

Why does ChatTTS-ui use the CPU even though I have an Nvidia GPU?

The README says GPU memory below 4G forces CPU, and that a source deployment with a card above 4G may still run on CPU. The suggested fix is to uninstall torch and torchaudio and reinstall the CUDA build, with CUDA 12.8 or newer installed.

Does the same voice value produce the same audio every time in ChatTTS-ui?

No. The README states that the same seed on different devices produces different audio, and that even on one device the timbre can change, especially pitch. Treat voice values as approximate rather than reproducible.

Official sources

  1. Issues
  2. jianchang512/ChatTTS-ui on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/jianchang512-chattts-ui.svg)](https://hysenlabs.com/projects/jianchang512-chattts-ui)