Model or dataset
2noise/ChatTTS avatar
2noise/ChatTTS

ChatTTS: deliberately degraded weights, a non-commercial model licence, and a package that reports 0.0.0

GitHub describes it as A generative speech model for daily dialogue.. The repository metadata lists Python as its primary language. The metadata lists the AGPL-3.0 license. This article stays within the project description and details documented in the GitHub repository README.

39,889 stars4,260 forksPythonAGPL-3.0

At a glance

What is it?
ChatTTS is an AGPL licensed text-to-speech model tuned for dialogue, with English and Chinese support and control over prosodic features like laughter and pauses. Two things dominate any evaluation: the public checkpoint is an earlier training stage, and the authors state they degraded the audio on purpose.
Who is it for?
ChatTTS fits research, prototyping and dialogue experiments where you need multiple speakers and control over pauses, laughter and interjections, and where a non-commercial licence is acceptable. It does not fit a product launch, because the model licence is non-commercial and the code licence is AGPL, and it does not fit anyone expecting broadcast audio, because the quality was deliberately reduced.
Can I use it commercially?
Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
Is it still maintained?
Yes. The repository last received commits 174 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The checkpoint on Hugging Face is a base model, not the fine-tuned one

The dataset and model section is short and the distinction in it is the whole story. The main model is trained with Chinese and English audio data of 100,000 or more hours. The open source version on Hugging Face is a 40,000 hour pre-trained model without SFT. So the weights anyone can download are from an earlier stage of training than the model the project describes, and the gap is the supervised fine-tuning stage. The roadmap lists what has actually been opened, which includes the 40k-hours base model and the spk_stats file, streaming audio generation, and the DVAE encoder with zero shot inferring code. Two items remain unchecked: multi-emotion controlling, and ChatTTS.cpp, for which the project is explicitly inviting a new repository in the 2noise organisation. The practical consequence is that the public model is a base model on a research footing, and two capabilities people ask for, emotion control and a C or C++ inference path, do not exist yet in this repository.

The authors added noise and MP3 compression to the released model on purpose

The disclaimer is the most unusual paragraph in this README and it is worth reading without softening it. To limit the use of ChatTTS, the authors say they added a small amount of high-frequency noise during the training of the 40,000-hour model, and compressed the audio quality as much as possible using MP3 format, to prevent malicious actors from potentially using it for criminal purposes. They add that they have internally trained a detection model and plan to open-source it in the future. Three consequences follow, and none of them are speculative. First, the audio quality ceiling is set by design rather than by accident, so artefacts you hear are an intended feature of the release and not something a configuration flag will remove. Second, the countermeasure against misuse is a detection model you cannot obtain yet. Third, anyone planning to train their own checkpoint from the same recipe inherits the same constraint, since it lives in the training data rather than in the inference code.

The code is AGPL and the weights are non-commercial, which are different obligations

There are two licences and they are not interchangeable. The code is published under AGPLv3+. The model is published under CC BY-NC 4.0, stated as intended for educational and research use and not to be used for any commercial or illegal purposes. The model section also says the authors do not guarantee the accuracy, completeness or reliability of the information, that the information and data in the repository are for academic and research purposes only, and that the data was obtained from publicly available sources with no claim of ownership or copyright over it. Put those together and you get a specific commercial position: the code you may use is copyleft, and the weights you may ship are not. A commercial deployment needs either a different checkpoint or a negotiated licence, which is what the [email protected] address is for, alongside four QQ groups and a Discord server. The repository itself is described as containing the algorithm infrastructure and some simple examples, with end-user products collected in a community index repository called Awesome-ChatTTS.

The README says DO NOT INSTALL twice, for two different reasons

The optional accelerator section is unusually candid. TransformerEngine for NVIDIA GPUs on Linux carries a warning reading DO NOT INSTALL, with the explanation that the adaptation is under development and cannot run properly now, that it should only be installed for developing purposes, and that the installation process is very slow. FlashAttention-2 carries the same warning for a different reason, namely that it will slow down generating speed according to a linked issue in the transformers repository. The one optional accelerator presented without a warning is vLLM, and it is Linux only and pinned to an exact version:

bash
pip install safetensors vllm==0.2.7 torchaudio

So every documented route to faster inference on this project is either unreleased, slower, or a single pinned Linux package. That is not a criticism so much as a map. If you are evaluating on a Mac or on CPU, none of these apply, and the plain requirements path is what you get.

setup.py reports version 0.0.0 unless an environment variable overrides it

The packaging metadata has one line that matters more than it looks:

python
version = "v0.0.0"

setup(
    name="chattts",
    version=os.environ.get("CHTTS_VER", version).lstrip("v"),

The declared default is a literal placeholder. The real version is only present if something exports CHTTS_VER in the environment before the build, and the leading v is stripped on the way in. So a source install reports 0.0.0, and package metadata cannot tell you what you are running. The README offers three installation routes with genuinely different version semantics: the stable release from PyPI with pip install ChatTTS, the latest from GitHub with pip install git+https://github.com/2noise/ChatTTS, and a local editable install. The published tags are v0.2.5 on 2026-04-10, v0.2.4 on 2025-05-23 and v0.2.3 on 2025-02-18, an irregular cadence with a gap of about eleven months before the most recent. The distribution name is lowercased to chattts while the project styles itself ChatTTS. Record the tag you used somewhere other than the package metadata.

requirements.txt is a superset of install_requires, and three pins only apply on Linux

Two dependency lists exist and they are not the same, which produces a very specific failure. The install_requires in setup.py is numba, numpy below 3.0.0, pybase16384, torch at or above 2.1.0, torchaudio, tqdm, transformers at or above 4.41.1, vector_quantize_pytorch and vocos. The requirements.txt used for a source clone adds IPython, gradio, av, pydub and requests, and then adds three text processing packages each gated on the platform: pynini pinned to exactly 2.1.5, WeTextProcessing and nemo_text_processing, all with sys_platform equal to linux. The consequences are concrete. Install the stable package from PyPI and gradio is absent, so the documented WebUI entry point, python examples/web/webui.py, will fail to import. And the text normalisation stack that the Chinese path depends on is not installed on macOS or Windows, so the same input text can be normalised differently depending on the operating system. If you are not on Linux, install from the clone and read requirements.txt rather than trusting the package metadata.

A bare except hides save failures, and a speaker embedding only exists if you save it

The basic usage example ends with a defensive block that swallows everything:

python
for i in range(len(wavs)):
    """
    In some versions of torchaudio, the first line works but in other versions, so does the second line.
    """
    try:
        torchaudio.save(f"basic_output{i}.wav", torch.from_numpy(wavs[i]).unsqueeze(0), 24000)
    except:
        torchaudio.save(f"basic_output{i}.wav", torch.from_numpy(wavs[i]), 24000)

The intent is clear, a version difference in torchaudio's expected shape, and the fallback drops the unsqueeze. But the except clause is bare, so it also catches a full disk, a bad path and a permissions problem, then tries the second call, which raises again and produces a traceback that points at audio handling rather than at the real cause. The sample rate is hard-coded to 24000 in both branches. Speaker identity is handled the same manual way: sample_random_speaker returns an embedding you are told to print and save for later timbre recovery, and nothing in the library persists it. Lose the embedding and you lose that voice. The examples directory offers six surfaces, api, cmd, ipynb, onnx and web, with an openai_api.ipynb notebook at the repository root, so the quickest route to serving this is a notebook rather than a documented service.

Editorial conclusion

ChatTTS fits research, prototyping and dialogue experiments where you need multiple speakers and control over pauses, laughter and interjections, and where a non-commercial licence is acceptable. It does not fit a product launch, because the model licence is non-commercial and the code licence is AGPL, and it does not fit anyone expecting broadcast audio, because the quality was deliberately reduced. Before you build on it, confirm which checkpoint you have, since the published one is a base model without supervised fine-tuning, read the disclaimer about the noise and MP3 compression that were added on purpose, expect the WebUI example to fail after a plain PyPI install because gradio is not a declared dependency, and check the language support beyond English and Chinese.

Frequently asked questions

how to install chattts

Three routes are given. pip install ChatTTS gets the stable version from PyPI, pip install git+https://github.com/2noise/ChatTTS gets the latest from GitHub, and pip install -e . installs a local checkout in development mode. From a clone you can instead run pip install --upgrade -r requirements.txt, or create a conda environment with python=3.11 first, which pulls in gradio and the Linux-only text processing packages that the package metadata omits.

what is chat tts

A generative speech model for daily dialogue, and a text-to-speech model designed specifically for dialogue scenarios such as an LLM assistant. It supports English and Chinese, handles multiple speakers for interactive conversations, and can predict and control fine-grained prosodic features including laughter, pauses and interjections. The repository itself is described as containing the algorithm infrastructure and some simple examples.

Is TTS free?

For ChatTTS specifically, the code is published under AGPLv3+ and the model under CC BY-NC 4.0, which the project states is intended for educational and research use and should not be used for commercial purposes. Formal enquiries about the model and roadmap go to [email protected]. The published checkpoint is a 40,000 hour pre-trained model without supervised fine-tuning, released for academic purposes only.

chattts vs kokoro

The README contains no comparison with Kokoro or with any other text-to-speech system, so there is nothing on this page to weigh against. What it does claim for ChatTTS is that it surpasses most open-source TTS models in terms of prosody, and what it discloses is that the released 40,000 hour model was deliberately trained with added high-frequency noise and MP3 compression to limit its misuse.

How do I use TTS with ChatTTS?

Instantiate ChatTTS.Chat, call load, then pass a list of strings to infer, which returns waveforms you save with torchaudio. Two ready-made entry points are documented as well: python examples/web/webui.py launches a web interface, and python examples/cmd/run.py takes text arguments and writes mp3 files to ./output_audio_n.mp3. Both commands must be run from the project root.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/2noise-chattts.svg)](https://hysenlabs.com/projects/2noise-chattts)