# Coqui TTS: setup.py blocks Python 3.12, and the last commit is from August 2024

> Coqui TTS is a large text-to-speech research toolkit with pretrained models in more than 1100 languages, a Trainer API, dataset curation tools, and implementations of dozens of published architectures. The details that decide whether it works for you are unglamorous: a hard Python version gate in setup.py, a requirements file where every language and model family is a mandatory dependency, a CUDA-based default Docker image, and a repository whose last push was 2024-08-16.

**coqui-ai/TTS** — GitHub describes it as 🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production. The repository metadata lists Python as its primary language. The metadata lists the MPL-2.0 license. This article stays within the project description and details documented in the GitHub repository README.

- Repository: https://github.com/coqui-ai/TTS
- Website: http://coqui.ai
- Stars: 46,080 · Forks: 6,167
- Language: Python
- License: MPL-2.0
- Published: 2026-08-13 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/coqui-ai-tts

## setup.py raises a RuntimeError outside Python 3.9 to 3.11, before anything installs

The version constraint is not a warning or an install hint, it is a gate that runs when the build script is imported. setup.py reads sys.version, compares it with packaging.version.Version, and if the interpreter is below 3.9 or at or above 3.12 it raises RuntimeError with the message TTS requires python >= 3.9 and < 3.12 but your Python version is, followed by the full version string it read. Consequence: on Python 3.12, which is a perfectly ordinary interpreter today, the install does not degrade or warn, it stops, and pip reports a build failure that looks like a broken package rather than a policy decision. The README states the same window and says the toolkit is tested on Ubuntu 18.04, which is a much older base than most people run now. There is no environment marker in requirements.txt that would have warned you earlier, and no alternate branch documented. If your project is already on 3.12, the realistic options are a separate interpreter for speech work, or a fork, because nothing in the repository offers a supported escape hatch.

## Every language and model family is a hard requirement rather than an extra

requirements.txt is one flat list with comments marking groups, and the groups are informational rather than optional. Under chinese g2p deps you get jieba and pypinyin. Under korean you get hangul_romanize, plus jamo, nltk and g2pkk further down. The gruut line carries its own language extras, gruut[de,es,fr]==2.2.3, pinned to an exact version. Under deps for bangla there are three packages, bangla, bnnumerizer and bnunicodenormalizer. Then the model families: deps for XTTS pulls unidecode, num2words and spacy[ja]>=3, deps for bark pulls encodec, and deps for tortoise pulls einops and transformers>=4.33.0. Consequence: a user who only wants to synthesize English with one released checkpoint installs a Japanese spaCy model, a Bangla normalizer, a Korean romanizer, and the entire Transformer stack. Disk and install time are the visible cost; the subtler cost is that any one of these packages failing to resolve blocks the whole installation, and you cannot narrow the list without editing the file.

## A synthesis-only user still gets torch, training plots, and notebook tooling

The dependency list is organised for the library's full scope, not for inference. The core group alone requires torch>=2.1, torchaudio, scipy>=1.11.2, librosa>=0.10.0, scikit-learn>=1.3.0, cython>=0.29.30, soundfile>=0.12.0, inflect, tqdm, anyascii, and pyyaml. Then there are groups labelled deps for examples with flask>=2.0.1, deps for notebooks with umap-learn>=0.5.1 and pandas>=1.4,<2.0, and deps for training with matplotlib>=3.7.0, all installed by default. Two pins are exact rather than ranged: numpy==1.22.0 for Python 3.10 and below, and mutagen==1.47.0. One marker is dead code, because numba==0.55.1 applies to python_version below 3.9 and setup.py already refuses anything under 3.9. And fsspec carries a floor of 2023.6.0 for a stated reason in the file, that anything at or below 2023.9.1 makes aux tests fail, so a lower bound exists to satisfy the test suite rather than the library. The two Coqui stack dependencies, trainer>=0.0.36 and coqpit>=0.0.16, are both pre-1.0.

## The default branch is dev, and the README's own links name dev, main, and master

Three branch names appear in a single page. The installation link points at tree/dev, the contributing link points at blob/main, and the code of conduct link points at blob/master, while the repository's own default branch is dev. Consequence: a reader clicking through the documentation from a fresh clone lands in different trees depending on which link they followed, and one of those branches may be a stale mirror rather than the working branch. The same split shows up in how models are distributed. Released models live in the GitHub releases, and anything experimental lives in a wiki page called Experimental-Released-Models, which is editable by hand and therefore not versioned with the code. The fine-tuning story is a directory rather than a feature list, with example recipes under recipes/ljspeech, and the project's plan is not a document at all but a GitHub issue numbered 378 titled as the main development plans. Treat all of these as moving targets: a recipe that worked against dev has no guarantee against a release, and the roadmap is a discussion thread.

## The performance section showcases models that are not released, and two language counts disagree

There is a section on TTS performance, and it opens by saying that the underlined entries TTS* and Judy* are internal models that are not released open-source, present to show the potential. Models prefixed with a dot, and the three named are .Jofish, .Abe and .Janice, are real human voices. So the first thing a reader learns about output quality is that the headline examples are partly unobtainable. Consequence: a voice you hear in a demo or a shared notebook may correspond to a checkpoint you cannot fetch, and the only published evidence of quality is a page whose best entries are not downloadable. The news list above it has a related inconsistency worth noting: one item announces XTTSv2 with 16 languages, and another describes XTTS as a production model that can speak 13 languages. Those are two different things, but nothing in the page labels which model a given figure belongs to, so check the per-model documentation on the readthedocs site before you commit to a language count.

## The Dockerfile defaults to a CUDA base and replaces the system llvmlite

The container build is not a neutral default. It starts from an ARG named BASE that defaults to nvidia/cuda:11.8.0-base-ubuntu22.04, so an unedited build pulls a GPU runtime image even on a machine that will never use one. The apt layer installs gcc, g++, make, python3, python3-dev, python3-pip, python3-venv, python3-wheel, espeak-ng and libsndfile1-dev, and that list is the most useful thing in the file for a bare host, because espeak-ng and libsndfile1 are the two system audio dependencies the toolkit expects. The next layer runs pip3 install llvmlite --ignore-installed, which deliberately discards the system copy of llvmlite, and the one after installs the CUDA build of the framework wheels:

```bash
pip3 install llvmlite --ignore-installed
pip3 install torch torchaudio --extra-index-url https://download.pytorch.org/whl/cu118
```

The build then copies the repository and runs make install, and the entrypoint is tts with --help as the default command. Consequence for a CPU deployment: you are expected to override BASE deliberately, and a related search for a CPU build is answered by that override, not by a flag.

## Breadth is counted in papers, and the per-model contract lives on the docs site

The model list is the project's strongest feature and its main source of confusion. Spectrogram models number thirteen, including Tacotron, Tacotron2, Glow-TTS, Speedy-Speech, Align-TTS, FastPitch, FastSpeech, FastSpeech2, SC-GlowTTS, Capacitron, OverFlow, Neural HMM TTS and Delightful TTS. End-to-end models add XTTS, VITS, YourTTS, Tortoise and Bark. Then six attention methods, two speaker encoder approaches, eight vocoders, and FreeVC for voice conversion. The 1100 plus language claim comes from the availability of roughly 1100 Fairseq models, so that number is a model collection size, not a quality statement about each language. Consequence: breadth here means many architectures are implemented and readable, and every one of them is a separate configuration, checkpoint and documentation page under tts.readthedocs.io, for example the xtts, bark and tortoise pages linked from the news list. A second-order cost is the build side, where pyproject.toml requires cython~=0.29.30 and numpy>=1.22.0 at build time, so compiled extensions are part of the install and a missing toolchain is a plausible failure.

## Conclusion

Coqui TTS fits a reader on Python 3.9 through 3.11 with a CUDA machine, who wants a research framework in which they can read the training code for a specific architecture, add their own, or fine-tune XTTS, and who is prepared to own the dependency surface themselves. It does not fit a reader on Python 3.12, a reader who wants a small inference-only dependency tree, or a reader who needs a project that will still be receiving fixes, because the newest release is v0.22.0 from 2023-12-12 and the repository was last pushed on 2024-08-16. Before you install, check five things: which Python you have, since setup.py raises a RuntimeError outside 3.9 to 3.11; whether you have espeak-ng and libsndfile1 present, which is what the Docker image installs; whether the model you want is released or sits in the wiki's experimental list; whether the quality you heard in a demo came from a model marked as internal and not released; and whether your GPU situation matches the default nvidia/cuda base, since overriding it is a deliberate step.

## FAQ

### What is Coqui TTS used for?

It is a library for advanced text-to-speech generation, with pretrained models in more than 1100 languages, tools for training new models and fine-tuning existing models in any language, and utilities for dataset analysis and curation. The released flagship is XTTS, described as a production model that can speak 13 languages, with XTTSv2 announced separately at 16 languages, and Bark is available for inference with unconstrained voice cloning.

### What is the quality of Coqui TTS?

The project answers this on a TTS Performance page whose own note says the underlined TTS* and Judy* entries are internal models that are not released open-source and are shown to indicate potential, while models prefixed with a dot, such as .Jofish, .Abe and .Janice, are real human voices. Quality is otherwise documented per model on the readthedocs site rather than by a single benchmark.

### Is Coqui TTS open-source?

The repository is under the MPL-2.0 with LICENSE.txt in the tree, and the code plus most models are published, although the performance page explicitly marks some models as unreleased. The more consequential detail is the timeline: the newest release is v0.22.0 from 2023-12-12 and the last push to the repository was on 2024-08-16, so an open licence sits on code that has not been touched in over two years.

### how to use coqui ai tts

Install from PyPI if you only want to synthesize speech with the released models, and check your interpreter first, because setup.py raises a RuntimeError unless you are on Python 3.9 up to but not including 3.12. The toolkit is tested on Ubuntu 18.04, and the system level dependencies it expects are espeak-ng and libsndfile1, both of which the Dockerfile installs alongside gcc, g++ and make before running make install.

### Is voice AI TTS free?

The library carries the MPL-2.0 licence and installs from PyPI with no paid tier, so the cost is compute rather than licence. The dependency list requires torch>=2.1 and torchaudio, the default container base is nvidia/cuda:11.8.0-base-ubuntu22.04, and the training story is built around detailed logs on the terminal and Tensorboard plus a Trainer API, which means hardware is where the bill lands.

### coqui ai tts alternative

The repository does not point at a replacement. What it does point at is the upstream projects behind two of the models it wraps, tortoise-tts and the Bark repository from suno-ai, plus the Fairseq collection of roughly 1100 models, so those are the places to look if you want one model rather than a research framework spanning dozens of architectures.

## Sources

- [Official documentation](http://coqui.ai)
- [Official README](https://github.com/coqui-ai/TTS#readme)
- [Project repository](https://github.com/coqui-ai/TTS)
- [Release notes](https://github.com/coqui-ai/TTS/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/coqui-ai-tts
