Open-source project
shivammehta25/Matcha-TTS avatar
shivammehta25/Matcha-TTS

Matcha-TTS: the build pins Cython and numpy exactly while the runtime leaves both open, and a fourth console command ships undocumented

[ICASSP 2024] 🍵 Matcha-TTS: A fast TTS architecture with conditional flow matching

1,362 stars218 forksJupyter NotebookMIT

At a glance

What is it?
Matcha-TTS is the official implementation of a 2024 conference paper on non-autoregressive speech synthesis with conditional flow matching, from a group at a Swedish university. It installs from the package index, synthesises from a one line command, trains through configuration overrides, and exports to a portable graph format for deployment without the framework. Its packaging is where the interesting details are, and so is its training documentation.
Who is it for?
Matcha-TTS fits someone who has read the paper and wants the reference implementation to reproduce a figure, or needs a small non-autoregressive synthesiser they can export for inference without a Python framework in the loop. It does not fit someone who wants a maintained package to depend on.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 6 days ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Four console commands ship and one is never mentioned

The packaging metadata declares four console scripts:

code
matcha-tts-get-durations=matcha.utils.get_durations_from_trained_model:main

alongside the synthesis command, the web interface command, and the statistics command used during training. The readme documents three of them. The synthesis command is covered in detail with a section per flag, the web interface gets one line, and the statistics generator appears in the training walkthrough because you are told to run it to get the normalisation numbers. The fourth, which extracts durations from a trained model, has no mention anywhere in the visible documentation. That is a small thing, and it is also the pattern worth noticing: the readme is organised around what you are trying to do, so anything that is a utility rather than a step has nowhere to live. If you are reading the source to work out how to get durations out of your own checkpoint, the command name is the only clue.

The build environment is pinned to two exact versions the runtime leaves open

Compare the two dependency lists. The build requirements are pinned hard:

code
requires = ["setuptools", "wheel", "cython==0.29.35", "numpy==1.24.3", "packaging"]

The runtime requirements list the same two libraries with no version at all, and pin other things instead: a training framework at a floor, a configuration library at an exact version with its logger and its sweeper alongside, and a diffusion library carrying a comment saying which version it was developed against. So the extension is compiled against one Cython and one numpy, and then linked into an environment where the resolver may have chosen quite different ones. That is not automatically wrong, since the compiled part only needs numpy's headers at build time, but it does mean the thing that builds your install and the thing that runs it are resolved by two different mechanisms. It also constrains what you can install this on: if you want a newer numpy, you are now editing a build requirement, not a runtime one, and the extension will be recompiled against it.

The pre-trained weights arrive from a shared Drive folder on first run

The models are not in the repository and not in a release. The readme links a folder on a consumer file-hosting service and states that the pre-trained models will be downloaded automatically by the command line or the web interface the first time either runs. So the very first command a new user types triggers a download from a folder whose contents the project does not version, whose contents can change, and whose availability is outside the project's control. The download helpers are runtime dependencies rather than install-time ones, which is why they appear in the requirement list alongside the training stack. There is a hosted demo on a model hub and a local browser demo linked as well, so a reader who wants to try the thing before committing to an install has somewhere to go. For a research implementation that is a reasonable trade, and for a deployment it is the first question to ask.

Normalisation statistics are computed by a command and pasted in by hand

Step five of the training walkthrough is the one that will bite someone following it quickly. You generate the dataset statistics with a command, which prints two numbers, a mean and a standard deviation for the mel spectrograms:

bash
matcha-data-stats -i ljspeech.yaml

and then you are told to update those values in the dataset configuration file under a specific key, with the output block showing the six decimal places to copy. Nothing enforces it. There is no check in the training script that compares the configured statistics against the data, and the walkthrough does not say what happens if you skip it. A run with missing or wrong statistics trains happily and produces a model that fails quietly, which is the worst failure mode in a training pipeline because it looks like success until you listen to it. The same walkthrough also makes you prepare file lists by following the setup of a different project, which is the normal way this documentation is written and is worth knowing before you start.

The alignment path is a compiled extension, so every install builds code

The packaging metadata declares a single Cython extension, targeting the module that does monotonic alignment, with a source file in that same directory, compiled at language level three, and it passes numpy's include directory into the build. That is the whole reason the build requirements pin an exact Cython and an exact numpy: the alignment code needs numpy's headers to compile. So this is not a pure Python install. On any platform without a published wheel you are compiling, which means a compiler, Python headers and a working numpy at build time, and on an interpreter or architecture combination nobody has built before you are the person finding out whether it works. The readme does say the conda environment is suggested rather than required and the metadata declares a floor of Python 3.9, so the project is reasonably open about the range, but the combination of a compiled extension and an exact build pin is the practical constraint on where this will install.

Pytest is pointed at a tests directory the repository does not contain

The test configuration in the packaging metadata points at a directory:

code
testpaths = "tests/"

and the top level of the repository has no such directory. What is there is the source package, the configuration directory, scripts, notebooks, a makefile, and a synthesis notebook. The packaging metadata also excludes tests from the package discovery, which implies the directory existed when that exclusion was written. Along with the path, the configuration does two things worth knowing about. It turns on doctest collection for modules, so documentation examples inside the package are executed as tests, which is the mechanism that catches a broken example. And it filters two categories of warning out of the output entirely, deprecation and user warnings, which for a package built on a stack of versioned scientific libraries means the loudest signal that its dependencies have moved on is suppressed by design.

The ONNX section asks for a pre-release build of the framework

The export instructions end with a requirement and an explanation. Exporting to a portable graph needs a version of the deep learning framework newer than what the requirements allow, because a particular attention operator cannot be exported from older versions, and the text says you must install that newer version manually as a pre-release. Then the closing clause, which dates the whole section: until the final version is released. This instruction has been sitting in the readme since before the newest release tag, and the newest release tag itself is two years old. The rest of the ONNX section is the most carefully written part of the documentation. It separates the case where you export only the acoustic model, which writes spectrogram graphs and arrays, from the case where you embed the vocoder in the exported graph, which writes audio files in one run, and it offers a third path of supplying an external vocoder at inference. It also makes the step count an export time parameter rather than an input, with a stated default, so the same graph produces the same audio.

The version is read from a file, and the newest tag is two years old

There is no version literal in the packaging metadata. The build script reads a version file inside the source package, strips it, and passes it to the packaging call. That is a sound approach when releases are cut from tags, and it means there is exactly one place the number lives. It also means the number in the file and the tag on the release have to agree, and the readme does not say which wins if they disagree. The release history is three tags: one in January 2024 whose title describes fixing a download bug, a refactor and a phonemizer change, then May 2024, then August 2024. The default branch received a commit on 2026-09-28. So the source tree has moved well past the newest tag, and anyone installing from the package index gets the August 2024 state while anyone cloning gets the current one. The platform also classifies this repository as a notebook project rather than a Python one, because the notebooks outweigh the source by volume, which is a small clue about where to look first.

Editorial conclusion

Matcha-TTS fits someone who has read the paper and wants the reference implementation to reproduce a figure, or needs a small non-autoregressive synthesiser they can export for inference without a Python framework in the loop. It does not fit someone who wants a maintained package to depend on. Four things to check first. Which version you are getting, since the newest release tag and the last commit to the default branch are two years apart, so the package index copy is much older than the repository. That the install compiles code, because the alignment module is a Cython extension, which means a compiler and a pinned build toolchain on every machine, and on a platform with no wheel you are building from source. Where the pre-trained weights come from, because they are pulled from a shared file-hosting folder on first use rather than from a model hub or a release asset. And whether the training steps were followed, because the mel statistics have to be computed by a command and pasted into the configuration by hand, and a run that skips that produces a trained model rather than an error.

Frequently asked questions

How do I install Matcha-TTS?

Install it from the package index with pip, or from source by installing the git URL and then running an editable install in the cloned directory. A conda environment on Python 3.10 is suggested but optional, and the packaging metadata declares a floor of Python 3.9.

How do I synthesise speech with Matcha-TTS?

Run the command with a text argument, or point it at a file. Add a batched flag for batch synthesis from a file, and there are separate flags for speaking rate, sampling temperature and the number of Euler solver steps, plus one for a custom checkpoint path.

What commands does the Matcha-TTS package install?

Four console scripts: the synthesis command, the web interface command, a data statistics generator used during training, and a command that extracts durations from a trained model. The readme documents the first three and never mentions the fourth.

Where do the pre-trained Matcha-TTS models come from?

A linked folder on a consumer file-hosting service, downloaded automatically the first time the command line or the web interface runs. The repository does not version the weights, and there is no release asset or model hub entry for them.

Does Matcha-TTS support ONNX export and inference?

Yes, with the export and inference both run as module invocations. You can embed the vocoder in the exported graph so one run produces audio, or export only the acoustic model and supply a vocoder separately at inference. The solver step count is an export time parameter with a default of five, not a runtime input.

How do I train Matcha-TTS on my own dataset?

Prepare train and validation file lists the way a related project's setup describes, point the dataset configuration at them, run the data statistics command and copy the two printed mel values into that configuration, then start training with a make target or a configuration override, including a low memory variant and a multi device form.

Official sources

  1. License: MIT
  2. Project website
  3. README
  4. Releases
  5. shivammehta25/Matcha-TTS on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/shivammehta25-matcha-tts.svg)](https://hysenlabs.com/projects/shivammehta25-matcha-tts)