KittenTTS: a 25 MB ONNX text-to-speech model that runs on CPU
State-of-the-art TTS model under 25MB 😻
At a glance
- What is it?
- KittenTTS is an Apache-2.0 Python library that runs ONNX text-to-speech models of 15M to 80M parameters without a GPU. It is a developer preview, and the int8 build is the one users have complained about.
- Who is it for?
- KittenTTS fits offline or CPU-only synthesis where 25 to 80 MB of model weights matter more than voice cloning or multilingual coverage: the eight built-in voices and the 24 kHz output are the whole feature set. It does not fit anyone who needs a stable API today, since the README labels it a developer preview, and it does not fit languages other than English, because the roadmap still lists multilingual TTS as unreleased.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 42 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem KittenTTS solves, and for whom
Most open text-to-speech stacks assume a GPU somewhere in the pipeline. KittenTTS takes the opposite position: the README describes it as an ONNX-based library whose models run from 15M to 80M parameters, 25 to 80 MB on disk, with synthesis on CPU and no GPU required. That size class is the point. A 25 MB model can sit inside a desktop application, a container image, or a machine with no accelerator, which is a different deployment story from a multi-gigabyte checkpoint that needs a card to answer in reasonable time.
The audience follows from that. Developers building offline narration, accessibility readouts, or test fixtures for a speech pipeline are the natural users. The README also points at a hosted option: a free Kitten TTS API at platform.kittenml.com, plus a Hugging Face Spaces demo for browser testing. If you only need a few seconds of speech to check a layout, the demo is faster than installing anything.
The library is English-only in practice. The roadmap lists multilingual TTS as an unchecked item, and the text normalizer defaults to en-US. Anyone who needs Chinese, Spanish, or code-switched output is looking at the wrong project today, and the repository does not claim otherwise.
How the ONNX pipeline actually runs
The dependency list in pyproject.toml tells you the architecture. phonemizer and espeakng_loader convert text to phonemes, onnxruntime executes the acoustic model, soundfile writes audio, numpy carries the sample array, and huggingface_hub downloads weights. There is no PyTorch or CUDA dependency in the base requirements, which is why the default path is CPU.
Data flow is linear. Text enters, the phonemizer produces a phoneme sequence, the ONNX graph turns that into audio, and generate() returns a NumPy array at 24 kHz. The same module exposes normalize_text(), which handles the text side on its own: numbers, currencies, units, abbreviations, dates, times, and URLs. Its signature takes locale="en-US" and return_spans=False, and with return_spans=True it returns character spans mapping original segments to their normalized form. That span output is unusual for a small TTS package and useful if you highlight text while it is spoken.
One inconsistency is worth flagging. pyproject.toml declares the license as Apache 2.0, matching the LICENSE badge in the README, but setup.py carries the classifier "License :: OSI Approved :: MIT License" and a Development Status of Alpha. If you are auditing a dependency for redistribution, the classifier in setup.py is stale and the Apache 2.0 text is the one the repository actually ships.
Installing KittenTTS and generating a first WAV
There is no PyPI release in the repository files. The README installs the library from a wheel attached to the 0.8.1 GitHub release, and the prerequisites are Python 3.8 or later with pip. Run this in a virtual environment, which the README recommends to avoid dependency conflicts:
pip install https://github.com/KittenML/KittenTTS/releases/download/0.8.1/kittentts-0.8.1-py3-none-any.whlThat command pulls in phonemizer, espeakng_loader, onnxruntime, soundfile, numpy, and huggingface_hub. The first model load downloads weights from the Hugging Face Hub, so the initial run needs network access; after that the files are cached.
A minimal synthesis script loads a model by repository ID, generates audio for one voice, and writes it to disk. Note that the model name is passed as a Hugging Face repo ID such as KittenML/kitten-tts-mini-0.8:
from kittentts import KittenTTS
import soundfile as sf
model = KittenTTS("KittenML/kitten-tts-mini-0.8")
audio = model.generate("This high-quality TTS model runs without a GPU.", voice="Jasper")
sf.write("output.wav", audio, 24000)You should get output.wav at 24 kHz. The README's advanced example adds speed control and a direct file writer, and both are worth knowing before you build a wrapper:
audio = model.generate("Hello, world.", voice="Luna", speed=1.2)
model.generate_to_file("Hello, world.", "output.wav", voice="Bruno", speed=0.9)
print(model.available_voices)available_voices returns ['Bella', 'Jasper', 'Luna', 'Bruno', 'Rosie', 'Hugo', 'Kiki', 'Leo']. For GPU execution the README points to requirements_gpu.txt and a backend argument: KittenTTS("KittenML/kitten-tts-mini-0.8", backend="cuda"), with example_cuda.py in the repository.
The int8 model, the preview label, and other limits
The most concrete warning in the README concerns the smallest build. It states that some users have reported issues with kitten-tts-nano-0.8-int8 and asks them to open an issue. That is the 25 MB variant, the one whose size makes the project interesting on constrained hardware. If your reason for choosing KittenTTS is the 25 MB figure, treat that variant as unverified and test it on your own text before designing around it. The fp32 nano build is listed at 56 MB, and micro at 41 MB.
Second, the project calls itself a developer preview and says APIs may change between releases. The version history supports that reading: 0.1 in August 2025, then 0.8 in February 2026 and 0.8.1 a few days later. Pinning 0.8.1 and reading release notes before upgrading is the sane approach.
Third, the default voice argument in the API tables is "expr-voice-5-m", which does not appear in the eight-name list returned by available_voices. The README does not explain the relationship between the two, so pass a voice name explicitly rather than relying on the default.
Finally, the model lineup is small and English-only, and the roadmap's unchecked items (optimized inference engine, mobile SDK, higher quality models, multilingual TTS, KittenASR) are all still open. KittenTTS is the wrong tool for voice cloning, for languages outside English, and for anyone who needs a stable interface with a deprecation policy.
KittenTTS compared with Piper and Coqui TTS
Piper is the closest alternative: a CPU-oriented neural TTS system distributed as standalone binaries with ONNX voice files, aimed at embedded and offline use. The difference in approach is packaging. Piper ships prebuilt executables and per-voice model files that you invoke from a shell or bind through a C API; KittenTTS is a Python package whose model is fetched from the Hugging Face Hub at runtime and whose text pipeline is exposed as normalize_text() with locale and span options. If your application is already Python, KittenTTS avoids a subprocess boundary. If it is not, Piper's binary model is the easier integration, and its voice catalogue is broader and multilingual, which KittenTTS does not offer yet.
Coqui TTS sits at the other end. It is a full training and synthesis framework with a large model zoo, and it expects a heavier dependency stack; the trade-off is that you can fine-tune voices rather than choose from eight fixed ones. KittenTTS does not document a training path at all, so custom voices are a commercial-support conversation, not a repository feature. The README lists commercial support for integration assistance, custom voices, and enterprise licensing.
There is also the hosted route. platform.kittenml.com offers a free API, and the Hugging Face Spaces demo runs the model in a browser. Both remove the install step, and both put the synthesis off your machine, which is the opposite of the reason most people pick a 25 MB local model.
Maintenance, licensing, and what an upgrade costs
The repository is not archived, and the last push was on 2026-08-19. Release cadence has been uneven: 0.1 in August 2025, then 0.8 and 0.8.1 in February 2026, with no release since. Recent commits do not imply a stable interface, and the README's own preview warning says APIs may change.
The upgrade surface is small, which keeps migration cheap. The public API is a constructor plus generate(), generate_to_file(), available_voices, and normalize_text(). If a release changes the voice list or the default voice, code that passes voice explicitly survives it. The real cost is the model download: swapping between nano, micro, and mini means new weights from the Hugging Face Hub, and cache_dir lets you control where they land.
Licensing is Apache 2.0 per pyproject.toml and the LICENSE file, which permits commercial use and modification with attribution and notice requirements. Two caveats are worth a lawyer's eye rather than mine: setup.py still declares an MIT classifier, so a scanner may report a mismatch, and the models live in separate Hugging Face repositories whose own licence terms the README does not restate. Check the model card for the variant you ship before assuming the library licence covers the weights.
Editorial conclusion
KittenTTS fits offline or CPU-only synthesis where 25 to 80 MB of model weights matter more than voice cloning or multilingual coverage: the eight built-in voices and the 24 kHz output are the whole feature set. It does not fit anyone who needs a stable API today, since the README labels it a developer preview, and it does not fit languages other than English, because the roadmap still lists multilingual TTS as unreleased. Before committing, check that the model you pick actually loads: the README warns that some users have reported issues with kitten-tts-nano-0.8-int8, so verify that variant on your own text first and fall back to kitten-tts-micro-0.8 if it fails.
Frequently asked questions
What is KittenTTS?
It is an open-source text-to-speech library built on ONNX, with models from 15M to 80M parameters that run on CPU without a GPU. It ships eight built-in voices and outputs 24 kHz audio.
How do I use KittenTTS?
Install the 0.8.1 wheel from the GitHub release, then load a model by its Hugging Face repository ID and call generate() with a voice name. The README's example writes the returned NumPy array to a WAV file with soundfile at 24000 Hz.
What are the KittenTTS voices?
The README lists eight: Bella, Jasper, Luna, Bruno, Rosie, Hugo, Kiki, and Leo, returned by model.available_voices. Speech speed is adjustable per call with the speed parameter.
What is a KittenTTS alternative?
Piper is the closest one: it also targets CPU inference with ONNX voice files, but it ships as standalone binaries rather than a Python package and covers more languages. Coqui TTS is the alternative if you need to train or fine-tune voices, which KittenTTS does not document.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/kittenml-kittentts)