CLI tool
kyutai-labs/pocket-tts avatar
kyutai-labs/pocket-tts

Pocket TTS: a hundred million parameter voice model built to run on two CPU cores

A TTS that fits in your CPU (and pocket)

9,813 stars1,033 forksPythonMIT

At a glance

What is it?
Kyutai's text-to-speech model skips the GPU entirely, streams audio about 200 milliseconds to the first chunk, and ships voice cloning alongside a catalog whose voices each carry their own license.
Who is it for?
Pocket TTS is one of the few speech models where the interesting constraint is the hardware rather than the model. A hundred million parameters, two CPU cores and roughly 200 milliseconds to the first audio chunk is a different deployment story from a GPU pipeline, and it means voice cloning runs on a laptop with no accelerator at all.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The headline claim is about the machine, not the model

The project description is one sentence: a TTS that fits in your CPU, and pocket. Everything else in the README is arranged around that idea.

The main takeaways list leads with runs on CPU, a small model size of 100M parameters, audio streaming, and low latency at about 200 milliseconds to get the first audio chunk. Then it claims faster-than-real-time performance at roughly 6x real-time on a MacBook Air M4 CPU, and states the resource budget as two CPU cores. A Python API and CLI, voice cloning, and six languages follow on the list.

Those are unusually concrete numbers for a speech project. Most readmes in this space describe architecture or output quality. This one tells you the machine, and two cores is a small enough claim to make that it is the interesting part. If it holds, you can run this in a container on a cheap VM, on a host with no GPU driver installed, or on hardware where asking for a GPU was never an option.

The README is also honest about the ceiling. Non-English languages ship in larger 24-layer variants that are higher quality but slower, selected with a suffix such as `--language italian_24l`. English is the default. Handling infinitely long text input is claimed as well, which matters more than it sounds, because streaming systems commonly buffer a whole utterance before emitting anything and fall over on a long document.

There is a hosted demo so you can hear the output before installing anything, plus a Hugging Face model card, a technical report from January 2026, and a paper on arXiv at 2509.06926.

Installing without accidentally pulling three gigabytes of CUDA

The Python side looks ordinary. It supports 3.10 through 3.14, requires PyTorch 2.5 or newer, and states that it does not require the GPU build of PyTorch.

bash
uvx pocket-tts generate
# or if you installed it manually with pip:
pocket-tts generate

The recommended installer is `uv`, because it resolves dependencies on the fly in an isolated environment. A plain `pip install pocket-tts` also works.

That last sentence hides a real trap, and release v3.1.0 exists partly because of it. The v3.1.0 notes document CPU-only installation and say plainly that pip pulls around 3 GB of CUDA wheels by default. So on Linux, installing this deliberately CPU-targeted project the ordinary way can bring in gigabytes of GPU libraries you will never call. The README links to a dedicated CPU-only installation section for exactly this case, and it is the first thing to read before the first install on a Linux host.

The dependency pins in `pyproject.toml` show why this happens. PyTorch is declared as `torch>=2.5.0`, which on the default index resolves to the CUDA build. The numpy entry carries a comment that gives a sense of how carefully these floors were chosen: a version was tried and rejected because the generated audio came out wrong, so a later minimum is required. That is the kind of note you only find in a project where somebody chased a numerical bug to its source.

Optional extras are declared separately rather than folded in. `audio` adds `soundfile`, and `quantize` adds `torchao`, so quantization is an explicit opt-in. The package version in the same file reads 3.1.0, matching the latest tag, and the build backend is Hatchling with `pocket_tts` as the packaged directory.

Running it as a service rather than a command

The second CLI command runs a local HTTP server instead of writing a file:

bash
uvx pocket-tts serve
# or if you installed it manually with pip:
pocket-tts serve

You then navigate to localhost on port 8000 for a web interface. The README gives the reason to prefer it: it is faster than the command line because the model stays in memory between requests. That is the whole argument. Loading a hundred million parameter model per invocation is the dominant cost, and the README says as much when it recommends serve for trying multiple voices and prompts quickly.

The container path is built on that same command. The `Dockerfile` is six meaningful lines: it starts from the `uv` image on Debian, copies `pyproject.toml`, `uv.lock`, `README.md`, `.python-version` and the `pocket_tts` package, runs `uv run pocket-tts --help` as a build-time smoke test, and sets the entrypoint to `uv run pocket-tts`.

The compose file maps port 8000, runs `serve --host 0.0.0.0`, restarts unless stopped, and mounts two named volumes: one for `/root/.cache/pocket_tts` and one for `/root/.cache/huggingface`. Caching the model outside the container is the detail that matters, since it means a redeploy does not re-download weights.

The repository also carries `deploy.sh` and `swarm-config.yaml`, so the authors expect real deployment rather than local experiments, and `mkdocs.yml` with a `docs/` directory for the documentation site the README links to.

Voice cloning, and the one caveat the README insists on

The `--voice` argument accepts either a name from the catalog or a plain wav file. The second case is the voice cloning path: you supply a short recording and the model reproduces that speaker.

The README attaches one instruction to this feature and it is worth taking seriously, because it is the kind of thing that gets skipped. It recommends cleaning the sample first, and gives the reason directly: the audio quality of the sample is also reproduced. That is the honest description of what the feature does. You are not getting a clean voice out of a noisy recording, you are getting a faithful copy of whatever you supplied, defects included.

The catalog runs to more than twenty named voices across English, Italian, Spanish, German, Portuguese and French. Their provenance is visible in the sample paths: several come from VCTK, at least one from the Expresso corpus, several from a voice-zero directory, and a few from a voice-donations folder. Those names matter because of what comes next.

Voice licensing is handled per voice, not once for the model. The README directs you to a Hugging Face repository, `kyutai/tts-voices`, specifically to read the licenses for each voice, and the individual voice entries link to sample files in that repository or to paths inside the model repository. If you are planning to ship audio, that page is the document that matters, not the model card.

There is also community extension. Release v3.0.2 added the Czech model to the README and introduced a section for models trained by the community, with an invitation to open a pull request. Release v3.1.0 took the training side further, adding a documented path for finetuning a released model into a new language.

Models, configuration paths, and where the defaults come from

A few configuration details determine what you actually get, and the README is specific about all of them.

Language is chosen with `--language`, and the default is `english`. The flag works with `generate`, `export-voice` and `serve`. Non-English languages have both a compact variant and a 24-layer variant, the latter described as higher quality but slower.

Weights come from `--config`, which accepts three forms: a local YAML path, an `https://` URL, or an `hf://` path in the form `hf://<repo_id>/<path>` with an optional `@revision`. That third form is what makes the community models section work, since a fine-tuned model is just a different repository and path.

Voice defaults have moved recently. Release v3.0.2 made `--voice` default to alba with voice cloning enabled on a custom config, which is a quiet behavioral change if you had a script relying on an earlier default. Release v3.0.1 was published the same day as v3.0.0 with the single-line description Build fixed, so if you jumped versions, 3.0.1 is the one you want.

The training code itself arrived in August 2026, announced in a note at the top of the README. It lives in `training/`, has its own README, and is excluded from the inference dependency set, which the comment in `pyproject.toml` states explicitly. Inference needs none of the dev group, so the training dependencies do not weigh on a deployment. The repository also carries `tests/`, `scripts/`, `AGENTS.md` and `CONTRIBUTING.md`, and the v3.1.0 notes record that the training tests were moved into CI with explicit zip lengths.

How to tell what is measured and what is projected

The speed claims are specific enough to be useful and narrow enough that you should know their conditions. The 6x real-time figure is stated for a MacBook Air M4 CPU, and the two-core budget is stated separately. Neither number says anything about older hardware, about many concurrent requests, or about the 24-layer language variants.

The 200 millisecond figure is a time to first audio chunk, not to a finished file. That distinction is the right one for a streaming model, and it is the number that matters for interactive use. It also depends on prompt length and on which language variant you loaded.

Against that, release v3.1.0 added a speed table whose rows were measured on H100 GPUs reading from local storage. That is a different question from the CPU figures, and it tells you the project is now tracking inference throughput in a data center configuration too. The same release includes a change to the lsd depth distillation training pass, which is about quality per unit of compute rather than about latency.

The 9,584 stars and 1,001 forks against 55 open issues suggest a project with broad adoption and a manageable issue load. The last push was on 2026-09-21, and the MIT license applies to the code. What the license does not settle is what you may do with the voices themselves, which is why the per-voice page in the voices repository is the place to look before this leaves your machine.

Editorial conclusion

Pocket TTS is one of the few speech models where the interesting constraint is the hardware rather than the model. A hundred million parameters, two CPU cores and roughly 200 milliseconds to the first audio chunk is a different deployment story from a GPU pipeline, and it means voice cloning runs on a laptop with no accelerator at all. The thing to understand before you commit is licensing, not capability. Every voice in the catalog carries its own terms on Hugging Face, voice cloning reproduces the quality of the sample you feed it, and the project's own advice is to clean that sample first. Start with the generate command on the default alba voice to hear what you get, read the per-voice license page before shipping anything, and prefer the serve command over repeated one-shot invocations if you are generating more than a line at a time.

Frequently asked questions

What is pocket TTS?

It is Kyutai's text-to-speech model, about 100M parameters, built specifically to run on CPUs rather than GPUs. The project states it uses two CPU cores, streams audio about 200 milliseconds to the first chunk, and reaches roughly 6x real-time on a MacBook Air M4 CPU. It ships a Python API, a CLI, and a server mode.

How do I install pocket-TTS?

The project recommends `uvx pocket-tts generate` for a one-off run, or `pip install pocket-tts` to install manually. Python 3.10 through 3.14 and PyTorch 2.5 or newer are required. On Linux, read the CPU-only installation section first, because a default pip install pulls around 3 GB of CUDA wheels even though the project does not need a GPU.

Can Pocket TTS clone a voice from my own audio?

Yes. The `--voice` argument accepts a plain wav file in addition to the named catalog voices. The README's advice is to clean the sample before using it, because the audio quality of the sample is reproduced along with the voice. Sample files for the cataloged voices live in a separate Hugging Face repository that also documents each voice's license.

Which languages does Pocket TTS support?

English, French, German, Portuguese, Italian and Spanish, with Czech added in release v3.0.2. English is the default for `generate`, `export-voice` and `serve`. Non-English languages also ship larger 24-layer variants that are higher quality but slower, selected with a suffix such as `--language italian_24l`.

How do I run Pocket TTS as a local server?

Use `uvx pocket-tts serve` or `pocket-tts serve`, then open port 8000 in a browser. The README recommends this over repeated command line calls because the model stays in memory between requests. The provided compose file runs the same command with `--host 0.0.0.0` and mounts caches for the model weights and the Hugging Face hub.

Official sources

  1. Issues
  2. kyutai-labs/pocket-tts on GitHub
  3. License: MIT
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/kyutai-labs-pocket-tts.svg)](https://hysenlabs.com/projects/kyutai-labs-pocket-tts)