# sanoTTS: a 1.4M-parameter neural voice for a $3 chip and the browser

> sanoTTS is a GPL-3.0 text-to-speech family from 294k to 2.3M parameters that runs on an ESP32-S3 or in WebAssembly. The interesting part is the constraint, not the voice quality.

**Ampixa/sanoTTS** — sanoTTS (सानो = 'small' in Nepali): a ~1.4M-param neural TTS that runs on a $3 chip or in the browser. Leads SCOREQ/UTMOS in the sub-15M class.

- Repository: https://github.com/Ampixa/sanoTTS
- Website: https://ampixa.github.io/sanoTTS/
- Stars: 755 · Forks: 76
- Language: C
- License: GPL-3.0
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/ampixa-sanotts

## What sanoTTS is actually for

The name is Nepali for "small," and the size claim is the whole product. sanoTTS is a family of neural text-to-speech voices between 294k and 2.3M parameters. That range is one to two orders of magnitude below the models people normally reach for when they want speech that does not sound like a 1990s screen reader.

The target user is not a voice-over studio. It is someone with an ESP32-S3 on a breadboard, a GPIO pin, an LM386 amplifier and a speaker, who wants the board to say something without a network round trip. The README states the measured real-time factor on an ESP32-S3 is 0.383 with the Xtensa LX7 SIMD kernels the Arduino library assembles by default, which is 2.6x faster than real time. The same README gives the scalar C figure on the same board as 1.58, slower than real time. That gap is the reason the project exists in C at all.

The second audience is web developers who want speech synthesis that never leaves the visitor's machine. The browser demo runs an espeak-ng phonemizer compiled to WebAssembly, then that voice's own neural stack, all client-side. No server, no upload. For anything with a privacy or latency constraint, that architecture is the point, not a bonus.

## How the pipeline is split between phonemizer and acoustic model

sanoTTS does not do end-to-end text-to-waveform. The repository layout shows two separate WebAssembly modules: `web/snt_g2p.js` with `snt_g2p.wasm` and a `snt_g2p.data` file, and `web/snt_voice.js` with `snt_voice.wasm`. The naming is explicit. G2P means grapheme to phoneme, and the `.data` sidecar is the phoneme table that espeak-ng needs.

So the data flow is: text goes into the espeak-ng phonemizer in WASM, phonemes come out, and the per-voice neural stack turns phonemes into audio. The voice weights are fetched lazily on first use of that voice and are never bundled with the runtime. That separation is why the wasm runtime is roughly 700 KB gzipped over the wire while a voice can be as small as 337 KB.

It also explains a licensing detail worth noticing. The phonemizer is espeak-ng, which the README lists as included. The project itself is GPL-3.0, and the repository root also contains a `LICENSE.MIT` file. The README does not explain which parts fall under which licence, and I would not guess.

The precision story is equally concrete. The README gives per-voice sizes of 337 KB for heart-nano in int8, about 3 MB for fp16, and 5.5 to 8.7 MB for fp32. For the ten voices added on 2026-09-08, fp16 is described as the reference rather than an approximation, because every number measured for them was measured on the fp16 package, and the browser widens them to fp32.

## Installing sanotts and getting a wav file out

The Python path is the shortest route to hearing the thing. The README gives `pip install sanotts` and a CLI that writes a wav file. Voices download on first use into `~/.cache/sanotts/`, fetched from Hugging Face with the GitHub voices-v1 and voices-v2 releases as fallback.

```bash
pip install sanotts

sanotts say "Hello from a two megabyte voice." --voice amy -o hello.wav
```

After that command you should have `hello.wav` in the current directory, synthesized by the 1.46M-parameter amy voice. The README also shows a Vietnamese example with `--voice vi -o xinchao.wav`, which is a useful check that the multilingual voices are actually wired into the CLI and not only the browser demo.

If the download host is the problem, the README documents an environment variable to pin one source:

```bash
SANOTTS_VOICE_SOURCE=hf sanotts say "Hello" --voice amy -o hello.wav
```

The accepted values are `hf` and `github`. There is also a library call, which returns numpy audio at 22.05 kHz:

```python
import sanotts
result = sanotts.synthesize("Hello world", voice="amy")
```

The README states inference is pure numpy, with no torch and no onnxruntime. That is a real deployment advantage: the runtime dependency list is short enough to fit in a Lambda or a small container without pulling a deep learning framework.

For the microcontroller, the library lives in `arduino/`. You add it to the Arduino IDE as a .zip library, or list it in `platformio.ini`:

```ini
lib_deps = https://github.com/Ampixa/sanoTTS.git
```

Board support, memory guidance and flashing the model blobs are in `arduino/README.md`, which the top-level README points to rather than duplicating. The ESP32-S3 is the board marked as supported.

## The browser deployment is static files, and that is the whole trick

There are two documented ways to self-host the web build, and they differ in how much you want to own.

Option A uses the npm package. You copy the package's `dist/` directory and this repository's `web/voices/` directory to your own static host, then point the loader at them:

```js
import { SanoTTS, playAudio } from 'sanotts-web';

const tts = await SanoTTS.load({
  assetBase: 'https://your-cdn.example.com/sanotts/',
});
const result = await tts.synthesize('Hello from my own server.', {
  voice: 'amy',
  voiceBase: 'https://your-cdn.example.com/sanotts/',
});
playAudio(result);
```

Option B skips npm entirely. You copy `snt_g2p.js`, `snt_g2p.wasm`, `snt_g2p.data`, `snt_voice.js`, `snt_voice.wasm` and the `voices/` directory onto a static host and load them the way `web/index.html` does. The README's own snippet is a stub that stops at initialization and tells you to read `web/index.html` for the full sequence.

That last sentence is the honest summary of the browser documentation. The loading contract is documented. The synthesis sequence is documented by example file, not by prose. If you are integrating this into an existing app rather than copying the demo, budget time to read the demo source.

One number to plan around: the wasm runtime is about 700 KB gzipped over the wire, of which the espeak-ng G2P ships roughly 2.5 MB uncompressed including its phoneme table, and the acoustic and decoder wasm adds about 40 KB. The phonemizer, not the neural model, dominates the initial payload.

## Where sanoTTS is the wrong tool

Naturalness is unmeasured. The README says this plainly: the word error rates reported for the ten languages added on 2026-09-08 run from 0.038 for Portuguese to 0.392 for Romanian, measured over 16 sentences per language, and the README describes that as intelligibility rather than naturalness and states that no one has run a listening test. If your product depends on prosody, emotion or long-form narration quality, you are choosing a model whose perceptual quality has not been evaluated by its own maintainers. Pick something with a published MOS study.

Language coverage is uneven. The browser demo carries 30 voices across 16 languages, but the pip package ships 27 voices across 13 languages, and Nepali, Hindi and Chinese are browser-only for now. If your deployment target is Python and your language is one of those three, the CLI will not serve you.

Platform support is narrow where it matters most. The README names the ESP32-S3 specifically, and the release notes for the Arduino library describe a 294k voice used in a board benchmark. Other microcontrollers are not claimed. The `esp32c3/` directory exists in the repository, but the README's board support line does not mention that part, so I would not assume it works without checking `arduino/README.md` and the board notes yourself.

Finally, the Python package name is not the project name. The distribution is `saanotts` in `pyproject.toml` while the CLI it installs is `sanotts`, and the package version there is `0.1.0.dev0` with `requires-python = ">=3.11"`. The README's install line and the packaging metadata are not identical strings, which is a small thing that will cost someone ten minutes.

## How it compares to Piper and to cloud TTS APIs

The closest open-source comparison is Piper, and the difference is architectural rather than cosmetic. The `pyproject.toml` describes the project as "Research and tooling for compact Piper/VITS-distilled TTS," and `piper-tts>=1.3` appears in the `research` optional dependency group. So sanoTTS is downstream of the Piper and VITS lineage rather than a parallel invention.

The divergence is where inference runs. Piper's own runtime targets CPUs on machines with an operating system. sanoTTS ships a portable C runtime with a golden-gate test (`make test-runtime` builds and runs it) and an Arduino library, so the same family of distilled models is being pushed down to a board with no OS. That is the claim worth evaluating: not that the voices are better, but that the deployment floor is lower. The scalar-versus-SIMD numbers in the README (1.58 versus 0.383 RTF) show how much of that floor depends on hand-written Xtensa kernels rather than on the model itself.

Against cloud APIs the trade is obvious and one-sided in each direction. A cloud API gives you better voices and no size budget. sanoTTS gives you no per-request cost, no network dependency, and no audio leaving the device. The README's browser demo makes the privacy argument implicitly by never uploading text. If your constraint is a data residency rule or an offline device, the cloud option is not on the table at all, and voice quality becomes secondary.

## Maintenance, licensing and what a fork costs you

The repository is not archived, and the last push was on 2026-09-16, one day before this writing. Recent activity is real: a voices-v2 release on 2026-09-04 adding heart and heart-nano, an Arduino library 0.1.0 release on 2026-09-03, and a 2.27M tiny model release on 2026-08-04. The ten-language expansion is dated 2026-09-08 in the README.

That is a fast-moving research repository, and the version string in `pyproject.toml` (`0.1.0.dev0`) is consistent with it. Expect interface churn in the Python package. The Makefile is the part that looks most stable and is the best entry point for a contributor: `make check` compiles the Python, runs the dependency-free package tests, and runs the MCU golden gate in one command. `make test-python` runs the unittest discovery, `make test-runtime` builds the portable C runtime, and `make preservation-verify RESTORE_ROOT=/tmp/recovery` verifies a local asset recovery root, which implies the maintainers treat voice blobs as something that can go missing.

On licensing: the project is GPL-3.0. If you link the Arduino library or the C runtime into a product, GPL-3.0 obligations travel with the distribution. The repository also contains a `LICENSE.MIT` file at the root, and the README does not state which components it covers. The `pyproject.toml` carries no license field at all. I am not giving legal advice, but the practical point is that you cannot determine the licence boundary from the README alone, and you should read the two licence files and the per-directory headers before shipping anything closed.

## Conclusion

Adopt sanoTTS if your problem is fitting speech into a microcontroller or a static site with no server and no NPU, and if you can accept that naturalness has not been measured by a listening test. Skip it if you need studio-quality prosody, a documented API surface, or permissive licensing, because the project is GPL-3.0 with an additional LICENSE.MIT file whose scope the README does not explain. Before committing, open arduino/README.md and confirm your board is listed, then run the Python CLI once and listen to the output yourself.

## FAQ

### What does sanoTTS do?

It is a family of tiny neural text-to-speech voices between 294k and 2.3M parameters that run without a cloud service or an NPU. The README states it runs real-time on an ESP32-S3 and in the browser through WebAssembly.

### Is sanoTTS worth adopting?

It depends on whether your constraint is size and offline operation rather than voice quality. The README states that intelligibility has been measured but that no one has run a listening test, so naturalness is unverified, while the deployment floor is genuinely low.

### Is sanoTTS a good long-term investment?

The repository is not archived and the last push was on 2026-09-16, with releases in September 2026, but the Python package version is 0.1.0.dev0 and the README describes a research project. Treat the API as unstable and pin a version.

### What is sanoTTS best known for?

The size claim. The README describes it as the smallest neural TTS family known, at 294k to 2.3M parameters, and reports an RTF of 0.383 on an ESP32-S3 with the Xtensa LX7 SIMD kernels the Arduino library assembles by default.

## Sources

- [Ampixa/sanoTTS on GitHub](https://github.com/Ampixa/sanoTTS)
- [License: GPL-3.0](https://github.com/Ampixa/sanoTTS/blob/master/LICENSE)
- [Project website](https://ampixa.github.io/sanoTTS/)
- [README](https://github.com/Ampixa/sanoTTS/blob/master/README.md)
- [Releases](https://github.com/Ampixa/sanoTTS/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ampixa-sanotts
