NeuTTS: Neuphonic's On-Device Text-to-Speech Models with Instant Voice Cloning
On-device TTS model by Neuphonic
At a glance
- What is it?
- NeuTTS is a collection of open-source, on-device text-to-speech models from Neuphonic, built on small LLM backbones and available in GGUF format for CPU and GPU inference, with instant voice cloning from three seconds of audio and multilingual support across English, French, German, and Spanish.
- Who is it for?
- NeuTTS is a practical choice for engineers building embedded voice agents, privacy-sensitive applications, or offline assistants who cannot route audio through a cloud API. The GGUF models run on mid-range consumer hardware and even on a Galaxy A25 5G.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 62 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What NeuTTS Addresses and Who It Is For
NeuTTS targets developers who need text-to-speech synthesis that runs entirely on-device, without sending audio data to a web API. The use cases described in the README include embedded voice agents, assistants on mobile devices, toys, and compliance-sensitive applications where audio content must not leave the local machine.
The library provides models at two size classes: NeuTTS-Air (approximately 360 million active parameters) and NeuTTS-Nano (approximately 120 million active parameters). A third model, NeuTTS-2E, targets approximately 125 million parameters and adds emotional control. All models support voice cloning from as little as three seconds of reference audio.
Architecture: LLM Backbones and the NeuCodec
NeuTTS models are built from small LLM backbones (lightweight language models optimized for text understanding and generation) combined with NeuCodec, Neuphonic's neural audio codec. NeuCodec operates at 50 Hz and uses a single codebook to achieve what the README describes as exceptional audio quality at low bitrates.
All models in the NeuTTS-Nano family accept phoneme input. NeuTTS-2E is the exception: it takes raw text input directly. The Air and Nano models require phoneme preprocessing using a phonemizer library, while 2E handles the text-to-phoneme step internally.
The context window is 2048 tokens, which the README states is sufficient for approximately 30 seconds of audio including prompt duration. Requests longer than that require either splitting the input or accepting that the model will truncate context.
Outputs are watermarked, which the README lists as a responsibility measure.
Model Variants and Language Coverage
NeuTTS comes in three backbone families:
- NeuTTS-Air (~360M active params): English only, Apache 2.0 license, supports voice cloning, GGUF streaming - NeuTTS-Nano multilingual collection (~120M active params each): English, French, German, and Spanish, each as a separate model, NeuTTS Open License 1.0, supports voice cloning, GGUF streaming - NeuTTS-2E (~125M active params): English with emotional control (6 emotions plus neutral), 4 fixed speakers, NeuTTS Open License 1.0
For the codec, any backbone can be paired with any codec variant. Decoder-only ONNX codecs require pre-encoded references but ship with the repository for all bundled voices, including NeuTTS-2E speakers. This means ONNX-only inference is possible without a GPU-backed encoder at inference time.
GGUF quantizations are provided at Q4 and Q8 precision for all models. Streaming inference is only available through the GGUF path, not through the PyTorch models.
Installing NeuTTS and Running Examples
The package is named `neutts` (version 1.4.1 per pyproject.toml) and is built using scikit-build-core. It requires Python 3.10 to 3.13. The core dependencies are torch >= 2.8.0, torchaudio >= 2.8.0, transformers == 5.1.0, phonemizer >= 3.0.0, librosa == 0.11.0, soundfile == 0.13.1, and neucodec >= 0.0.4.
Three optional dependency groups are defined in pyproject.toml: `llama` adds llama-cpp-python and gguf for GGUF inference via llama.cpp; `onnx` adds onnxruntime for ONNX codec decoding; `all` includes both. The pyproject.toml entry looks like:
[project.optional-dependencies]
llama = ["llama-cpp-python", "gguf"]
onnx = ["onnxruntime"]
all = ["llama-cpp-python", "gguf", "onnxruntime"]The `examples/` directory contains runnable scripts: `basic_example.py`, `basic_streaming_example.py`, `basic_example_emotions.py`, and `basic_streaming_example_emotions.py`. The Jupyter notebook `examples/interactive_example.ipynb` provides an interactive walkthrough. `examples/encode_reference.py` handles the voice cloning workflow for encoding a reference audio clip. A fine-tuning script and its configuration file are at `examples/finetune.py` and `examples/finetune_config.yaml`.
The TRAINING.md file at the repository root contains fine-tuning documentation. The repository also includes a `samples/` directory with reference audio files and an `output.wav` example output for verifying the installation.
Throughput on Real Hardware
The README includes CPU benchmarks for Q4_0 quantizations of NeuTTS-Air and NeuTTS-Nano across four devices. These numbers reflect the Speech Language Model only and do not include the time for the codec decode step, which runs separately.
For NeuTTS-Nano Q4_0 on CPU: 45 tokens per second on a Galaxy A25 5G (6 threads), 221 tokens per second on an AMD Ryzen 9HX 370 (14 threads prefill, 16 threads decode), and 195 tokens per second on an iMac M4 16GB (14 threads). On a GPU (RTX 4090 via vLLM), NeuTTS-Nano reaches 19,268 tokens per second.
For NeuTTS-Air Q4_0 on CPU: 20 tokens per second on the Galaxy A25 5G and 119 tokens per second on the Ryzen 9HX 370. The Air model is approximately 3x slower than Nano on CPU for the language model step, but it has approximately 3x more active parameters.
These benchmarks give a basis for hardware selection, but add codec latency to any real-world estimate.
Limitations and When to Use Something Else
The primary limitation of the NeuTTS Open License 1.0, which governs NeuTTS-Nano and NeuTTS-2E, is that its terms are not the same as Apache 2.0 or MIT. Teams planning commercial deployment should read that license carefully before proceeding. Only NeuTTS-Air carries the Apache 2.0 license.
The 2048-token context window imposes a ceiling of approximately 30 seconds of audio per generation call. For longer text passages, you must split the input and stitch the outputs, which may produce inconsistency at join points.
Air and Nano take phoneme input, not raw text. The `phonemizer` dependency requires espeak-ng on the host system. This is an additional system-level dependency that can complicate containerized deployments.
ElevenLabs is a commercial cloud TTS API with a broader set of voices and managed infrastructure. It requires sending audio data and text to their servers. Piper is an open-source on-device TTS system from the Rhasspy project with MIT-licensed models. Piper models are smaller and faster on very low-power hardware, but the README does not provide a direct quality comparison.
Repository Maintenance and License Notes
The repository is not archived and its last push was on 2026-07-30. The package version is 1.4.1 per `pyproject.toml`. There are no GitHub releases; versioning is managed through PyPI.
The license field in the repository metadata is listed as NOASSERTION. The actual licensing is model-specific: NeuTTS-Air uses Apache 2.0 and the Nano and 2E models use NeuTTS Open License 1.0. The README explicitly warns that websites such as neutts.com are not affiliated with Neuphonic; the official site is neuphonic.com.
Editorial conclusion
NeuTTS is a practical choice for engineers building embedded voice agents, privacy-sensitive applications, or offline assistants who cannot route audio through a cloud API. The GGUF models run on mid-range consumer hardware and even on a Galaxy A25 5G. The right variant depends on your constraints: NeuTTS-Air (Apache 2.0) is the only model with an unambiguously permissive license; NeuTTS-Nano and NeuTTS-2E carry the NeuTTS Open License 1.0, which you should read before deploying commercially. NeuTTS-2E supports text input directly, while Air and Nano require phoneme preprocessing. Verify your phonemizer dependencies (`phonemizer>=3.0.0`) work in your target environment before committing to those models.
Frequently asked questions
Is NeuTTS good for production use?
NeuTTS runs in real time on mid-range consumer hardware and mobile devices, supports voice cloning from three seconds of audio, and includes GGUF quantizations for efficient on-device inference. For production use, note that NeuTTS-Nano and NeuTTS-2E are under the NeuTTS Open License 1.0 rather than a permissive open-source license, and the benchmark numbers exclude codec latency.
NeuTTS vs Piper: how do they compare?
Piper is an open-source MIT-licensed on-device TTS system from the Rhasspy project, optimized for very low-power embedded hardware. NeuTTS is built on LLM backbones and supports instant voice cloning and multilingual models. The README does not benchmark NeuTTS against Piper directly. NeuTTS-Air uses Apache 2.0 but NeuTTS-Nano and 2E use the NeuTTS Open License 1.0.
NeuTTS vs ElevenLabs: what is the difference?
ElevenLabs is a commercial cloud TTS API that processes audio server-side. NeuTTS runs entirely on-device with no external API calls, which suits privacy-sensitive or offline applications. ElevenLabs offers a broader voice library and managed infrastructure. NeuTTS voice cloning requires only three seconds of reference audio and runs locally.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/neuphonic-neutts)