MOSS-TTS-Nano: Real-Time Multilingual TTS on CPU
A 100M-parameter multilingual TTS model for real-time CPU inference, voice cloning, and 48 kHz stereo generation
At a glance
- What is it?
- MOSS-TTS-Nano is an open-source 0.1B parameter text-to-speech model from OpenMOSS and MOSI.AI, designed to run on CPU hardware without a GPU. It supports 20 languages, produces 48kHz stereo output, and provides voice cloning through a pure autoregressive Audio Tokenizer plus LLM pipeline. An ONNX version offers roughly twice the processing efficiency of the PyTorch version.
- Who is it for?
- MOSS-TTS-Nano is a practical choice for teams that need a self-hosted TTS system with multilingual support and voice cloning, without requiring a GPU. The ONNX version removes the PyTorch dependency from the inference path, which reduces the deployment footprint and runs on a single CPU core.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 24 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What MOSS-TTS-Nano is and who it is designed for
MOSS-TTS-Nano is a tiny open-source text-to-speech model released by MOSI.AI and the OpenMOSS team. At 0.1B parameters it is deliberately small: the README frames the design goal as producing speech quality that is good enough for real-time product use while keeping the deployment footprint minimal enough for CPU-only hardware.
The target users are developers who want local TTS without renting GPU cloud instances, teams building web demos or lightweight integrations that need to run on commodity hardware, and researchers who want a starting point for finetuning a multilingual model on custom voice data. The README notes that the model handles streaming inference with low first-audio latency and can generate audio on a 4-core CPU, which covers a wide range of server and desktop hardware configurations.
The project is Apache-2.0 licensed, with all model weights available on both Hugging Face and ModelScope. A live demo is accessible at the Hugging Face Space under OpenMOSS-Team/MOSS-TTS-Nano.
Architecture: Audio Tokenizer and LLM pipeline
The README describes MOSS-TTS-Nano as built on a pure autoregressive architecture: an audio tokenizer converts speech into discrete tokens, and a language model generates new token sequences conditioned on the input text and a reference audio clip. The generated tokens are then decoded back into audio waveforms by the tokenizer.
The companion component is MOSS-Audio-Tokenizer-Nano, which handles both the encoding and decoding steps. The README describes it as a standalone component with its own evaluation results, including reconstruction quality comparisons on speech, audio, and music benchmarks against other open-source tokenizers on the LibriSpeech dataset. The tokenizer is also published separately on Hugging Face.
This Audio Tokenizer plus LLM design is a departure from older cascade TTS pipelines (acoustic model plus vocoder) and from diffusion-based TTS systems. The advantage is that an autoregressive LLM-based TTS can benefit from the same inference optimizations developed for language models, including token streaming and batched generation. The tradeoff is that autoregressive generation is inherently sequential: each token depends on the previous one, which is why the streaming mode matters for keeping first-audio latency low.
Installation and running the PyTorch version
The package is installable via pip. The pyproject.toml declares the package name as moss-tts-nano with a minimum Python version of 3.10. The dependencies include pinned versions of the core inference stack:
dependencies = [
"numpy>=1.24",
"fastapi>=0.110.0",
"python-multipart>=0.0.9",
"sentencepiece>=0.1.99",
"torch==2.7.0",
"torchaudio==2.7.0",
"transformers==4.57.1",
"uvicorn>=0.29.0",
"onnxruntime>=1.20.0",
]For local setup from source, the README describes a conda environment approach under the quickstart section. After installing dependencies, the direct inference script is run with `python infer.py` for voice cloning from a reference audio file, and `python app.py` for a local Gradio web demo. These two scripts cover the two primary usage modes: scripted batch inference and interactive browser-based testing. The package also registers a command-line entry point through the pyproject.toml scripts section:
[project.scripts]
moss-tts-nano = "moss_tts_nano.cli:main"Installing the package via pip makes the `moss-tts-nano` command available in the shell.
CLI commands: generate and serve
The pyproject.toml entry point (`moss_tts_nano.cli:main`) exposes two CLI subcommands documented in the README quickstart: `moss-tts-nano generate` and `moss-tts-nano serve`. The generate subcommand handles text-to-speech generation with voice cloning from a reference audio. The serve subcommand starts a local HTTP server backed by FastAPI and uvicorn, which is the same server used by the web demo scripts.
The serve mode makes MOSS-TTS-Nano accessible as an HTTP endpoint, which is how it integrates into larger applications: a parent service can send text to the TTS server and receive audio in response without embedding the Python inference code directly. The fastapi and uvicorn dependencies in pyproject.toml are what back this server. The python-multipart dependency handles multipart form data, which is the standard format for uploading the reference audio file along with the text input in a single HTTP request.
For the ONNX version, a parallel set of scripts is provided: `infer_onnx.py` for direct scripted inference and `app_onnx.py` for the ONNX-backed web demo. These use the `ort_cpu_runtime.py` module, which wraps ONNX Runtime rather than PyTorch.
ONNX CPU version and the Reader extension
The ONNX version of MOSS-TTS-Nano was released in April 2026 and removes PyTorch from the inference path. The model weights are published as two separate Hugging Face repositories: MOSS-TTS-Nano-100M-ONNX for the TTS model and MOSS-Audio-Tokenizer-Nano-ONNX for the tokenizer. The README states that in testing, the ONNX version delivers nearly twice the processing efficiency of the PyTorch version, and it runs smoothly on a single CPU core on a MacBook Air M4.
The practical implication is that a deployment using the ONNX version can serve the same voice cloning quality on significantly cheaper hardware, and the absence of PyTorch as a runtime dependency simplifies the container image for teams that only need inference, not training.
Built on the ONNX version, the OpenMOSS team also released MOSS-TTS-Nano-Reader, a browser extension that runs MOSS-TTS-Nano locally without requiring a separate inference service. The extension approach is made possible by the ONNX version's ability to run in JavaScript environments through ONNX Runtime Web. This makes it practical to use in contexts where installing a local Python server is not an option.
For Android deployment, the repository includes an example in `examples/android_onnx_runtime/` that demonstrates running the ONNX model using the Android ONNX Runtime. The mlx-audio integration (added in May 2026) extends support to Apple Silicon environments running MLX.
Supported languages and audio output specifications
MOSS-TTS-Nano supports 20 languages. The README lists them with their ISO codes: Chinese (zh), English (en), German (de), Spanish (es), French (fr), Italian (it), Japanese (ja), Korean (ko), Portuguese (pt), Russian (ru), Arabic (ar), Hindi (hi), Turkish (tr), Dutch (nl), Polish (pl), Swedish (sv), Indonesian (id), Thai (th), Vietnamese (vi), and Czech (cs).
The audio output format is 48kHz with 2 channels (stereo). 48kHz stereo is a higher sample rate than many TTS systems target; the consequence is that each second of generated audio requires more data than 16kHz mono output, but the audio quality ceiling is correspondingly higher. The README does not document loudness normalization or specific codec output formats; the inference scripts produce WAV files directly.
Long-text input is handled through automatic chunked voice cloning: the model splits long input text into segments, processes each segment with the reference voice, and concatenates the outputs. This avoids the context length limitations of the autoregressive model architecture while preserving voice consistency across chunks. Voice cloning requires a reference audio clip from the target speaker; no training is needed to clone a new voice, only a reference sample.
Finetuning and the broader MOSS-TTS family
The finetuning code was released in April 2026. The README points to `finetuning/README.md` for training and usage details. Finetuning allows teams to adapt MOSS-TTS-Nano to a specific voice or language domain on their own data, rather than relying solely on zero-shot voice cloning from a reference audio clip. Finetuning is relevant for production deployments where voice consistency is critical and where a reference audio may not always be available.
The requirements.txt for the base package includes WeTextProcessing>=1.0.4.1 for text normalization, soundfile for audio I/O, and sentencepiece for tokenization. These cover the text preprocessing pipeline that converts raw input text into a form the model can accept.
The README describes a MOSS-TTS family that includes MOSS-TTS (the full-size model) and MOSS-TTS-Nano. The distinction the README draws is that Nano is optimized for low-latency CPU deployment, while the full MOSS-TTS model targets higher quality at the cost of more compute. The 2.0 version of MOSS-TTS was announced as coming in April 2026. MOSS-TTS-Nano also integrates with mlx-audio, which provides an alternative inference backend for Apple Silicon hardware that uses the MLX framework instead of PyTorch or ONNX Runtime.
Editorial conclusion
MOSS-TTS-Nano is a practical choice for teams that need a self-hosted TTS system with multilingual support and voice cloning, without requiring a GPU. The ONNX version removes the PyTorch dependency from the inference path, which reduces the deployment footprint and runs on a single CPU core. The last push was on 2026-09-06. The project is Apache-2.0 licensed, has a demo available at the Hugging Face Space, and the finetuning code and Android ONNX runtime example are included in the repository for teams that need custom voices or mobile deployment.
Frequently asked questions
What hardware is required to run MOSS-TTS-Nano?
The README states that streaming generation can run on a 4-core CPU without a GPU. The ONNX version runs on a single CPU core on a MacBook Air M4 in testing. The PyTorch version requires torch==2.7.0 and torchaudio==2.7.0, which need at least Python 3.10 per the pyproject.toml.
What is the difference between the PyTorch and ONNX versions of MOSS-TTS-Nano?
The PyTorch version uses torch==2.7.0 and torchaudio==2.7.0 as its runtime and is the primary development version. The ONNX version removes PyTorch from the inference path and uses ONNX Runtime instead, which the README states delivers nearly twice the processing efficiency. The ONNX version also enables the browser extension (MOSS-TTS-Nano-Reader) and the Android ONNX Runtime example.
Does MOSS-TTS-Nano support voice cloning without finetuning?
Yes. The README describes zero-shot voice cloning: passing a reference audio clip from the target speaker to `python infer.py` or the `moss-tts-nano generate` CLI command produces output in that voice without any training. For production use where a reference clip is not always available or where higher voice consistency is needed, the finetuning code in `finetuning/README.md` allows training on custom voice data.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/openmoss-moss-tts-nano)