# VieNeu-TTS: on-device Vietnamese text to speech with instant voice cloning

> VieNeu-TTS is an Apache-2.0 Vietnamese TTS toolkit that runs on CPU through ONNX Runtime and clones a voice from a 3 to 8 second clip. Its v3 Turbo model is documented as early access, which matters more than the feature list.

**pnnbao97/VieNeu-TTS** — Vietnamese TTS with instant voice cloning • On-device • Real-time CPU inference • 24kHz audio quality • Chuyển văn bản thành giọng nói tiếng Việt • Text to speech tiếng Việt • TTS tiếng Việt

- Repository: https://github.com/pnnbao97/VieNeu-TTS
- Website: https://www.vieneu.io
- Stars: 2,673 · Forks: 818
- Language: Python
- License: Apache-2.0
- Published: 2026-08-08 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/pnnbao97-vieneu-tts

## The gap VieNeu-TTS fills for Vietnamese speech

Most open text to speech stacks treat Vietnamese as an afterthought. The voices are trained on English or Mandarin corpora, and Vietnamese input either comes out with wrong tones or is romanized badly. VieNeu-TTS is built the other way around: the project describes itself as Vietnamese TTS with bilingual English-Vietnamese training, and the pronunciation layer is a separate project by the same author, sea-g2p, which the README credits for handling code-switching between the two languages.

The intended reader is a developer or researcher who wants speech output for Vietnamese text without sending audio or text to a hosted API. The README's own framing is on-device: the minimal install is torch-free, so no PyTorch dependency is pulled in, and inference runs on ONNX Runtime on the CPU. That is a real constraint change, not a marketing line. It means a laptop without a discrete GPU can generate speech, and it means the model weights have to fit in memory alongside whatever else the process is doing.

The secondary audience is people building dialogue or podcast audio. There is a documented Podcast and Conversation mode with multi-speaker support and automatic character detection, and the v3 Turbo notes describe a batched Conversation mode that groups the whole script by speaker rather than by line.

## How the synthesis path is assembled

VieNeu-TTS is not one model. The README describes a stack: a text front end handled by sea-g2p for phonemization and English-Vietnamese code-switching, a codec, and a backbone that produces the waveform. For v3 Turbo the README names the codec explicitly as MOSS-Audio-Tokenizer-Nano and states the architecture was designed and trained from scratch by Phạm Nguyễn Ngọc Bảo.

The runtime has two engines behind one API. On CPU the SDK runs ONNX Runtime and never imports PyTorch. On a CUDA machine the SDK auto-detects and switches to a PyTorch engine, and the README states that inference is batched automatically on CUDA without a code change. That auto-detection is convenient, but it also means the same script can take a different execution path on two machines, which is worth knowing before you compare timings across a team.

Precision is a CPU-side knob. The README states that on CPU the backbone runs int8 by default, described as roughly 1.6 times faster and about 4 times smaller than fp32, and that `Vieneu(precision="fp32")` trades speed for fidelity. The same paragraph notes that `precision` is ignored on GPU because that path is PyTorch. Model selection is a constructor argument: the SDK defaults to v3 Turbo, and `Vieneu(mode="v3turbo")` selects it explicitly, with v2 and v3 variants published as separate Hugging Face repositories.

## Installing VieNeu-TTS and generating a first clip

The README recommends uv for dependency management and gives separate install paths for CPU and GPU. The CPU path is the default and the one the project steers most users toward. It installs only the ONNX stack and never installs PyTorch.

```bash
git clone https://github.com/pnnbao97/VieNeu-TTS.git
cd VieNeu-TTS
uv sync
```

After `uv sync` finishes, start the bundled Gradio interface. The README says the UI is served at `http://127.0.0.1:7860`.

```bash
uv run vieneu-web
```

The README notes that `uv sync` reproduces a locked environment pinning an optimized ONNX Runtime build, and explicitly says to prefer it over `pip install` for the fastest CPU inference. That is a stronger claim than most projects make about their own lockfile, and it is the reason the Makefile exposes a matching `make setup` target described as a torch-free install of v3 Turbo on ONNX plus v3 Nano on CPU.

For programmatic use, the SDK is a separate install with a small surface. The README's quick start constructs the object with no arguments, which selects v3 Turbo at 48 kHz, and calls `infer` with a Vietnamese string that includes an inline emotion cue.

```python
from vieneu import Vieneu

vieneu = Vieneu()
audio = vieneu.infer("[cười] Trời ơi, cái giọng nó tự nhiên mà nó mượt mà dã man")
```

That first call downloads model weights, so expect a delay before any audio appears. The README also documents a streaming path for v3 Turbo that starts audio in roughly 300 ms and keeps generation ahead of playback with an RTF below 1.

## Where the CPU-first design stops paying off

The README is unusually candid that GPU is not automatically the faster choice. It states that the GPU win comes from batching, so it only pays off on long text, and that for short text the torch-free CPU path is usually faster because there is no batch to fill and no kernel-launch overhead. Read that as a design boundary: if your workload is one sentence at a time in a request handler, adding a CUDA GPU buys you very little and adds an install path.

The other boundary is maturity. The package metadata classifies the project as Development Status 4 - Beta, and the README labels v3 Turbo as early access and a preview, with the full v3 release described as coming in the next few weeks. The SDK nonetheless defaults to that preview model. A default that points at a preview is a deliberate choice, and it means the behaviour you get from `Vieneu()` may change when the full v3 lands. If you need reproducible output, pin the model variant rather than relying on the default.

There is also a deprecation to track. The v3 Turbo notes state that the `style` argument is deprecated and ignored, and that reading style now follows the reference voice instead. Code written against v2 that passes `style` will still run, but it will silently do nothing on v3 Turbo. The README does not document a rollback path between model versions, and it does not describe how to pin a specific weight revision beyond naming the Hugging Face repository.

## VieNeu-TTS compared with F5-TTS

F5-TTS is the natural comparison because both do zero-shot voice cloning from a short reference clip, and F5-TTS appears often enough in searches around this project to be the reference point people already have in mind. The difference is in what the two projects optimize for.

F5-TTS is a general multilingual flow-matching model. You bring your own reference audio and it clones whatever language the reference is in, which makes it flexible but leaves Vietnamese-specific problems, notably tone accuracy and English words embedded in Vietnamese sentences, to the model's general training. VieNeu-TTS narrows the scope. It ships a phonemizer built for the language pair, publishes Vietnamese-specific model variants, and offers built-in default voices on v3 Turbo so a reference clip is optional. The README states that built-in voices are stable and consistent and need no reference clip, which is a workflow F5-TTS does not offer in the same form.

The trade-off runs the other way too. F5-TTS is not Vietnamese-only, so if you need a second or third language, one model covers more ground. VieNeu-TTS also carries a heavier dependency story for the GPU path: a pinned CUDA wheel index, a pinned `transformers==4.57.6`, and an explicit note that this pin is the most stable version for the GPU SDK. F5-TTS users are not managing a project-specific pin like that.

## Licence, packaging and the cost of keeping up

The repository is Apache-2.0, and `pyproject.toml` declares the licence as a file reference to LICENSE. Apache-2.0 is permissive, so commercial use and modification are permitted under its terms. Two things sit outside that grant and are worth checking yourself: the v3 Turbo codec is a separate model published by another team, and the weights live on Hugging Face under their own model cards. The repository licence covers the code, not necessarily every artifact the code downloads. This is a factual boundary in the project layout, not legal advice.

The upgrade cost is real and visible in the release cadence. The desktop app for Windows shipped app-v0.7.8, app-v0.7.9 and app-v0.8.0 between 2026-08-20 and 2026-08-25, and the v0.7.8 release note states it was replaced by 0.7.9. Releases that supersede each other within days mean anyone tracking the app should follow the release feed rather than assume a downloaded build is current. The Python package is versioned separately at 3.6.4 in `pyproject.toml`.

Installation surface is broader than the README's happy path suggests. The repository root carries `requirements_xpu.txt`, `run_xpu.bat` and `setup_xpu_uv.bat` for Intel XPU, a `docker/` directory with CPU and GPU compose profiles exposed through `make docker-cpu` and `make docker-gpu`, a `finetune/` directory with a LoRA path described as requiring peft and accelerate, and a `.env.example` that sets `GRADIO_SERVER_PORT=7860` and points `PHONEMIZER_ESPEAK_LIBRARY` at a system eSpeak NG shared library. That last variable is the one most likely to break a container build, because it hardcodes a Linux library path. The last push to the default branch was on 2026-08-25.

## Conclusion

Adopt VieNeu-TTS if you need Vietnamese speech synthesis that stays on your own machine and you can accept a beta-stage package with a preview model. Do not adopt it if you need a stable, versioned synthesis API for production pipelines, or if your target language is not Vietnamese or English. Before committing, run `uv sync` and `uv run vieneu-web`, confirm that v3 Turbo quality on your own text is acceptable, and read the LICENSE file plus the model cards on Hugging Face, since the code and the weights are licensed separately.

## FAQ

### Is voice cloning illegal?

The repository does not address the legality of cloning a voice. It documents how to clone from a 3 to 8 second reference clip, and the code is Apache-2.0, which says nothing about consent or personality rights in your jurisdiction.

### What is TTS and how does it work?

Text to speech converts written text into an audio waveform. In VieNeu-TTS the README describes a pipeline where sea-g2p handles phonemization and English-Vietnamese code-switching, and a backbone generates the 24 kHz or 48 kHz waveform on ONNX Runtime or PyTorch.

### Is Google TTS free?

The repository does not discuss Google TTS or its pricing. What it does state is that VieNeu-TTS runs fully offline on your own machine, which removes per-request API cost and keeps text and reference audio local.

## Sources

- [Official documentation](https://www.vieneu.io)
- [Official README](https://github.com/pnnbao97/VieNeu-TTS#readme)
- [Project repository](https://github.com/pnnbao97/VieNeu-TTS)
- [Release notes](https://github.com/pnnbao97/VieNeu-TTS/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/pnnbao97-vieneu-tts
