Open-source project
netease-youdao/EmotiVoice avatar
netease-youdao/EmotiVoice

EmotiVoice: prompt-controlled TTS with emotion, 2000 voices and a Docker path

EmotiVoice 😊: a Multi-Voice and Prompt-Controlled TTS Engine

8,535 stars758 forksPythonApache-2.0

At a glance

What is it?
EmotiVoice is a Python text-to-speech engine from NetEase Youdao that takes a speaker id and a style or emotion prompt alongside the text. It installs from pip, a Docker image or a Mac build, and it expects an NVIDIA GPU for the container route.
Who is it for?
Adopt EmotiVoice if you need English or Chinese speech with per-utterance emotion control and you can supply an NVIDIA GPU or a Mac. Skip it if you need Japanese or Korean, a hosted service with an uptime commitment, or a project whose last release is recent: v0.3 shipped on 2023-12-28 even though the repository was pushed on 2026-09-03.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 27 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem EmotiVoice targets: one voice, one flat tone

Most open text-to-speech stacks give you a voice and a speed slider. If you want the same sentence read as happy in one take and angry in the next, you either record it yourself or you train a separate model per emotion. EmotiVoice folds that choice into the input string. The README describes the synthesis text format as `<speaker>|<style_prompt/emotion_prompt/content>|<phoneme>|<content>`, so a single inference call carries the speaker identity, the emotional style and the words. The project says it speaks English and Chinese and ships over 2000 voices, with a wiki page listing them. That combination is the pitch: a fixed catalogue of speakers, plus a prompt dimension you change per request rather than per deployment. It is aimed at people building dubbing, game dialogue, audiobook drafts or accessibility tools who want variation without assembling a training pipeline for every mood. The audience is Python developers with a GPU, not end users looking for a web app they can sign into.

How the joint acoustic model and vocoder pipeline fits together

The repository layout tells you most of the architecture. Two inference entry points exist: `inference_am_vocoder_joint.py` and `inference_am_vocoder_exp.py`, which the names present as a joint acoustic-model-plus-vocoder run. The joint script takes `--logdir prompt_tts_open_source_joint`, `--config_folder config/joint` and a `--checkpoint g_00140000` argument, and writes audio into `outputs/prompt_tts_open_source_joint/test_audio`. Text does not go straight in. A frontend step converts it to phonemes: `frontend.py` reads a text file and prints a phoneme-annotated version, and the README suggests `frontend_en.py` for English and `frontend_cn.py` for Chinese, with `frontend_en.py` also invoked inside the Dockerfile at build time. A separate `predict.py` and a `text/` directory sit alongside, and `prepare_for_training.py` plus `train_am_vocoder_joint.py` cover the fine-tuning path. The style prompt is a free-text string such as `Happy`, placed between the speaker id and the phoneme block. That is a prompt-conditioned model, not a classifier with a fixed emotion enum, which is why the README can list happy, excited, sad and angry as examples rather than as the complete set. The trade-off is that a free-text prompt gives you no guarantee of a stable mapping: the same word may drift across checkpoints, and nothing in the repository documents a canonical emotion vocabulary.

Installing EmotiVoice: Docker, conda and the Mac build

The README calls the Docker image the easiest way to try it, and it requires a machine with an NVIDIA GPU plus the NVIDIA container toolkit. The image is published as `syq163/emoti-voice:latest`. The single-port command exposes only the Streamlit demo; the second form adds port 8000 for the OpenAI-compatible API, and the README states that image was updated on January 4th, 2024.

sh
docker pull syq163/emoti-voice:latest
docker run -dp 127.0.0.1:8501:8501 -p 127.0.0.1:8000:8000 syq163/emoti-voice:latest

After that, the README says to open http://localhost:8501 for the web interface and that the OpenAI-compatible TTS API answers on http://localhost:8000/. If you prefer a native install, the README gives a conda environment pinned to Python 3.8, followed by the pip set and an NLTK data download.

sh
conda create -n EmotiVoice python=3.8 -y
conda activate EmotiVoice
pip install torch torchaudio
pip install numpy numba scipy transformers soundfile yacs g2p_en jieba pypinyin pypinyin_dict
python -m nltk.downloader "averaged_perceptron_tagger_eng"

Model files are separate. The README points at the wiki page "How to download the pretrained model files" and shows two routes for the Chinese SimBERT component, a git-lfs clone from Hugging Face or a plain clone from ModelScope.

sh
git lfs install
git lfs clone https://huggingface.co/WangZeJun/simbert-base-chinese WangZeJun/simbert-base-chinese

For the synthesis checkpoints themselves the README gives `git clone https://www.modelscope.cn/syq163/outputs.git`, and it also links a Google Drive folder. There is a Mac app: the release notes for v0.3 on December 28th, 2023 include a `emotivoice-1.0.0-arm64.dmg` download, which is the only non-GPU route the README advertises.

A first real synthesis run, from text file to audio

The workflow is three steps: phonemize, format, synthesize. Put your text in a file, then run the frontend to get the phoneme form, which the README shows as a redirect into a second file.

sh
python frontend.py data/my_text.txt > data/my_text_for_tts.txt

Each line of the resulting file must follow the pipe-delimited format. The README's English example uses speaker `8051`, the style prompt `Happy`, a phoneme sequence with `<sos/eos>` markers and `engsp4` separators, and the content string `Emoti-Voice - a Multi-Voice and Prompt-Controlled T-T-S Engine`. Then point the joint inference script at that file.

sh
TEXT=data/inference/text
python inference_am_vocoder_joint.py \
--logdir prompt_tts_open_source_joint \
--config_folder config/joint \
--checkpoint g_00140000 \
--test_file $TEXT

The README states the output lands in `outputs/prompt_tts_open_source_joint/test_audio`. If you would rather click than script, `pip install streamlit` followed by `streamlit run demo_page.py` opens the interactive page, and the OpenAI-compatible API is a third entry point: install the extra packages and start uvicorn.

sh
pip install fastapi pydub uvicorn[standard] pyrubberband
uvicorn openaiapi:app --reload

The release notes credit community pull requests for voice speed tuning in that API and for the API itself, which is worth knowing: parts of the serving layer came from outside the core team.

Where EmotiVoice is the wrong tool

Two limitations are stated by the project itself. Japanese and Korean are listed under "Features under development" with issue references, so they are not available today. Everything else follows from the deployment shape. The Docker path requires an NVIDIA GPU; the README does not document a CPU-only container, and the Dockerfile installs `torch==1.11.0`, which is older than the `torch>=2.1` floor in setup.py's infer extras, so the container and the pip path do not share a dependency set. The Mac app is arm64 and tied to the v0.3 release from December 2023. If you need a hosted endpoint with a service level agreement, this is the wrong layer: the demo is hosted on Replicate, and the README mentions an HTTP API with over 13,000 free calls plus additional voices from Zhiyun, which are commercial services rather than guarantees about this repository. Emotion control is also not a solved interface. Because the style is a free-text prompt, reproducibility across runs depends on the checkpoint and on the prompt wording, and the README does not document a fixed emotion list or a way to verify that `Happy` at one checkpoint means the same thing at another. Finally, the last tagged release is v0.3 from 2023-12-28. The repository was pushed on 2026-09-03, so work continues, but anyone pinning to a release should expect the release itself to be old.

EmotiVoice against ChatTTS, GPT-SoVITS, IndexTTS2 and ElevenLabs

The alternatives people search for alongside EmotiVoice differ in what they optimise. ChatTTS and GPT-SoVITS are conversational and cloning-oriented respectively; GPT-SoVITS pairs a paper with a voice-cloning recipe, and the EmotiVoice README points to its own wiki page "Voice Cloning with your personal data" plus DataBaker and LJSpeech recipes under `data/`, so both projects expect you to bring audio if you want a new voice. IndexTTS2 and Qwen3-TTS appear in the same search space as newer entrants, and the EmotiVoice README does not compare itself to any of them. ElevenLabs is the hosted option: you send text and get audio, with no GPU, no checkpoint download and no conda environment, in exchange for per-character cost and an external dependency. EmotiVoice's actual difference is the prompt slot. A style string sits inside the inference line, so emotional variation is a per-request parameter rather than a separate model or a separate account tier. If your requirement is a single stable narrator voice at the lowest possible operational cost, the hosted route removes an entire class of problems that this project hands you: CUDA versions, checkpoint downloads and a Python 3.8 environment.

Licence, maintenance and what an upgrade costs

The repository is Apache-2.0, and setup.py declares `license="Apache Software License"`. That is permissive for the code. The catch is the file `EmotiVoice_UserAgreement_易魔声用户协议.pdf` in the repository root. A user agreement sitting next to an Apache-2.0 licence is a signal that the project intends additional terms around use, and the README does not explain how the two interact. I am not a lawyer and this is not legal advice: read that PDF before shipping anything commercial, and note that the README's own voice list and pretrained checkpoints may carry terms of their own. On maintenance, the facts are narrow. The last push was 2026-09-03. The last tagged release is v0.3 from 2023-12-28, preceded by v0.2 on 2023-11-17. So the release cadence is slow even though the branch moves. Upgrade cost concentrates in dependencies: setup.py pins `transformers==4.26.1` and `numba==0.58.1`, the Dockerfile pins `torch==1.11.0`, and the requirements.txt at the root lists the same packages unpinned. Three dependency sets that disagree means a container upgrade and a pip upgrade are separate projects. The README does not document a rollback procedure or a migration path between checkpoints, so treat a checkpoint change as a re-validation of your audio, not a routine bump.

Editorial conclusion

Adopt EmotiVoice if you need English or Chinese speech with per-utterance emotion control and you can supply an NVIDIA GPU or a Mac. Skip it if you need Japanese or Korean, a hosted service with an uptime commitment, or a project whose last release is recent: v0.3 shipped on 2023-12-28 even though the repository was pushed on 2026-09-03. Before committing, run the Docker image on your own hardware, confirm the pretrained checkpoints download from ModelScope or Google Drive, and read the user agreement PDF in the repository root, since the Apache-2.0 LICENSE sits next to a separate EmotiVoice user agreement.

Frequently asked questions

Is voice cloning illegal?

The repository does not make a legal claim about voice cloning. It ships an Apache-2.0 LICENSE alongside a separate file, EmotiVoice_UserAgreement_易魔声用户协议.pdf, in the repository root, and it links a wiki page titled "Voice Cloning with your personal data". Read both before cloning a voice you do not own.

What is an emotive voice?

In EmotiVoice the term maps to the style slot in the inference format `<speaker>|<style_prompt/emotion_prompt/content>|<phoneme>|<content>`. The README names happy, excited, sad and angry as examples of what that prompt can carry, and calls emotional synthesis the project's most prominent feature.

How does voice-based emotion detection work?

EmotiVoice does not detect emotion from audio. The emotion is supplied by you as a text prompt in the inference line, and the README's example uses `Happy` for speaker `8051`. Detection is a different task from what this repository implements.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. netease-youdao/EmotiVoice on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/netease-youdao-emotivoice.svg)](https://hysenlabs.com/projects/netease-youdao-emotivoice)