Model or dataset
netease-youdao/Confucius4-TTS avatar
netease-youdao/Confucius4-TTS

Confucius4-TTS: Zero-Shot Voice Cloning Across 14 Languages

Confucius4-TTS: a Multilingual and Cross-Lingual Zero-Shot TTS Engine

824 stars88 forksPythonNOASSERTION

At a glance

What is it?
Confucius4-TTS is an LLM-based text-to-speech engine from NetEase Youdao that clones a speaker's voice across 14 languages from a single reference audio clip, with no reference transcript required.
Who is it for?
Confucius4-TTS is appropriate for researchers and developers building multilingual voice applications where preserving speaker identity across languages is a requirement. The CUDA 12.6 dependency and the Python 3.10 floor narrow the hardware pool significantly; users on older drivers or consumer GPUs without 12.6-compatible drivers should not attempt the install without checking compatibility first.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 27 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Confucius4-TTS Solves in Multilingual Voice Synthesis

Most traditional text-to-speech systems are monolingual or require separate voice models for each language. Cross-lingual transfer, making the same voice speak a different language without a foreign accent, is a distinct and harder problem. Confucius4-TTS targets this problem directly.

The engine takes a reference audio clip and a target text, then generates speech in any of its 14 supported languages using the voice characteristics from the reference. The 14 languages are Chinese, English, Japanese, Korean, German, French, Spanish, Indonesian, Italian, Thai, Portuguese, Russian, Malay, and Vietnamese, with the README noting that more are coming.

The key design claim is unconstrained voice cloning: no transcript of the reference audio is required. The speaker's voice characteristics are extracted from the audio signal alone. The README also describes emotion transfer, where the emotional coloring of the reference audio carries over into the synthesized output.

The primary audience is developers building multilingual content production pipelines, dubbing tools, or voice-interface applications where a consistent speaker identity must span language boundaries.

Architecture: Speech Encoder and LLM in Two Stages

Confucius4-TTS uses a two-stage pipeline described in the README as a speech encoder combined with a large language model. The first stage is the Text2Semantic (T2S) autoregressive stage, which maps text and the extracted speaker representation to a discrete semantic token sequence. The second stage is an acoustic decoder that converts those tokens to the final waveform.

The T2S stage runs as a HuggingFace Transformers model by default. The repository includes an alternative vLLM backend for this stage that uses PagedAttention to accelerate the autoregressive generation. Switching to vLLM affects only the T2S computation; the acoustic decoder is unchanged.

The repository layout reflects this architecture: configuration lives in config/, the model implementation is in confuciustts/, external components (including third-party code) are in external/, and inference entry points are example.py, example_vllm.py, webui.py, and server.py. The checkpoints directory and the HuggingFace repository (netease-youdao/Confucius4-TTS) hold the model weights.

Installing and Running a First Synthesis

The README states Python 3.10 and CUDA 12.6 as requirements. Installation starts with cloning the repository and creating a dedicated conda environment:

bash
git clone https://github.com/netease-youdao/Confucius4-TTS.git
cd Confucius4-TTS
bash
conda create -n confuciustts python=3.10 -y
conda activate confuciustts

Then install the Python dependencies:

bash
pip install -r requirements.txt

For HuggingFace-restricted environments, the README recommends setting a mirror endpoint before running:

bash
export HF_ENDPOINT=https://hf-mirror.com

Once installed, zero-shot synthesis using example.py takes a reference audio file, a text string, a language code, and an output path:

bash
python example.py \
    --prompt_wav path/to/reference.wav \
    --text "Hello, this is a test of zero-shot voice cloning." \
    --lang en \
    --out output.wav \
    --config config/inference_config.yaml

The Python API is also available through the ConfuciusTTS class in confuciustts.cli.inference, accepting the same parameters and returning a torchaudio-compatible tensor. The README gives a complete code example for this path, including saving the output with torchaudio.

The vLLM Backend: Faster T2S at a Narrow Version Constraint

For production workloads where the autoregressive T2S stage is the throughput bottleneck, Confucius4-TTS supports a vLLM backend that accelerates the LLM with PagedAttention. The README makes the version requirement explicit: vLLM 0.16.0 using the V1 engine. The README states that older vLLM versions have not been tested and may not work because the model registration and GPUModelRunner patches target V1 engine internals.

The recommended setup is a separate conda environment cloned from the base installation:

bash
conda create -n confuciustts_vllm --clone confuciustts
conda activate confuciustts_vllm
pip install -r requirements_vllm_add.txt

The README warns that vLLM has specific requirements on GPU architecture, driver version, and CUDA, and that installing it into the base environment can break other packages. Both non-streaming (return a complete WAV) and streaming (generate chunk by chunk) modes are available via the `--stream` flag in example_vllm.py.

This vLLM dependency is a real maintenance risk. The vLLM project releases frequently and V1 internals can change between minor versions. Any vLLM upgrade that modifies GPUModelRunner would silently break the integration until the patches in Confucius4-TTS are updated.

Web UI, FastAPI Server, and Online Demo

For interactive testing without writing code, Confucius4-TTS provides a Gradio web interface:

bash
python webui.py --port 7860

The README states that the reference audio is uploaded to the server via HTTP, so the browser client does not need to share a filesystem with the server. This makes the web UI usable for remote inference on a GPU server.

For programmatic access, a FastAPI server exposes two endpoints:

bash
python server.py --port 8000

The `/api/tts` endpoint (POST, multipart) returns a complete WAV in PCM 16-bit format. The `/api/tts/stream` endpoint (POST, multipart) returns raw int16-LE PCM chunks with the sample rate in the `X-Sample-Rate` response header. Both endpoints accept the reference audio as an HTTP file upload, so client and server do not need to share a filesystem.

An online demo is available at confucius4-tts.youdao.com/gradio for evaluating synthesis quality without any local setup.

Limitations: Hardware Floor, Version Pinning, and Licence Ambiguity

The CUDA 12.6 requirement is stricter than many other Python deep learning projects, which typically allow CUDA 11.8 or 12.1. Users with older GPU drivers or systems that cannot be upgraded to a 12.6-compatible environment cannot run the model.

The PyTorch version is pinned at exactly 2.7.0 in requirements.txt. This makes the environment more reproducible but creates friction when the user's other projects require a different PyTorch version. The vLLM path compounds this by requiring its own isolated environment.

The licence field in the repository metadata shows NOASSERTION, which typically means the automated licence detector could not identify the licence. The repository does include a LICENSE file. Anyone considering commercial deployment should read it directly before proceeding.

A practical alternative for zero-shot multilingual TTS is Coqui TTS (XTTSv2 specifically), which also supports multiple languages and zero-shot cloning from a reference audio clip. XTTSv2 runs on CUDA 11.8 and later, which is a lower hardware bar. The trade-off is that Confucius4-TTS covers 14 languages with a specific emphasis on Asian languages such as Chinese, Japanese, Korean, Thai, and Malay that are not in XTTSv2's coverage.

Editorial conclusion

Confucius4-TTS is appropriate for researchers and developers building multilingual voice applications where preserving speaker identity across languages is a requirement. The CUDA 12.6 dependency and the Python 3.10 floor narrow the hardware pool significantly; users on older drivers or consumer GPUs without 12.6-compatible drivers should not attempt the install without checking compatibility first. The licence is listed as NOASSERTION in the repository metadata, and the included LICENSE file should be reviewed carefully before any commercial or redistributed use. The online Gradio demo at confucius4-tts.youdao.com/gradio allows functional testing without a local GPU.

Frequently asked questions

What languages does Confucius4-TTS support?

The README lists 14 languages: Chinese, English, Japanese, Korean, German, French, Spanish, Indonesian, Italian, Thai, Portuguese, Russian, Malay, and Vietnamese, with more noted as coming soon.

Does Confucius4-TTS require a transcript of the reference audio?

No. The README describes the voice cloning as unconstrained, meaning no reference transcript is required. The model extracts speaker characteristics from the audio signal alone.

What GPU and Python version does Confucius4-TTS require?

The README states Python 3.10 and CUDA 12.6 are required. The requirements.txt pins PyTorch at version 2.7.0.

Official sources

  1. Issues
  2. netease-youdao/Confucius4-TTS on GitHub
  3. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/netease-youdao-confucius4-tts.svg)](https://hysenlabs.com/projects/netease-youdao-confucius4-tts)