Confucius4-TTS: Zero-Shot Voice Cloning Across 14 Languages
Confucius4-TTS: a Multilingual and Cross-Lingual Zero-Shot TTS Engine
At a glance
- What is it?
- Confucius4-TTS is an LLM-based text-to-speech engine from NetEase Youdao that clones a speaker's voice across 14 languages from a single reference audio clip, with no reference transcript required.
- Who is it for?
- Confucius4-TTS is appropriate for researchers and developers building multilingual voice applications where preserving speaker identity across languages is a requirement. The CUDA 12.6 dependency and the Python 3.10 floor narrow the hardware pool significantly; users on older drivers or consumer GPUs without 12.6-compatible drivers should not attempt the install without checking compatibility first.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 27 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Confucius4-TTS Solves in Multilingual Voice Synthesis
Most traditional text-to-speech systems are monolingual or require separate voice models for each language. Cross-lingual transfer, making the same voice speak a different language without a foreign accent, is a distinct and harder problem. Confucius4-TTS targets this problem directly.
The engine takes a reference audio clip and a target text, then generates speech in any of its 14 supported languages using the voice characteristics from the reference. The 14 languages are Chinese, English, Japanese, Korean, German, French, Spanish, Indonesian, Italian, Thai, Portuguese, Russian, Malay, and Vietnamese, with the README noting that more are coming.
The key design claim is unconstrained voice cloning: no transcript of the reference audio is required. The speaker's voice characteristics are extracted from the audio signal alone. The README also describes emotion transfer, where the emotional coloring of the reference audio carries over into the synthesized output.
The primary audience is developers building multilingual content production pipelines, dubbing tools, or voice-interface applications where a consistent speaker identity must span language boundaries.
Architecture: Speech Encoder and LLM in Two Stages
Confucius4-TTS uses a two-stage pipeline described in the README as a speech encoder combined with a large language model. The first stage is the Text2Semantic (T2S) autoregressive stage, which maps text and the extracted speaker representation to a discrete semantic token sequence. The second stage is an acoustic decoder that converts those tokens to the final waveform.
The T2S stage runs as a HuggingFace Transformers model by default. The repository includes an alternative vLLM backend for this stage that uses PagedAttention to accelerate the autoregressive generation. Switching to vLLM affects only the T2S computation; the acoustic decoder is unchanged.
The repository layout reflects this architecture: configuration lives in config/, the model implementation is in confuciustts/, external components (including third-party code) are in external/, and inference entry points are example.py, example_vllm.py, webui.py, and server.py. The checkpoints directory and the HuggingFace repository (netease-youdao/Confucius4-TTS) hold the model weights.
Installing and Running a First Synthesis
The README states Python 3.10 and CUDA 12.6 as requirements. Installation starts with cloning the repository and creating a dedicated conda environment:
git clone https://github.com/netease-youdao/Confucius4-TTS.git
cd Confucius4-TTSconda create -n confuciustts python=3.10 -y
conda activate confuciusttsThen install the Python dependencies:
pip install -r requirements.txtFor HuggingFace-restricted environments, the README recommends setting a mirror endpoint before running:
export HF_ENDPOINT=https://hf-mirror.comOnce installed, zero-shot synthesis using example.py takes a reference audio file, a text string, a language code, and an output path:
python example.py \
--prompt_wav path/to/reference.wav \
--text "Hello, this is a test of zero-shot voice cloning." \
--lang en \
--out output.wav \
--config config/inference_config.yamlThe Python API is also available through the ConfuciusTTS class in confuciustts.cli.inference, accepting the same parameters and returning a torchaudio-compatible tensor. The README gives a complete code example for this path, including saving the output with torchaudio.
The vLLM Backend: Faster T2S at a Narrow Version Constraint
For production workloads where the autoregressive T2S stage is the throughput bottleneck, Confucius4-TTS supports a vLLM backend that accelerates the LLM with PagedAttention. The README makes the version requirement explicit: vLLM 0.16.0 using the V1 engine. The README states that older vLLM versions have not been tested and may not work because the model registration and GPUModelRunner patches target V1 engine internals.
The recommended setup is a separate conda environment cloned from the base installation:
conda create -n confuciustts_vllm --clone confuciustts
conda activate confuciustts_vllm
pip install -r requirements_vllm_add.txtThe README warns that vLLM has specific requirements on GPU architecture, driver version, and CUDA, and that installing it into the base environment can break other packages. Both non-streaming (return a complete WAV) and streaming (generate chunk by chunk) modes are available via the `--stream` flag in example_vllm.py.
This vLLM dependency is a real maintenance risk. The vLLM project releases frequently and V1 internals can change between minor versions. Any vLLM upgrade that modifies GPUModelRunner would silently break the integration until the patches in Confucius4-TTS are updated.
Web UI, FastAPI Server, and Online Demo
For interactive testing without writing code, Confucius4-TTS provides a Gradio web interface:
python webui.py --port 7860The README states that the reference audio is uploaded to the server via HTTP, so the browser client does not need to share a filesystem with the server. This makes the web UI usable for remote inference on a GPU server.
For programmatic access, a FastAPI server exposes two endpoints:
python server.py --port 8000The `/api/tts` endpoint (POST, multipart) returns a complete WAV in PCM 16-bit format. The `/api/tts/stream` endpoint (POST, multipart) returns raw int16-LE PCM chunks with the sample rate in the `X-Sample-Rate` response header. Both endpoints accept the reference audio as an HTTP file upload, so client and server do not need to share a filesystem.
An online demo is available at confucius4-tts.youdao.com/gradio for evaluating synthesis quality without any local setup.
Limitations: Hardware Floor, Version Pinning, and Licence Ambiguity
The CUDA 12.6 requirement is stricter than many other Python deep learning projects, which typically allow CUDA 11.8 or 12.1. Users with older GPU drivers or systems that cannot be upgraded to a 12.6-compatible environment cannot run the model.
The PyTorch version is pinned at exactly 2.7.0 in requirements.txt. This makes the environment more reproducible but creates friction when the user's other projects require a different PyTorch version. The vLLM path compounds this by requiring its own isolated environment.
The licence field in the repository metadata shows NOASSERTION, which typically means the automated licence detector could not identify the licence. The repository does include a LICENSE file. Anyone considering commercial deployment should read it directly before proceeding.
A practical alternative for zero-shot multilingual TTS is Coqui TTS (XTTSv2 specifically), which also supports multiple languages and zero-shot cloning from a reference audio clip. XTTSv2 runs on CUDA 11.8 and later, which is a lower hardware bar. The trade-off is that Confucius4-TTS covers 14 languages with a specific emphasis on Asian languages such as Chinese, Japanese, Korean, Thai, and Malay that are not in XTTSv2's coverage.
Editorial conclusion
Confucius4-TTS is appropriate for researchers and developers building multilingual voice applications where preserving speaker identity across languages is a requirement. The CUDA 12.6 dependency and the Python 3.10 floor narrow the hardware pool significantly; users on older drivers or consumer GPUs without 12.6-compatible drivers should not attempt the install without checking compatibility first. The licence is listed as NOASSERTION in the repository metadata, and the included LICENSE file should be reviewed carefully before any commercial or redistributed use. The online Gradio demo at confucius4-tts.youdao.com/gradio allows functional testing without a local GPU.
Frequently asked questions
What languages does Confucius4-TTS support?
The README lists 14 languages: Chinese, English, Japanese, Korean, German, French, Spanish, Indonesian, Italian, Thai, Portuguese, Russian, Malay, and Vietnamese, with more noted as coming soon.
Does Confucius4-TTS require a transcript of the reference audio?
No. The README describes the voice cloning as unconstrained, meaning no reference transcript is required. The model extracts speaker characteristics from the audio signal alone.
What GPU and Python version does Confucius4-TTS require?
The README states Python 3.10 and CUDA 12.6 are required. The requirements.txt pins PyTorch at version 2.7.0.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/netease-youdao-confucius4-tts)