CosyVoice: an Apache-2.0 multilingual TTS stack for zero-shot voice cloning
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
At a glance
- What is it?
- CosyVoice 3.0 is a 0.5B text-to-speech model with zero-shot cloning across nine languages and 18+ Chinese dialects, shipped with training, inference and deployment scripts. The install is a conda recipe plus a model download, and the licence is Apache-2.0.
- Who is it for?
- Adopt CosyVoice if you need zero-shot cloning across Chinese, English and the dialects it lists, and you are willing to run a conda environment with pinned torch 2.3.1 and onnxruntime-gpu 1.18.0 on Linux. Do not adopt it if you need a hosted endpoint, a Windows training path, or a licence review that survives a strict legal read of the weights.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 127 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 26, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What CosyVoice solves, and who it is actually for
CosyVoice is a text-to-speech system built on a large language model backbone. The problem it targets is specific: producing speech in a voice that was never in the training set, from a short reference clip, across languages the reference speaker does not speak. The README frames Fun-CosyVoice 3.0 as designed for zero-shot multilingual speech synthesis in the wild, and lists 9 common languages (Chinese, English, Japanese, Korean, German, Spanish, French, Italian, Russian) plus 18+ Chinese dialects and accents, including Guangdong, Minnan, Sichuan, Dongbei, Shanghai and Tianjin.
The audience is narrower than the model card suggests. This is a Python project that expects a conda environment, a GPU with CUDA 12.1 wheels, and a model download before the first line of inference runs. Someone who wants a hosted API with a billing page is not the target. Someone building a Chinese-language product with dialect coverage, or a research group fine-tuning on their own recordings, is.
The evaluation table in the README is the clearest statement of intent. It compares Fun-CosyVoice3-0.5B-2512 against Seed-TTS, MiniMax-Speech, F5-TTS, Spark TTS, FireRedTTS2, Index-TTS2, VibeVoice, HiggsAudio-v2, VoxCPM and GLM-TTS on character error rate and speaker similarity. The base model reports 1.21 CER and 78.0 SS on test-zh, and the RL variant reports 0.81 CER. Those numbers come from the project's own table; treat them as claims to reproduce on your data, not as settled results.
The mechanism: LLM backbone, flow matching, and a 150ms streaming claim
The architecture is visible in the repository layout rather than in prose. There is a cosyvoice/ package, a runtime/ directory, a third_party/ directory pulled in as a git submodule, and separate vllm_example.py and webui.py entry points at the top level. The README's roadmap traces the history: flow matching training arrived in 2024/07, streaming inference with kv cache and sdpa for RTF optimisation in 2024/08, a 25Hz CosyVoice2-0.5B in 2024/12, vLLM support for CosyVoice2-0.5B in 2025/05, and triton/trtllm runtime support contributed by NVIDIA in 2025/08.
Two mechanisms deserve attention because they shape what you can build. First, bi-streaming: the README states support for both text-in streaming and audio-out streaming, with latency as low as 150ms. That is a claim about the streaming path, not about batch synthesis, and it is the feature that makes the project usable for interactive voice agents rather than offline dubbing. Second, pronunciation inpainting: the README says the model supports inpainting of Chinese Pinyin and English CMU phonemes, which gives you a way to correct a mispronounced word without regenerating the whole utterance.
Text normalisation is handled without a traditional frontend module, per the README. That is a design choice with consequences. It removes the rule-based pipeline that usually sits in front of a TTS model, but it also means number and symbol handling is learned behaviour, and the README does not document how to override it when it gets a date or a unit wrong.
Installing CosyVoice and running a first synthesis
The README gives a conda recipe. Clone with submodules, because third_party/ is a submodule and a plain clone leaves it empty. The README notes that if the submodule clone fails on network errors, you should rerun the update command until it succeeds.
git clone --recursive https://github.com/FunAudioLLM/CosyVoice.git
cd CosyVoice
git submodule update --init --recursiveThen create the environment. The README pins python=3.10 and installs requirements.txt through the Aliyun mirror with a trusted-host flag. If sox causes compatibility problems, the README gives apt-get on Ubuntu and yum on CentOS.
conda create -n cosyvoice -y python=3.10
conda activate cosyvoice
pip install -r requirements.txt -i https://mirrors.aliyun.com/pypi/simple/ --trusted-host=mirrors.aliyun.com
sudo apt-get install sox libsox-devModels are not bundled. The README recommends downloading the pretrained checkpoints and the CosyVoice-ttsfrd resource, and shows the ModelScope SDK path first, with a Hugging Face path for users outside China.
from modelscope import snapshot_download
snapshot_download('FunAudioLLM/Fun-CosyVoice3-0.5B-2512', local_dir='pretrained_models/Fun-CosyVoice3-0.5B')
snapshot_download('iic/CosyVoice2-0.5B', local_dir='pretrained_models/CosyVoice2-0.5B')The repository ships example.py at the top level, alongside examples/ directories for grpo, libritts and magicdata-read. The README does not print the contents of example.py, so read that file for the exact call signature rather than guessing at argument names. Expect the first run to spend its time loading checkpoints from pretrained_models/ before any audio appears.
Where CosyVoice breaks: platform limits, licence exposure, and the streaming claim
The dependency list is the first hard constraint. requirements.txt pins torch==2.3.1 and torchaudio==2.3.1 against a CUDA 12.1 extra index, and pins onnxruntime-gpu==1.18.0 for Linux while falling back to onnxruntime==1.18.0 on darwin and win32. deepspeed==0.15.1 and the tensorrt-cu12 packages are marked linux-only. The README does not document a supported Windows or macOS training path, and the presence of Linux-only training dependencies suggests there is not one. If your team runs Windows workstations, plan for a Linux box or WSL.
The second constraint is the licence boundary. The repository carries Apache-2.0, which is permissive for code. The README links model checkpoints on ModelScope and Hugging Face under separate model pages, and the repository does not state in the documentation available here whether those checkpoint licences match the code licence. For a commercial deployment, that distinction matters more than the code licence does.
The third is the latency figure. The README states 150ms latency for the bi-streaming path. It does not state the hardware, the utterance length, or whether that figure includes the text frontend. Do not put that number in a capacity plan without measuring it yourself on your own hardware.
Finally, the README does not document rollback, checkpoint compatibility between CosyVoice 1.0, 2.0 and 3.0, or a migration path for fine-tuned models. Upgrading a checkpoint in production is an open question.
CosyVoice against F5-TTS, Index-TTS2 and VoxCPM
The README's own evaluation table makes the comparison concrete, which is unusual and useful. F5-TTS is listed at 0.3B with 1.52 CER and 74.1 SS on test-zh; Index-TTS2 at 1.5B with 1.03 CER and 76.5 SS; VoxCPM at 0.5B with 0.93 CER and 77.2 SS. Fun-CosyVoice3-0.5B-2512 is listed at 0.5B with 1.21 CER and 78.0 SS, and the RL variant at 0.81 CER and 77.4 SS.
The difference in approach is not just the numbers. CosyVoice ships training, inference and deployment scripts together, with GRPO training examples under examples/grpo/ and a runtime/ directory for optimised serving. A project like F5-TTS is primarily an inference-oriented release. If you intend to fine-tune on your own speakers, the training path in this repository is the reason to pick it. If you only need to synthesise text with a pretrained voice, the extra training machinery is weight you carry without using.
The dialect coverage is the other differentiator that the table does not capture. Nine languages plus 18+ Chinese dialects and accents is a specific bet on Chinese-language deployment. If your product is English-only, that coverage is not worth the dependency chain.
Maintenance status, upgrade cost, and licence implications
The repository is not archived, and the last push was on 2026-05-25. That is the fact to work from. The roadmap entries run through 2025/12, when Fun-CosyVoice3-0.5B-2512, its RL model and the training/inference scripts were released, along with a ModelScope Gradio space. There is no release list in the documentation available here, so versioning is effectively by checkpoint name and roadmap date.
Upgrade cost is dominated by the pinned dependency set. torch 2.3.1, transformers 4.51.3, onnxruntime-gpu 1.18.0 and the tensorrt-cu12 10.13.3.9 trio are all pinned to exact versions. Moving any one of them means revalidating the others, and the Linux-only markers on deepspeed and tensorrt mean the validation has to happen on Linux. Budget for a rebuild of the environment rather than an in-place upgrade.
On licence: the repository is Apache-2.0, which permits commercial use and modification of the code. The checkpoints are hosted separately on ModelScope and Hugging Face, and the documentation available here does not state their terms. A team shipping a cloned voice in a commercial product should read the model page for the specific checkpoint before building on it, and should treat voice cloning of real speakers as a separate consent question from the software licence. This is not legal advice.
Editorial conclusion
Adopt CosyVoice if you need zero-shot cloning across Chinese, English and the dialects it lists, and you are willing to run a conda environment with pinned torch 2.3.1 and onnxruntime-gpu 1.18.0 on Linux. Do not adopt it if you need a hosted endpoint, a Windows training path, or a licence review that survives a strict legal read of the weights. Before anything else, verify that your GPU and driver can run the pinned CUDA 12.1 wheels, and confirm the exact snapshot_download identifier for the checkpoint you intend to use, since the README lists several.
Frequently asked questions
Is CosyVoice open source?
Yes. The repository is licensed Apache-2.0 and the primary language is Python. The pretrained checkpoints are hosted separately on ModelScope and Hugging Face, and the documentation available here does not state their licence terms.
How to install CosyVoice?
The README gives a conda recipe: clone with --recursive, create an environment with python=3.10, and install requirements.txt. Models are downloaded separately through the ModelScope or Hugging Face SDK before inference will run.
How to use CosyVoice?
The repository ships example.py at the top level, plus vllm_example.py and webui.py. The README does not print the contents of example.py, so read that file for the exact call signature and point it at a checkpoint under pretrained_models/.
What is voice cloning and how does it work?
Voice cloning synthesises speech in a target voice from a reference recording. CosyVoice supports zero-shot cloning, meaning the reference speaker does not need to be in the training set, and the README states it works across languages and Chinese dialects.
What are the newest text-to-speech models available?
The README's evaluation table lists Fun-CosyVoice3-0.5B-2512 alongside F5-TTS, Spark TTS, FireRedTTS2, Index-TTS2, VibeVoice, HiggsAudio-v2, VoxCPM and GLM-TTS. Those are the projects the CosyVoice maintainers chose to benchmark against.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/qwenaudio-cosyvoice)