Model or dataset
QwenLM/Qwen3-TTS avatar
QwenLM/Qwen3-TTS

Qwen3-TTS: Open-Source Text-to-Speech Models with Streaming and Voice Clone

Project brief: Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streaming speech generation, free-form voice design, and vivid voice cloning.

13,596 stars1,763 forksPythonApache-2.0

At a glance

What is it?
Qwen3-TTS is an Apache 2.0 licensed series of text-to-speech models from Alibaba Cloud's Qwen team. It covers 10 languages, provides voice design from text descriptions, 3-second voice cloning, and a streaming mode with end-to-end latency as low as 97 milliseconds.
Who is it for?
Qwen3-TTS is the right choice for teams that need open-source, locally runnable TTS with multilingual coverage across 10 languages and built-in streaming generation. The 0.6B models reduce hardware requirements while keeping voice clone and custom timbre capabilities.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Qwen3-TTS Is and Who It Targets

Qwen3-TTS is a series of text-to-speech models developed by the Qwen team at Alibaba Cloud. The README describes it as providing comprehensive support for voice clone, voice design, ultra-high-quality speech generation, and natural language-based voice control. It targets developers building speech synthesis pipelines who want local inference, streaming output, and multilingual coverage without licensing restrictions.

The supported languages are Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian. The README notes that multiple dialectal voice profiles are also available to cover regional speech variation within supported languages.

The README was released with the model announcement on 2026-01-22. The Apache 2.0 license covers both research and commercial use. The package is installable via pip as qwen-tts and integrates with Hugging Face Transformers and ModelScope as the distribution channels for model weights.

Model Architecture: Tokenizer, Streaming, and End-to-End Design

The Qwen3-TTS architecture has two main components. The Qwen3-TTS-Tokenizer-12Hz encodes input speech into discrete codes and decodes codes back into audio. The README describes it as achieving efficient acoustic compression and high-dimensional semantic modeling, preserving what it calls paralinguistic information and acoustic environmental features.

The generative model uses a discrete multi-codebook language model architecture. The README contrasts this with a traditional LM+DiT (diffusion transformer) approach, stating that the end-to-end LM architecture bypasses information bottlenecks and cascading errors inherent in cascaded systems.

Streaming generation is handled by a Dual-Track hybrid streaming architecture. The README states this allows the same model to support both streaming and non-streaming output. In streaming mode, the first audio packet is output immediately after a single character of text is processed, achieving end-to-end synthesis latency as low as 97 milliseconds. This latency figure is stated as meeting real-time interactive scenario requirements.

Model Variants: 0.6B vs 1.7B and Capability Differences

Five models are in the current release, organized by size and task:

The 1.7B series has three variants: Qwen3-TTS-12Hz-1.7B-VoiceDesign generates speech from a user-provided voice description; Qwen3-TTS-12Hz-1.7B-CustomVoice provides style control over 9 premium timbres covering combinations of gender, age, language, and dialect; Qwen3-TTS-12Hz-1.7B-Base performs 3-second rapid voice cloning from user audio and is designed for fine-tuning.

The 0.6B series has two variants: Qwen3-TTS-12Hz-0.6B-CustomVoice supports the same 9 premium timbres as the 1.7B variant but without instruction-based control; Qwen3-TTS-12Hz-0.6B-Base matches the 1.7B-Base in voice clone capability at lower hardware cost.

All five models support streaming generation and all 10 languages. The key capability gap between the 0.6B and 1.7B models is instruction control: only the 1.7B-VoiceDesign and 1.7B-CustomVoice models include the instruction control column. The README describes instruction control as supporting flexible adjustment of acoustic attributes such as timbre, emotion, and prosody via natural language.

Installing Qwen3-TTS and Downloading Models

The Python package installs from PyPI:

bash
pip install qwen-tts

Model weights load automatically when a model is first initialized by name. For environments where network access during execution is restricted, the README provides manual download commands using the ModelScope CLI:

bash
pip install -U modelscope
modelscope download --model Qwen/Qwen3-TTS-Tokenizer-12Hz  --local_dir ./Qwen3-TTS-Tokenizer-12Hz
modelscope download --model Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice --local_dir ./Qwen3-TTS-12Hz-1.7B-CustomVoice

The Hugging Face collection at huggingface.co/collections/Qwen/qwen3-tts is the alternative download source for users outside mainland China. The README also describes a DashScope API option for users who prefer cloud inference without local hardware. The DashScope API is documented at the Aliyun model studio help page linked in the README.

A local web UI demo can be launched via the qwen-tts-demo entry point registered by the package. The vllm/ section of the README describes running models with vLLM as an inference server, which is useful for serving the models at higher throughput than single-request inference. The finetuning/ directory in the repository contains fine-tuning scripts for adapting the Base models to custom voices beyond the built-in 9 timbres.

Limitations: Repository Activity and Hardware Requirements

The last push to the repository was on 2026-03-17, which is more than six months before the review date of 2026-09-28. The repository has no GitHub releases; the model weights are distributed through Hugging Face and ModelScope rather than through GitHub. The pyproject.toml records version 0.1.1. Teams adopting Qwen3-TTS should confirm the current state of the repository against their requirements before committing to it in production.

The pyproject.toml lists transformers 4.57.3, accelerate 1.12.0, torchaudio, and onnxruntime as dependencies. Pinned transformers and accelerate versions can conflict with other packages in environments that maintain multiple model libraries. The torchaudio dependency implies a CUDA-capable GPU is expected for full-speed inference, though the README does not document minimum GPU memory requirements for each model variant. The librosa and soundfile dependencies cover audio file reading and writing. The sox dependency handles audio format conversion.

Voice design, the natural-language description feature, is only available in the 1.7B-VoiceDesign model. The 0.6B models do not include a voice design variant, so teams that need natural-language voice specification must use the larger model.

The qwen_tts/cli/demo.py entry point provides the web UI demo, and the examples/ directory contains four Python script examples: test_model_12hz_base.py, test_model_12hz_custom_voice.py, test_model_12hz_voice_design.py, and test_tokenizer_12hz.py. These cover the four main usage modes. The finetuning/ directory provides fine-tuning scripts for teams that need to adapt the Base model to custom voice characteristics beyond the 9 built-in timbres.

Qwen3-TTS vs Coqui TTS: Different Architectural Approaches

Coqui TTS is an open-source text-to-speech framework that supports a variety of neural TTS model architectures including Tacotron 2, VITS, and XTTS. It is a general-purpose TTS training and inference toolkit that allows training custom models on user data and supports voice cloning through XTTS.

The difference in approach is scope versus depth. Coqui TTS is a framework for training and running many model architectures across many languages. Qwen3-TTS is a specific model series with a fixed architecture optimized for the features listed in the README: multilingual support across exactly 10 languages, streaming with low latency, and natural-language voice control. Qwen3-TTS does not provide a general training framework for new architectures.

Qwen3-TTS's natural-language voice control is a feature the README describes as allowing flexible control over acoustic attributes such as timbre, emotion, and prosody through text instructions. The README states the model achieves this by deeply integrating text semantic understanding, adaptively adjusting tone, rhythm, and emotional expression based on the meaning of the input text. Coqui TTS does not offer this instruction-based voice control at the model level.

For teams that need to train a TTS model on proprietary voice data from scratch, Coqui TTS's training pipeline is more relevant. For teams that need out-of-the-box multilingual TTS with streaming and voice cloning, Qwen3-TTS offers those features in a ready-to-use package without requiring a training run. The finetuning/ directory in the repository provides a path for adapting the base models.

Editorial conclusion

Qwen3-TTS is the right choice for teams that need open-source, locally runnable TTS with multilingual coverage across 10 languages and built-in streaming generation. The 0.6B models reduce hardware requirements while keeping voice clone and custom timbre capabilities. The 1.7B models add instruction-based voice control and voice design. The last push to the repository was on 2026-03-17, and the project has no GitHub releases; verify that the repository state matches current deployment needs before committing to it in production. The Apache 2.0 license covers commercial use.

Frequently asked questions

Is Qwen3-TTS free?

Yes. Qwen3-TTS is released under the Apache 2.0 license, which permits commercial use. The Python package qwen-tts is available on PyPI. Model weights are available for free download from Hugging Face and ModelScope.

What is Qwen3-TTS?

Qwen3-TTS is a series of text-to-speech models developed by Alibaba Cloud's Qwen team, released in January 2026. It provides 0.6B and 1.7B model variants supporting 10 languages, streaming generation with latency as low as 97 milliseconds, 3-second voice cloning, and natural-language voice design.

Is Qwen3-TTS open source?

Yes. The model code and training configuration are published on GitHub under the Apache 2.0 license. Model weights are distributed through Hugging Face and ModelScope. The repository has no GitHub releases; version 0.1.1 of the qwen-tts Python package reflects the current state.

how to use qwen3-tts

Install the qwen-tts Python package from PyPI. Model weights are downloaded automatically by name on first use, or can be manually downloaded from Hugging Face or ModelScope. The qwen_tts package provides Python examples for custom voice generation, voice design, voice cloning, and tokenizer encode/decode. A local web UI demo launches via the qwen-tts-demo CLI entry point.

Official sources

  1. Official README
  2. Project repository