GLM-TTS: zero-shot voice cloning with a two-stage LLM and flow-matching pipeline
GLM-TTS: Controllable & Emotion-Expressive Zero-shot TTS with Multi-Reward Reinforcement Learning
At a glance
- What is it?
- GLM-TTS is an Apache-2.0 Python text-to-speech system from zai-org that clones a voice from 3 to 10 seconds of prompt audio, generates speech tokens with a Llama-based LLM, and converts them to waveforms with a flow-matching model. The interesting part is not the cloning; it is the phoneme-level control and the reinforcement-learning alignment, both of which are still partly unreleased.
- Who is it for?
- Adopt GLM-TTS if you need Chinese-first synthesis with a voice clone from a short prompt and you are willing to run a two-stage GPU pipeline yourself. Do not adopt it if you need a hosted API, a stable release tag, or English-only quality guarantees; the repository has no releases, the RL weights are listed as coming soon, and the README documents no rollback or versioning scheme.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 173 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap GLM-TTS targets: pronunciation control, not just cloning
Most open zero-shot TTS systems solve one problem well: given a short reference clip, produce speech in that voice. GLM-TTS treats cloning as table stakes and aims at two harder targets. The first is pronunciation accuracy for polyphones and rare characters. The README names the character 行, which can be read xíng or háng depending on context, as the motivating example, and describes a Phoneme-in mechanism that lets you override a specific word's reading while leaving the rest of the sentence as plain text. The second target is emotional expressiveness, addressed through a Multi-Reward Reinforcement Learning framework that the README says produces more natural emotional expression and prosody control.
The intended users are therefore not general app developers looking for a drop-in speech endpoint. They are teams building educational assessment tools, audiobook pipelines, or any Chinese-language product where a mispronounced surname or a flat delivery is a defect rather than a cosmetic issue. The README explicitly frames the phoneme feature around educational assessments and audiobooks. If your use case tolerates occasional pronunciation errors and does not need prosody control, the extra machinery here is overhead you will pay for in setup complexity.
How the two-stage pipeline moves from text to waveform
The architecture is split into two models rather than one end-to-end network. The first stage is an LLM based on the Llama architecture, which converts input text into speech token sequences. The second stage is a Flow Matching model that converts those tokens into a mel-spectrogram, after which a vocoder produces the waveform. The repository layout reflects this split directly: there are top-level llm/, flow/, frontend/, and cosyvoice/ directories, plus a grpo/ directory for the reinforcement-learning code.
Zero-shot cloning works by extracting speaker features from prompt audio rather than fine-tuning per speaker. The README states that 3 to 10 seconds of prompt audio is enough. That is a design constraint worth taking literally: the system is not doing speaker adaptation, so the quality ceiling is set by how well the prompt clip represents the target voice.
The Phoneme-in path adds a preprocessing step before the LLM sees the text. Inference follows a G2P to table lookup replacement to hybrid input workflow. A global grapheme-to-phoneme conversion produces a full phoneme sequence, a dynamic controllable dictionary identifies polyphones or rare characters and swaps in the specified target phonemes, and the mix of replaced phonemes and original text is fed to the model. Training uses random G2P conversion on parts of the text so the model learns to handle mixed input while retaining pure-text understanding. That is a sensible way to avoid a separate phoneme-only mode, but it also means pronunciation control depends on the quality of the G2P conversion and the dictionary, neither of which the README documents in detail.
Installing GLM-TTS and running the first inference
The README pins Python 3.10 through 3.12. Start by cloning the repository and installing the base requirements, which pin torch 2.3.1 and torchaudio 2.3.1 among roughly two dozen other packages.
git clone https://github.com/zai-org/GLM-TTS.git
cd GLM-TTS
pip install -r requirements.txtThe reinforcement-learning dependencies are separate and described as optional. They live under grpo/modules and involve cloning two external repositories plus downloading a wavlm_large_finetune.pth checkpoint into grpo/ckpt. If you only want inference, skip this step; the README marks it optional, and the RL-optimized weights are listed as coming soon anyway.
Next, fetch the model weights. The README says the complete set covers Tokenizer, LLM, Flow, Vocoder, and Frontend, and offers two sources.
mkdir -p ckpt
pip install -U huggingface_hub
huggingface-cli download zai-org/GLM-TTS --local-dir ckptModelScope is the alternative if HuggingFace is not reachable from your network. Then run the command-line demo. The README's example uses the bundled example_zh dataset and enables the cache.
python glmtts_inference.py \
--data=example_zh \
--exp_name=_test \
--use_cacheAdd --phoneme to enable phoneme capabilities. There is also a shell wrapper, glmtts_inference.sh, and a Gradio interface started with python -m tools.gradio_app, which is the fastest way to hear a result without writing any calling code. For NPU hardware the README gives a separate path: a CANN container image plus requirements_npu.txt installed with --no-build-isolation. The example container is quay.io/ascend/cann:8.5.1-910b-ubuntu22.04-py3.11, and the DEVICE variable is set to a /dev/davinci device node.
Where GLM-TTS gets thin: releases, languages, and unreleased weights
The most concrete limitation is that the repository has no releases. There is no tagged version to pin, so the practical unit of upgrade is a commit on main. The last push was on 2026-04-10, which is more than five months before today's date, so this is not a project you should describe as actively developed; treat the current state as a snapshot and plan accordingly.
The README lists two items as coming soon: a 2D Vocos vocoder update and model weights optimized via reinforcement learning. That second one matters. The headline claim is emotion and prosody control through multi-reward RL, but the RL-optimized weights are not in the download yet. You can run the RL code path, and the paper is on arXiv as 2512.14291, but the shipped checkpoints are not the RL-tuned ones. Anyone evaluating GLM-TTS specifically for expressive delivery should verify which weights they are actually loading before drawing conclusions.
Language coverage is also narrower than the feature list implies. The README says the system primarily supports Chinese and also supports English mixed text. That is not the same as bilingual synthesis. If your product is English-first, the examples directory does include example_en.jsonl and example_en1.jsonl, but the stated design center is Chinese, and the phoneme tooling is built around Chinese polyphones with pypinyin and jieba in the dependency list. A team needing balanced multilingual output should treat GLM-TTS as a Chinese system with English tolerance.
Finally, the README does not document rollback, version compatibility between checkpoints and code, or a supported hardware matrix beyond the GPU and NPU paths. The dependency pins are exact, which helps reproducibility but means a torch upgrade is a deliberate migration rather than a routine bump.
GLM-TTS compared with CosyVoice, which it vendors
The repository contains a top-level cosyvoice/ directory and lists HyperPyYAML, which is a CosyVoice-adjacent dependency. CosyVoice is the natural comparison point: it is also an Apache-licensed Chinese-first zero-shot TTS project, and GLM-TTS appears to reuse parts of its stack for the frontend or vocoder side rather than reimplementing everything.
The difference in approach is the acoustic model. CosyVoice's lineage centers on a supervised speech tokenizer feeding an autoregressive or flow-matching decoder trained on large speech corpora. GLM-TTS instead puts a Llama-architecture LLM in the first stage to generate speech tokens, then uses flow matching for the mel-spectrogram. The LLM stage is what makes the phoneme-in hybrid input workable, because the model can accept a mixed token stream of phonemes and text without a separate acoustic pathway.
The second difference is the alignment strategy. GLM-TTS adds a multi-reward RL stage on top of the base model, which CosyVoice does not ship in the same form. In principle that targets expressiveness rather than intelligibility, which is a different optimization target from what most TTS projects chase. In practice, as noted above, the RL-optimized weights are still listed as coming soon, so today the comparison rests mostly on the architecture rather than on shipped checkpoints.
If you want a project with tagged releases and a longer track record, CosyVoice is the safer starting point. If you specifically need phoneme-level override of Chinese polyphones, GLM-TTS is the one that documents that mechanism.
Licence and the cost of keeping GLM-TTS running
GLM-TTS is Apache-2.0, which permits commercial use, modification, and redistribution provided you keep the licence and notice files and state significant changes. That is a permissive licence and removes the most common legal blocker for product teams. It is not legal advice, and the model weights are downloaded separately from HuggingFace or ModelScope, so check the model card terms independently; the repository licence covers the code, and the README does not state that the weights carry identical terms.
The upgrade cost is the real maintenance burden. With no releases, staying current means tracking main and re-validating after each pull. The dependency file pins torch 2.3.1, torchaudio 2.3.1, transformers 4.57.3, and roughly 40 other packages at exact versions, so any environment drift shows up as an install failure rather than a subtle quality change. The optional RL path adds two git clones and a manual checkpoint download into grpo/ckpt, which is a step that will not survive an automated rebuild unless you script it yourself.
On hardware, the README gives a GPU path and an NPU path. The NPU path mounts driver libraries and device nodes from the host, which means the container is not self-contained; the Ascend driver version on the host becomes part of your reproducibility surface. Budget for that if you are targeting Ascend hardware.
Editorial conclusion
Adopt GLM-TTS if you need Chinese-first synthesis with a voice clone from a short prompt and you are willing to run a two-stage GPU pipeline yourself. Do not adopt it if you need a hosted API, a stable release tag, or English-only quality guarantees; the repository has no releases, the RL weights are listed as coming soon, and the README documents no rollback or versioning scheme. Before committing, verify that ckpt contains all five components (Tokenizer, LLM, Flow, Vocoder, Frontend), confirm your Python version is inside the 3.10 to 3.12 window, and run glmtts_inference.py with --data=example_zh to see whether the output matches your latency budget.
Frequently asked questions
What does GLM-TTS need to clone a voice?
The README states that zero-shot voice cloning works from 3 to 10 seconds of prompt audio, with speaker features extracted from that clip rather than fine-tuning per speaker. Example prompt audio is bundled under examples/prompt.
Which Python versions does GLM-TTS support?
The README says to use Python 3.10 through Python 3.12. Outside that window the pinned requirements, which include torch 2.3.1 and torchaudio 2.3.1, are not documented as supported.
How do I try the GLM-TTS demo without writing code?
The README gives an interactive web interface started with python -m tools.gradio_app, after the model weights have been downloaded into ckpt. There is also a shell script, glmtts_inference.sh, for the command-line path.
Does GLM-TTS support English?
The README says the system primarily supports Chinese and also supports English mixed text. The examples directory contains example_en.jsonl and example_en1.jsonl, but the stated design center is Chinese.
Where can I find the GLM-TTS paper?
The README links the technical report on arXiv under the identifier 2512.14291, published on 2025-12-17 according to the news section. The paper link is also in the header of the README alongside the HuggingFace and ModelScope model pages.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/zai-org-glm-tts)