VoxCPM2: Tokenizer-Free Text-to-Speech with 30-Language Support and Voice Cloning
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
At a glance
- What is it?
- VoxCPM2 is a 2B-parameter text-to-speech model built by OpenBMB that generates speech without discrete tokenization, using a diffusion autoregressive architecture to support 30 languages, real-time voice cloning, and 48kHz output. It is Apache-2.0 licensed and can be installed with a single pip command.
- Who is it for?
- Developers who need multilingual TTS with voice cloning in a single open-source package will find VoxCPM2 a direct option, given its Apache-2.0 license, pip installation, and support for Voice Design and Controllable Cloning. Projects that cannot meet the CUDA 12 requirement or cannot use Python 3.10 through 3.12 will need a different path.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 27 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
What Problem VoxCPM2 Addresses and Who It Is For
Most text-to-speech systems convert text into discrete tokens representing phonemes or sub-word units, then decode those tokens into audio. This tokenization step introduces a bottleneck: the quality of the synthesis is limited by how well the token vocabulary captures continuous speech characteristics.
VoxCPM is described in the README as 'a tokenizer-free Text-to-Speech system that directly generates continuous speech representations via an end-to-end diffusion autoregressive architecture, bypassing discrete tokenization.' The practical effect is that the model can generate highly natural, expressive audio without the quantization artifacts that discrete token systems introduce.
The primary audience is developers building multilingual voice applications, content creation tools, accessibility tools, and synthetic voice products. The Apache-2.0 license makes the model available for commercial use without royalty payments. Researchers working on speech synthesis and developers who want to integrate voice generation into Python applications are the other main groups this project targets.
Architecture: Diffusion Autoregressive Design and MiniCPM-4 Backbone
VoxCPM2 is a 2B parameter model trained on over 2 million hours of multilingual speech data. The README states it is built on a MiniCPM-4 backbone. The diffusion autoregressive architecture allows the model to generate speech in an autoregressive manner while using diffusion-based refinement, which the README credits with achieving natural and expressive synthesis.
The model outputs audio at 48kHz. The README notes that it accepts 16kHz reference audio and outputs 48kHz studio-quality audio using an AudioVAE V2 asymmetric encode/decode design with built-in super-resolution. This means there is no need for a separate upsampling step when using lower-quality reference recordings.
Streaming is supported. The README gives a real-time factor of approximately 0.3 on an NVIDIA RTX 4090. With acceleration via Nano-vLLM or vLLM-Omni, the RTF drops to approximately 0.13. An OpenAI-compatible API for VoxCPM2 is available through vLLM-Omni using PagedAttention.
Installing VoxCPM2 and Running Your First Synthesis
Installation requires Python 3.10 to 3.12 (not 3.13 or later), PyTorch 2.5.0 or higher, and CUDA 12.0 or higher. The package is on PyPI:
pip install voxcpmThe model weights are hosted on Hugging Face at openbmb/VoxCPM2. The Python API loads them automatically on first use via VoxCPM.from_pretrained:
from voxcpm import VoxCPM
import soundfile as sf
model = VoxCPM.from_pretrained(
"openbmb/VoxCPM2",
load_denoiser=False,
)
wav = model.generate(
text="VoxCPM2 is the current recommended release for realistic multilingual speech synthesis.",
cfg_value=2.0,
inference_timesteps=10,
seed=42,
)
sf.write("demo.wav", wav, model.tts_model.sample_rate)The generate method returns a waveform array. The soundfile library writes it to a WAV file. The cfg_value parameter controls classifier-free guidance strength, and inference_timesteps controls the number of diffusion steps. The README also shows how to download weights from ModelScope for users who cannot access Hugging Face.
Voice Design: Creating a Voice from a Text Description
VoxCPM2 supports a mode the README calls Voice Design, in which a new voice is generated from a natural-language description alone, with no reference audio required. The format is straightforward: place the voice description in parentheses at the start of the text parameter:
wav = model.generate(
text="(A young woman, gentle and sweet voice)Hello, welcome to VoxCPM2!",
cfg_value=2.0,
inference_timesteps=10,
seed=42,
)
sf.write("voice_design.wav", wav, model.tts_model.sample_rate)The description can include attributes such as gender, age, tone, emotion, and pace. The model interprets the description and synthesizes a voice matching those characteristics without any stored voice profile. This is useful for applications where a specific named voice is not required, such as audiobook narration with varied character voices or accessibility tools that need distinct speaker identities.
Where VoxCPM2 Has Limitations
The CUDA 12 requirement is a hard constraint. CPU inference and older CUDA versions are not supported according to the installation requirements documented in the README. This excludes macOS environments and machines without a compatible NVIDIA GPU, which limits deployment to Linux systems with supported hardware.
The model is 2B parameters. Downloading weights from Hugging Face requires sufficient disk space and bandwidth. The initial load time depends on network conditions and local storage speed. The README does not document a quantized or reduced-size variant for lower-resource environments.
The README includes a Risks and Limitations section in the table of contents, but the README does not document the specific risks in that section. The project is Apache-2.0 licensed, which permits commercial use, but any deployment that reproduces another person's voice without their consent raises legal and ethical questions independent of the software license. The README does not document VoxCPM's behavior when given a reference audio of an identifiable person without their consent.
Supported Languages and the CLI
VoxCPM2 supports 30 languages: Arabic, Burmese, Chinese, Danish, Dutch, English, Finnish, French, German, Greek, Hebrew, Hindi, Indonesian, Italian, Japanese, Khmer, Korean, Lao, Malay, Norwegian, Polish, Portuguese, Russian, Spanish, Swahili, Swedish, Tagalog, Thai, Turkish, and Vietnamese. The README also lists nine Chinese dialects including Cantonese, Wu, and Sichuan dialect.
A command-line interface is available as the voxcpm command, registered as a project script in pyproject.toml. The CLI entry point is voxcpm.cli:main. The pyproject.toml also lists a timestamps optional dependency group that adds stable-ts, which the README does not describe in the visible text but the dependency name suggests it supports timestamped output.
An alternative to VoxCPM2 for open-source multilingual TTS is Kokoro or Coqui TTS. Both are permissively licensed but do not offer the same architecture (diffusion autoregressive vs autoregressive or flow-matching). VoxCPM2's differentiator is the tokenizer-free approach and the Voice Design feature, which the alternatives do not replicate with the same natural-language prompt interface.
Maintenance Status and Licensing
The last push to the repository was on 2026-09-02. The most recent release, v2.0.3, was published on 2026-05-11 with changes described as fine-tuning validation, runtime stability, and streaming improvements. VoxCPM2 itself was released in April 2026 as a major update from the earlier VoxCPM1.5.
The repository is licensed under Apache-2.0, which permits commercial use, modification, and redistribution without royalty payments. The model weights on Hugging Face are also listed under Apache-2.0 in the README. This is a permissive choice that makes VoxCPM2 available for production products.
The pyproject.toml includes a dev dependency group with pytest and pre-commit tools. A fine-tuning web UI is available as lora_ft_webui.py at the repository root, and a Gradio demo app is at app.py, allowing local testing without writing Python code directly.
Editorial conclusion
Developers who need multilingual TTS with voice cloning in a single open-source package will find VoxCPM2 a direct option, given its Apache-2.0 license, pip installation, and support for Voice Design and Controllable Cloning. Projects that cannot meet the CUDA 12 requirement or cannot use Python 3.10 through 3.12 will need a different path. Before deploying a cloned voice in a production product, verify that the intended use complies with applicable regulations on synthetic voice reproduction in the relevant jurisdiction.
Frequently asked questions
How do I use VoxCPM2 for text-to-speech?
Install VoxCPM with pip install voxcpm, then load the model with VoxCPM.from_pretrained('openbmb/VoxCPM2') and call model.generate() with your text. The method returns a waveform array that can be written to a WAV file using the soundfile library. Voice Design is enabled by placing a parenthesized description at the start of the text parameter.
What is VoxCPM?
VoxCPM is a tokenizer-free text-to-speech system from OpenBMB that generates speech via a diffusion autoregressive architecture without discrete tokenization. VoxCPM2 is the latest version, a 2B parameter model supporting 30 languages, 48kHz audio output, voice cloning, and Voice Design from natural-language descriptions.
Is VoxCPM free to use commercially?
VoxCPM2 is licensed under Apache-2.0, which permits commercial use, modification, and redistribution without royalty payments. The model weights on Hugging Face are also listed under Apache-2.0. Independent of the software license, deployment that involves reproducing a specific person's voice may be subject to legal requirements in the relevant jurisdiction.
How do I install VoxCPM?
Run pip install voxcpm in a Python 3.10 to 3.12 environment with PyTorch 2.5.0 or higher and CUDA 12.0 or higher. Users who cannot access Hugging Face can download the model weights from ModelScope using the modelscope package and pass the local path to VoxCPM.from_pretrained.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/openbmb-voxcpm)
Community notes