CLI tool
k2-fsa/OmniVoice avatar
k2-fsa/OmniVoice

OmniVoice: zero-shot voice cloning TTS across 600+ languages

High-Quality Voice Cloning TTS for 600+ Languages

14,083 stars2,114 forksPythonApache-2.0

At a glance

What is it?
OmniVoice is a Python TTS model from k2-fsa that clones a voice from a 3 to 10 second clip and covers more than 600 languages. The install is straightforward; the hardware and the licensing of cloned voices are where adopters should slow down.
Who is it for?
Adopt OmniVoice if you are building multilingual speech output in Python and can supply a CUDA, MPS or XPU device, since the API is a few lines and the reference clip only needs to be 3 to 10 seconds. Do not adopt it if you need a maintained Studio-style web product, a ComfyUI node, or a guarantee about the rights attached to a cloned voice, because the repository documents none of those.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What OmniVoice solves, and for whom

Most open TTS stacks force a choice: a model with wide language coverage but a generic voice, or a cloning model that only speaks a handful of languages well. OmniVoice targets the overlap. The README describes it as a zero-shot text-to-speech model supporting over 600 languages, with a full list kept in docs/languages.md, and it clones a speaker from a short reference clip rather than requiring fine-tuning per voice.

The intended user is a Python developer who already has a reference recording and wants speech out of it. The README gives the practical shape of that input: use a 3 to 10 second reference clip, and note that longer audio slows down inference. That constraint matters more than it looks. It means the model is aimed at people who can produce a clean short sample, not at people who want to feed in an hour of archive audio and get a studio voice back.

A second group of users is served by voice design, where the README says you control a voice through assigned speaker attributes such as gender, age, pitch, dialect or accent, and whisper. That is a different workflow from cloning: no reference audio, just attributes. The two modes sit behind the same Python API, which is why the project reads as a toolkit rather than a single demo.

The diffusion language model design and what it changes at runtime

The README calls the architecture a diffusion language model-style design and presents it as the reason for both quality and speed, with an RTF as low as 0.025, described as 40 times faster than real time. Treat that number as a claim from the project rather than a measured result; it is not accompanied in the README by the hardware it was measured on, and the repository does not state the test conditions.

The data flow visible in the Python API is short. You load the model with OmniVoice.from_pretrained, passing a model id, a device_map and a dtype. You call generate with text plus either a reference clip or a saved prompt. The return value is a list of numpy arrays with shape (T,) at 24 kHz, which the README writes out with soundfile. There is no separate vocoder step exposed to the caller, and no tokenizer to configure.

One design decision worth flagging: if you omit ref_text, the model auto-transcribes the reference with Whisper. That is convenient, but it makes an ASR model part of your inference path. The README notes you can point asr_model_name at a local copy or a different Whisper model, and asr_device at cuda:1 or cpu. In a multi-GPU box that is the difference between a working pipeline and an out-of-memory error, so the escape hatch is real rather than decorative.

Installing OmniVoice with pip or uv, and a first clone

PyTorch comes first and OmniVoice second, and the README recommends a fresh virtual environment. The NVIDIA path pins a CUDA build explicitly:

bash
pip install torch==2.8.0+cu128 torchaudio==2.8.0+cu128 --extra-index-url https://download.pytorch.org/whl/cu128

Apple Silicon users install the plain wheels instead, and Intel Arc users go through Intel's XPU wheel index and then verify the backend:

bash
python -c "import torch; print(torch.xpu.is_available(), torch.xpu.device_count())"

With torch in place, the package itself has three routes. The stable release, the current master, or an editable clone:

bash
pip install omnivoice
pip install git+https://github.com/k2-fsa/OmniVoice.git

The uv route is a clone plus uv sync, and the README mentions a mirror flag, uv sync --default-index "https://mirrors.aliyun.com/pypi/simple", for users who need one. Note that pyproject.toml requires Python 3.10 or newer and pins transformers>=5.3.0, so an existing environment with an older transformers will need attention rather than a silent upgrade.

Before writing code, the fastest check is the bundled web UI, which the README launches with omnivoice-demo --ip 0.0.0.0 --port 8001. If that page loads and synthesises, your drivers and weights are fine. The first real use from Python is the cloning call:

python
from omnivoice import OmniVoice
import soundfile as sf
import torch

model = OmniVoice.from_pretrained(
    "k2-fsa/OmniVoice",
    device_map="cuda:0",
    dtype=torch.float16
)

audio = model.generate(
    text="Hello, this is a test of zero-shot voice cloning.",
    ref_audio="ref.wav",
    ref_text="Transcription of the reference audio.",
)

sf.write("out.wav", audio[0], 24000)

Expect a wav at 24 kHz in the working directory. On Apple Silicon the README says to use device_map="mps" and on Intel Arc device_map="xpu". If HuggingFace is unreachable, the README gives export HF_ENDPOINT="https://hf-mirror.com" as the workaround. If you plan to reuse the same voice, create_voice_clone_prompt encodes the clip once, prompt.save writes a .pt file, and VoiceClonePrompt.load brings it back in a later session without re-running the audio loading or the Whisper transcription.

Where OmniVoice is the wrong tool

The first limitation is hardware. Every documented path assumes an accelerator: CUDA, Apple MPS, or Intel XPU. The README does not describe a CPU-only inference path, and the XPU notes mention testing on an Arc A310 with 4 GB and an Arc Pro B50 with 16 GB, which tells you the memory floor is not trivial once the model and a Whisper pass are both resident. If your deployment target is a CPU-only container, this is the wrong project.

The second is the absence of a packaged application. People search for OmniVoice Studio, an OmniVoice app, and a ComfyUI node. The repository contains a Python package, a Gradio demo entry point, and a Colab notebook. Nothing in the README describes a Studio product or a ComfyUI integration, so anyone arriving with that expectation will be disappointed by what the repository actually ships.

The third is the reference audio itself. Auto-transcription with Whisper means a noisy or accented reference clip can be transcribed incorrectly, and the wrong transcript then conditions the clone. The README's own tips push toward short, clean clips. There is also no documented rollback or quality gate: the README does not describe how to detect a bad clone before it reaches a listener.

Finally, quality is not uniform across 600+ languages. The README lists coverage and points at docs/languages.md, but coverage is a count, not a per-language evaluation, and the README does not present per-language results. If you need a specific low-resource language, generate samples in that language before you plan around it.

How OmniVoice differs from Coqui TTS and other cloning stacks

The natural comparison is Coqui TTS. Coqui's cloning workflow is built around speaker embeddings and an encoder that conditions a vocoder, and its practical path to a new voice is fine-tuning on a dataset of that speaker. OmniVoice takes the zero-shot route: no training step, a 3 to 10 second clip, and the conditioning happens at inference. That difference decides your data budget. A few seconds versus minutes to hours of clean audio.

The trade-off runs the other way too. Fine-tuning on a specific speaker generally buys you consistency across long passages and unusual text, because the model has seen more of that voice. Zero-shot cloning depends entirely on what the encoder can extract from a short clip, which is why the README's advice about clip length is not a suggestion so much as an operating condition. If you have a large corpus for one speaker and need that exact voice to hold up over a long audiobook, a fine-tuning-first stack may serve you better.

OmniVoice does offer a middle path that the README documents: examples/run_finetune.sh and examples/run_finetune_lora.sh, with LoRA support behind the lora extra (peft>=0.20.0) and an omnivoice-merge-lora script to fold the adapter back in. So the choice is not strictly zero-shot versus fine-tuned; it is whether you want to start from a clone and improve it, or start from a corpus.

Maintenance, upgrade cost and the Apache-2.0 licence

The repository is not archived, and the last push was on 2026-09-14, which is recent. Releases are spaced rather than continuous: 0.1.5 on 2026-04-28, 0.2.0 on 2026-07-06, and 0.2.1 on 2026-07-16. That cadence suggests a research-led project where the paper and the model come first and the package follows, so budget for reading release notes rather than assuming drop-in upgrades.

The dependency surface is the real upgrade cost. pyproject.toml pins transformers>=5.3.0 and torch>=2.4, and the install instructions pin specific torch builds such as 2.8.0+cu128. Torch and transformers move quickly, and a project that tracks both will occasionally require you to move with them. The optional extras widen this further: eval pulls in funasr, s3prl and jiwer==3.1.0, while tn pulls in WeTextProcessing and num2words for text normalisation. A production deployment should install only the extras it needs rather than the full set.

On licensing, the package is Apache-2.0, and the LICENSE file sits at the repository root. That covers the code. It does not settle the rights in a cloned voice. The model card and the README do not grant you permission to imitate a specific person, and nothing in the repository addresses consent, attribution or jurisdiction. That question is separate from the software licence and depends on where you operate and whose voice you use. This is not legal advice; if you plan to ship cloned speech commercially, get an answer from someone qualified.

Editorial conclusion

Adopt OmniVoice if you are building multilingual speech output in Python and can supply a CUDA, MPS or XPU device, since the API is a few lines and the reference clip only needs to be 3 to 10 seconds. Do not adopt it if you need a maintained Studio-style web product, a ComfyUI node, or a guarantee about the rights attached to a cloned voice, because the repository documents none of those. Before committing, verify three things on your own hardware: that the pinned torch and transformers versions in pyproject.toml resolve in your environment, that your reference clips are clean enough for Whisper auto-transcription, and that you have written permission for every voice you intend to clone.

Frequently asked questions

What is OmniVoice?

It is a zero-shot text-to-speech model from k2-fsa that supports over 600 languages and can clone a voice from a short reference clip. The README also describes a voice design mode where you control gender, age, pitch, dialect or accent, and whisper.

How do I install OmniVoice?

Install PyTorch first for your accelerator, then install the package with pip install omnivoice, or pip install git+https://github.com/k2-fsa/OmniVoice.git for the latest source. The README also documents a uv route using git clone and uv sync, and recommends a fresh virtual environment either way.

How do I use OmniVoice?

The quickest start is the bundled Gradio UI, launched with omnivoice-demo --ip 0.0.0.0 --port 8001. From Python, load the model with OmniVoice.from_pretrained and call generate with your text plus a ref_audio clip and its ref_text.

Is OmniVoice free and open source?

The package is published under Apache-2.0, with the LICENSE file at the repository root, and the source is on GitHub under k2-fsa. The README does not describe a paid tier.

Is OmniVoice safe to use?

The repository does not document a safety filter or consent mechanism for cloned voices. The README covers clip length and transcription tips, but nothing about restricting whose voice can be cloned, so that responsibility sits with the user.

How big is OmniVoice?

The README does not state a parameter count or checkpoint size. The closest hardware signal is the Intel XPU note, which says the model was tested on an Arc A310 with 4 GB and an Arc Pro B50 with 16 GB.

Official sources

  1. Issues
  2. k2-fsa/OmniVoice on GitHub
  3. License: Apache-2.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/k2-fsa-omnivoice.svg)](https://hysenlabs.com/projects/k2-fsa-omnivoice)