Open-source project
breezeblue-ai/breeze-tts avatar
breezeblue-ai/breeze-tts

Breeze TTS 2: PyTorch inference for a bilingual, instruction-following speech model

Official PyTorch inference for Breeze TTS 2

557 stars93 forksPythonApache-2.0

At a glance

What is it?
BreezeBlue's inference repository wraps the Breeze TTS 2 checkpoint in a CLI, a streaming FastAPI endpoint and a Docker build. It is a research-oriented stack with a non-commercial weights licence and a 12 GB GPU floor.
Who is it for?
Adopt Breeze TTS 2 if you are doing research or internal prototyping on English or Chinese speech and you already have an NVIDIA GPU with at least 12 GB, plus a 24 GB card if you want the fast path.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 7 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Breeze TTS 2 solves, and who the inference repo is for

Breeze TTS 2 is an open-weight text-to-speech model that BreezeBlue describes as built for real-time interaction. The repository breezeblue-ai/breeze-tts is not the model itself; it is the official PyTorch inference code that loads the Breeze TTS 2 checkpoint and turns it into something you can actually call. The README states that the checkpoint is published on Hugging Face under BreezeBlue/breeze-tts-2 and that all required model components are included in it, so the repository and the weights are two separate downloads that have to meet on disk.

The audience is narrow and technical. You need a CUDA-capable NVIDIA GPU, Linux, and Python 3.10 or newer. The README puts eager inference at roughly 7.7 GiB of GPU memory and recommends a 12 GB card as the minimum, with the --fast-all path needing about 14.4 GiB and a 24 GB GPU. That rules out CPU-only laptops and most shared CI runners.

What the model offers is four related capabilities rather than one. Voice Clone reproduces a speaker from reference audio plus its exact transcript. Voice Design invents a voice from a natural-language description with no reference audio at all. Voice Direction keeps a reference speaker's identity but lets an instruction steer tone, emotion, pace and delivery. Vocal events are a smaller feature layered on top: inline markers such as (laugh), (cough), (clears throat) and (sigh) in English, or [笑], [咳嗽], [清嗓子] and [叹气] in Chinese. The project is bilingual, generating English and Chinese from a single model, and the README's own examples keep the instruction language matched to the target text.

How the inference path is wired: checkpoints, CLI flags and the streaming API

The repository layout is small and readable. infer.py sits at the top level as the command-line entry point. breeze_infer/ holds the package, including the API module that the README invokes as python -m breeze_infer.api. configs/ and models/ hold configuration and model-side code, docker/ holds the image build, and tests/ holds the test suite. requirements.txt pins the runtime: torch==2.9.1, torchaudio==2.9.1, qwen-tts==0.1.1, transformers==4.57.3, numpy>=2.0 and soundfile>=0.13, with fastapi>=0.115, uvicorn>=0.30 and python-multipart>=0.0.18 for the server, and pytest>=8.0 and ruff>=0.12 for development.

The interesting part is how the CLI decides which mode you asked for. There is no --mode flag. The README states plainly that Voice Clone does not use an instruction, and that adding --instruction selects Voice Direction instead. So a reference audio file plus --ref-text plus --instruction is Voice Direction; the same reference inputs without --instruction are Voice Clone; and --instruction with no reference audio is Voice Design. That is a compact design, and it also means a stray --instruction silently changes what you are running.

Instruction following is tunable. The README directs you to use --cfg-scale 4 to strengthen instruction-following for both Voice Design and Voice Direction, and the streaming curl example sends cfg_scale as a form field alongside ref_audio, ref_text, text, instruction and seed. The API is single-concurrency, per the README's description of the start command, and its response is not a WAV file: it is streaming mono 24 kHz signed 16-bit little-endian PCM, which the example writes to voice_direction.pcm. Any client has to know that format up front.

The performance figures in the README are specific and worth reading carefully. Under 40 ms time to first audio and a 0.32 real-time factor, generating at roughly 3.1x real time, are both attributed to the warmed-up fast path on an NVIDIA H100. Those are the fast path's numbers on that GPU, not a general promise about your machine.

Installing Breeze TTS 2 and running a first voice clone

The install is two commands and a checkpoint download. Clone the inference code, then install the pinned dependencies. The README also points at a Docker build for the tested CUDA environment, which matters because the dependency list includes flash-attention-style architecture targeting through the docker build script.

bash
git clone https://github.com/breezeblue-ai/breeze-tts.git
cd breeze-tts
python -m pip install -r requirements.txt

The default Docker image targets H100/Hopper, which the README identifies as sm90. If you are on an A100, the build script takes an architecture override:

bash
FLASH_ATTN_CUDA_ARCHS=80 bash docker/build.sh

For a first real run, use Voice Clone, because it needs no instruction and the failure mode is easy to hear. You supply a clean reference clip and its exact transcript. The README's English example looks like this:

bash
python infer.py ../breeze-tts-2 \
  --ref-audio reference_en.wav \
  --ref-text "This is the exact transcript of the English reference audio." \
  --text "(sigh) It is good to hear your voice again after all this time." \
  --output outputs/voice_clone_en.wav

The first argument is the path to the checkpoint directory, here ../breeze-tts-2, which is why the quick start downloads the code and the weights as siblings. The output is a WAV file at the path you give. Listen for whether the timbre and rhythm match the reference; if they do not, the transcript is the first thing to check, because the README says --ref-text should match the complete spoken content of the reference audio, including any repetitions, and that the reference should be clean, non-looping speech with minimal background noise. A Chinese clone works the same way with a Chinese transcript and Chinese text, and the README's example uses [叹气] as an inline event.

If you would rather drive it over HTTP, start the server and post a form to the speech endpoint:

bash
python -m breeze_infer.api ../breeze-tts-2 --host 0.0.0.0 --port 7860
bash
curl -X POST http://127.0.0.1:7860/v1/audio/speech \
  -F "cfg_scale=4" \
  -F "[email protected]" \
  -F "ref_text=This is the exact transcript of the reference audio." \
  -F "text=(clears throat) We need to discuss what happened last night." \
  -F "instruction=Speak slowly with a restrained, serious tone." \
  -F "seed=42" \
  --output voice_direction.pcm

The server listens on port 7860 by default in these examples, and the response body is raw PCM, not a container format. You will need to play it back with a tool that accepts 24 kHz mono signed 16-bit little-endian samples, or convert it first.

Where Breeze TTS 2 gets in your way

The licence is the first constraint, and it is not a footnote. The repository's source code is Apache-2.0, but the README carries an explicit notice that the Breeze TTS 2 model weights, derivative models and self-hosted outputs are for research and non-commercial use only. A permissive code licence on a repository whose weights are restricted is a common pattern in this space, and it means the thing you actually want to ship is the restricted part. If your plan ends in a paid product or a hosted service, this checkpoint is the wrong tool until you have a separate agreement.

The second constraint is hardware. The eager path fits in about 7.7 GiB, which the README maps to a 12 GB GPU minimum. The fast path needs roughly 14.4 GiB and a 24 GB card. The headline latency and throughput numbers belong to that fast path on an H100, so a reader planning to run eager inference on a smaller card should not carry those figures over. The README does not publish a latency table for the eager path.

The third is the API shape. The streaming server is described as single-concurrency. That is a deliberate fit for the model's real-time interaction goal, but it is a hard ceiling for anything multi-tenant. There is no documented queueing, worker pool or horizontal scaling story in the README, and no documented rollback or checkpoint-versioning procedure either. If you need to serve several users at once, you are building that layer yourself.

Finally, the reference-audio requirements are stricter than they first look. Voice Clone wants clean, non-looping speech with minimal background noise, and the transcript must match the complete spoken content including repetitions. Noisy interview audio, music beds or clips with overlapping speakers are outside what the README describes as supported input. Garbage in produces a clone that drifts, and the README gives no cleanup or preprocessing step to rescue it.

Breeze TTS 2 against a conventional hosted TTS API

The obvious alternative is a hosted commercial TTS API, where you send text and receive audio and never touch a GPU. The difference in approach is where the model runs and who controls the voice. A hosted API gives you a fixed catalogue of voices and a per-character bill; Breeze TTS 2 gives you the checkpoint, the PyTorch runtime and the ability to design a voice from a sentence of description or clone one from a reference clip, at the cost of owning the hardware and the serving stack.

That trade is sharper here than usual because of the licence split. With a hosted API you are buying a commercial right along with the inference. With Breeze TTS 2 you get the code under Apache-2.0 but the weights under a research and non-commercial restriction, so the hosted route is the one that clears the commercial question without a negotiation. The counter-argument is control: reference-free voice design and reference-guided direction are model behaviours you configure with --instruction and --cfg-scale 4 rather than features a vendor has to expose. If your product depends on a specific voice character, that is the part you cannot rent.

Within the open-weight space, the README positions Breeze TTS 2 against other open-weight models by citing the Artificial Analysis TTS leaderboard, where it says the model ranks first among open-weight entries. That is the project's own claim about a third-party ranking, and the leaderboard is a moving target, so treat it as a starting point for your own listening test rather than a settled result. What is more checkable from the repository itself is the bilingual scope: one model for English and Chinese, with inline vocal events in both.

Maintenance, upgrades and what the licence asks of you

The repository is not archived, and the last push was on 2026-09-09, which is recent enough that the code is moving. There are no releases retrieved, so the practical upgrade unit is the commit on main plus the pinned requirements.txt. That pinning is a real benefit: torch==2.9.1, torchaudio==2.9.1, transformers==4.57.3 and qwen-tts==0.1.1 are fixed, so an upgrade is a deliberate act rather than something that happens when you reinstall. The cost is that you have to move those pins yourself when you want newer libraries, and the Docker build script is where CUDA architecture targeting lives, so a GPU generation change probably means rebuilding the image with a different FLASH_ATTN_CUDA_ARCHS value.

The checkpoint is a separate artefact from the code. The README says the weights live on Hugging Face at BreezeBlue/breeze-tts-2 and that all required model components are in that checkpoint. Nothing in the README describes versioning, rollback or compatibility guarantees between a given checkpoint and a given commit of the inference code, so keep a copy of the checkpoint you validated alongside the commit you validated it with. That is the only rollback story the documentation supports.

On licensing, the split is the thing to internalise. Apache-2.0 on the source code is a permissive licence with the usual patent grant and notice requirements. The weights, derivative models and self-hosted outputs carry a research and non-commercial restriction that the README states directly. Whether your specific use counts as commercial is a question for your own counsel, not something this repository answers; the README's License and responsible use section is the authoritative text to read before you build anything on top of it.

Editorial conclusion

Adopt Breeze TTS 2 if you are doing research or internal prototyping on English or Chinese speech and you already have an NVIDIA GPU with at least 12 GB, plus a 24 GB card if you want the fast path. Do not adopt it for a commercial product: the Apache-2.0 licence covers the source code, while the model weights, derivative models and self-hosted outputs are restricted to research and non-commercial use, so a paid product built on this checkpoint needs a separate arrangement with BreezeBlue. Before committing, verify three things on your own hardware: that your GPU has enough memory for the mode you plan to run, that your reference clips give clean clone output when the transcript matches the audio exactly, and that the streaming API's single-concurrency limit fits your traffic. The first concrete step is to run infer.py once with a short reference clip and listen to the result.

Frequently asked questions

Which TTS model is considered the best?

The README states that Breeze TTS 2 ranks first among open-weight models on the Artificial Analysis TTS leaderboard while outperforming frontier proprietary systems. That is the project's own description of a third-party ranking, and leaderboards change, so it is not a substitute for your own listening test.

Is voice cloning illegal?

The repository does not address the legality of voice cloning. It does restrict use: the README says Breeze TTS 2 model weights, derivative models and self-hosted outputs are for research and non-commercial use only, so the licence limits what you may do with cloned output regardless of other law.

Is there any free TTS?

Breeze TTS 2 has no cost stated in the repository, and its source code is Apache-2.0. The weights are not unrestricted, though: the README limits the model weights, derivative models and self-hosted outputs to research and non-commercial use.

Official sources

  1. breezeblue-ai/breeze-tts on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/breezeblue-ai-breeze-tts.svg)](https://hysenlabs.com/projects/breezeblue-ai-breeze-tts)