# Qwen3-TTS on Apple Silicon: Local Voice Cloning with MLX

> A Python menu-driven wrapper that runs Qwen3-TTS voice cloning, voice design and preset voices offline on M-series Macs through MLX, with manual model downloads and no licence file in the repository.

**kapi2800/qwen3-tts-apple-silicon** — Run Qwen3-TTS text-to-speech locally on Mac (M1/M2/M3/M4). Voice cloning, voice design, custom voices. 100% offline using MLX.

- Repository: https://github.com/kapi2800/qwen3-tts-apple-silicon
- Stars: 572 · Forks: 83
- Language: Python
- License: not declared
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/kapi2800-qwen3-tts-apple-silicon

## What qwen3-tts-apple-silicon actually is

The repository is a small Python front end. The top level holds main.py, requirements.txt, an outputs/ directory, a voices/ directory and the README. Everything that produces speech comes from the mlx-audio package, which requirements.txt pins to a specific git commit of Blaizzy/mlx-audio, plus the Qwen3-TTS model weights that you download yourself from mlx-community on Hugging Face. The project's own contribution is the menu, the model selection logic and the wiring between them.

That matters when you decide whether to adopt it. You are not adopting a text-to-speech engine; you are adopting a convenience layer over mlx-audio and the MLX-converted Qwen3-TTS checkpoints. If the menu suits you, the wrapper saves real setup work. If you need programmatic control, the wrapper is something you will end up reading or replacing.

The intended user is a Mac owner with Apple Silicon who wants speech synthesis without sending text to a cloud service. The README states the offline property plainly: no cloud, no API keys, completely offline. It targets M1, M2, M3 and M4 machines and lists Python 3.10+ as a requirement.

## Voice cloning, voice design and the three model roles

The project exposes three capabilities, and each one maps to a different set of weights. Custom Voice uses preset speakers with emotion and speed controls. Voice Design generates a voice from a text description, and the README's own example is "calm British narrator". Voice Cloning takes a reference audio clip and reproduces that speaker.

The README says cloning works from a 5-second sample and recommends clean 5-10 second clips for best results. That is the honest constraint of the approach: the reference clip is the entire conditioning signal, so background noise or a clipped sample degrades the output in ways the tool cannot correct. There is no documented denoising step.

Each capability exists in two sizes. The 1.7B models are labelled Pro and "Best Quality"; the 0.6B models are labelled Lite and "Faster, Less RAM". The README's requirement section puts memory at roughly 3GB for Lite and roughly 6GB for Pro, which is the number to check against your machine before downloading anything. Speed control is limited to three documented settings: Normal at 1.0x, Fast at 1.3x and Slow at 0.8x.

## Why MLX changes the memory and heat profile

The README's central technical claim is a comparison table between standard PyTorch models and MLX-converted models: 10+ GB of RAM against 2-3 GB, and 80-90 degrees Celsius against 40-50 degrees. The note under the table attributes those figures to testing on an M4 MacBook Air with 1.7B models. Those are the project's numbers, not an independent measurement, and the README does not describe the test conditions or the text length used, so treat them as indicative rather than reproducible.

The mechanism behind the claim is stated in the README: MLX runs natively on the Apple Neural Engine and GPU. The 8-bit quantisation in the model names, such as Qwen3-TTS-12Hz-1.7B-CustomVoice-8bit, is the other half of the explanation for the lower memory figure. A fanless MacBook Air is the case where this matters most, because sustained synthesis on a thermally limited chassis is exactly where a 40-degree difference changes whether you can keep working.

One detail worth noticing is the 12Hz token rate in every model name. The MLX checkpoints are all 12Hz variants, which is a property of the released weights rather than a setting this project exposes.

## Installing qwen3-tts-apple-silicon on macOS

The README gives a five-minute quick start. It clones the repository, creates a virtual environment, installs the pinned requirements and installs ffmpeg through Homebrew. The virtual environment is not optional in practice: the troubleshooting table's first entry is mlx_audio not found, and the fix is to activate the environment before running anything.

```bash
git clone https://github.com/kapi2800/qwen3-tts-apple-silicon.git
cd qwen3-tts-apple-silicon
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
brew install ffmpeg
```

After that, the models are a manual step. The README does not provide a download command. It tells you to open the Hugging Face page for the model you want, click Download, and place the resulting folder in models/. The expected layout is shown in the README as three sibling directories under models/, named Qwen3-TTS-12Hz-1.7B-CustomVoice-8bit, Qwen3-TTS-12Hz-1.7B-VoiceDesign-8bit and Qwen3-TTS-12Hz-1.7B-Base-8bit. The troubleshooting table warns that a Model not found error usually means the folder names do not match exactly, so copy the names from the model pages rather than typing them.

```bash
source .venv/bin/activate
python main.py
```

Running main.py prints the manager menu with six numbered options, three for the Pro models and three for the Lite models, plus q to exit. Selecting Custom Voice asks you to pick a preset speaker and set emotion and speed. Selecting Voice Design asks for a description. Selecting Voice Cloning asks for a reference clip. For long passages, the README notes that you can drag a .txt file directly into the terminal instead of pasting text, and that typing q or exit returns you to the previous menu. Generated audio lands in the outputs/ directory that already exists in the repository.

## Where qwen3-tts-apple-silicon is the wrong tool

The entry point is an interactive menu, and the README documents no command-line flags, no configuration file and no importable API. If you want to synthesise a thousand lines in a batch job, or call the model from another Python program, this repository does not give you a documented path. You would be reading main.py and reimplementing the call against mlx-audio directly, at which point the wrapper has stopped being useful.

The dependency situation is the second constraint. The requirements file pins mlx-audio to a git URL at commit 9349644ccbd62eb10900852228f7b952c566def3, and it also pins transformers==5.0.0rc3, a release candidate. Pinning a commit is reproducible, but it means you get whatever that commit contained, with no version number to reason about and no upgrade path described in the README. The README does not document rollback, and there is no changelog or release history to consult.

The third constraint is the licence. The repository has no licence file at the top level, and the README does not state one. The Qwen3-TTS weights and the mlx-audio dependency carry their own terms, which this project does not restate. If you need a clear legal position before shipping something, this repository does not give you one, and you would have to check the upstream projects yourself.

Finally, the hardware requirement is absolute. This runs on Apple Silicon only. An Intel Mac, a Linux workstation with an NVIDIA card, or a Windows machine is out of scope, and no part of the README suggests otherwise.

## How it compares with the upstream Qwen3-TTS and mlx-audio

The README lists Qwen3-TTS by Alibaba as the original project and mlx-audio as the MLX framework for audio models. Those are the two real alternatives, and they differ from this repository in a specific way: both are libraries, and this is an application.

Going directly to Qwen3-TTS means working with the original implementation rather than the MLX-converted checkpoints. The README's comparison table is the reason not to: it claims 10+ GB of RAM and 80-90 degrees Celsius for standard models against 2-3 GB and 40-50 degrees for MLX. On a fanless MacBook Air, that gap decides whether the task is practical at all.

Going directly to mlx-audio means skipping the menu and calling the audio pipeline yourself. That is the better choice if you need scripting, batching or integration into an existing service, because you control the interface and you are not dependent on a wrapper whose only documented interface is a numbered prompt. The cost is that you do the wiring that main.py currently does for you, including model loading and output handling. The trade is convenience against control, and the README makes no argument for the wrapper beyond the five-minute setup.

## Maintenance, upgrades and licence exposure

The last push to the repository was on 2026-05-02. There are no releases in the repository. That combination means upgrades are manual: you pull the repository, reinstall from requirements.txt, and hope the pinned mlx-audio commit still builds against whatever else changed on your machine. Because mlx-audio is pinned to a commit rather than a version, there is no "newer version available" signal to act on, and the README does not describe an upgrade procedure. The practical cost of maintaining a working install is therefore tied to the pinned commit and to the transformers release candidate, both of which you would have to move deliberately.

On licensing, the repository itself has no licence file, so the code's terms are unstated. The model weights come from mlx-community on Hugging Face and the audio pipeline comes from Blaizzy/mlx-audio, each with its own terms that this project does not reproduce. Voice cloning adds a separate question that no licence file answers: whether you have permission to clone the voice in your reference clip. That is a consent and rights issue rather than a software one, and the README does not raise it.

## Conclusion

Adopt qwen3-tts-apple-silicon if you have an M-series Mac, roughly 3GB of free RAM for the 0.6B models and a text-to-speech task that must stay offline, such as cloning a consenting speaker's voice for internal narration. Do not adopt it if you need a supported library with a versioned API, a documented licence, or a non-interactive pipeline, because the entry point is a menu, the MLX audio dependency is pinned to a git commit, and the repository ships no licence file. Verify three things before you commit: that the model folder names under models/ match the README's names exactly, that the pinned mlx-audio commit still resolves, and that you have the right to use the Qwen3-TTS weights and any voice you clone.

## FAQ

### Can I use Python on Apple Silicon Macs?

Yes, and this project depends on it: the README requires Python 3.10+ on an M1, M2, M3 or M4 Mac, and the quick start creates a virtual environment with python3 -m venv .venv before installing requirements.txt.

### Is TTS available on Mac?

This repository provides text-to-speech on Apple Silicon Macs by running Qwen3-TTS locally through MLX, with no cloud service or API key involved. It covers preset voices, voice design and voice cloning from a reference clip.

### How to TTS on Mac?

Clone the repository, create a virtual environment, run pip install -r requirements.txt, install ffmpeg with brew, place a downloaded model folder under models/, then run python main.py and pick one of the six menu options.

### How to change TTS voice on Mac?

The manager menu offers three routes: Custom Voice for the preset speakers with emotion and speed settings, Voice Design to generate a voice from a text description, and Voice Cloning to reproduce a speaker from a 5 to 10 second reference clip.

## Sources

- [Issues](https://github.com/kapi2800/qwen3-tts-apple-silicon/issues)
- [kapi2800/qwen3-tts-apple-silicon on GitHub](https://github.com/kapi2800/qwen3-tts-apple-silicon)
- [README](https://github.com/kapi2800/qwen3-tts-apple-silicon/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/kapi2800-qwen3-tts-apple-silicon
