Model or dataset
Blaizzy/mlx-audio avatar
Blaizzy/mlx-audio

mlx-audio: Speech Processing for Apple Silicon

A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.

7,962 stars733 forksPythonMIT

At a glance

What is it?
mlx-audio is a Python library for text-to-speech, speech-to-text, speech-to-speech, and music generation on Apple Silicon Macs. It wraps multiple open-source models and provides an OpenAI-compatible REST API and web interface.
Who is it for?
Use mlx-audio if you own an Apple Silicon Mac and want to run speech processing models locally without cloud dependencies. Do not use it if you need GPU support beyond Apple Metal, if you are on Intel Macs or Linux, or if you need production-grade audio quality comparisons across models.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Local speech processing on Apple Silicon without cloud APIs

Speech synthesis and transcription are computationally intensive tasks, and cloud APIs charge per request or subscription. mlx-audio brings these capabilities to Apple Silicon Macs, enabling offline operation and eliminating per-request costs. The README describes mlx-audio as providing fast and efficient text-to-speech, speech-to-text, speech-to-speech, music generation, and more on Apple Silicon. The library runs inference directly on the M1, M2, M3, and later chips that power modern Macs. The project released v0.5.7 on 2026-09-28, indicating active maintenance with frequent updates. Dependencies include MLX >= 0.31.1, NumPy, SciPy, and transformers, which install via pip. The library can run on iPad Pros and iPhones via Swift integration, extending offline capabilities beyond desktop Macs.

Installation and quick start

mlx-audio installs via pip. The basic installation provides command-line tools:

bash
pip install mlx-audio

For development or running the web interface with 3D visualization, install with additional dependencies:

bash
git clone https://github.com/Blaizzy/mlx-audio.git
cd mlx-audio
pip install -e ".[dev,server]"

Basic text-to-speech from the command line:

bash
mlx_audio.tts.generate --model mlx-community/Qwen3-TTS-12Hz-0.6B-CustomVoice-8bit --text 'Hello, world!' --voice Vivian

The command downloads the model from Hugging Face, generates speech in 12 kHz, and saves wav files numbered as audio_000.wav and audio_001.wav. Add `--play` to play audio immediately, or `--join_audio` to save one combined file instead of numbered segments. Add `--lang_code English` to specify the language hint.

From Python:

python
from mlx_audio.tts.utils import load_model

model = load_model("mlx-community/Qwen3-TTS-12Hz-0.6B-CustomVoice-8bit")
for result in model.generate("Hello from MLX-Audio!", voice="Vivian", lang_code="English"):
    print(f"Generated {result.audio.shape[0]} samples")

Over 15 TTS models with varying capabilities and languages

mlx-audio bundles over 15 text-to-speech models with varying capabilities. Kokoro is a compact 82 million parameter model supporting 8 languages (English, Japanese, Chinese, French, Spanish, Italian, Portuguese, Hindi) and providing fast, high-quality synthesis. Qwen3-TTS from Alibaba supports Chinese, English, Japanese, Korean, and additional languages with voice design control. OmniVoice handles zero-shot multilingual synthesis with voice cloning across 646+ languages without training. Higgs Audio v3 is a 4 billion parameter conversational TTS model supporting 100 languages with voice cloning and inline control tokens for pauses and emphasis. Ming Omni TTS offers multimodal generation with voice cloning and style control, supporting both English and Chinese. Voxtral is Mistral's 4 billion parameter model with 20 predefined voices and 9 languages.

For speech-to-text, the library supports Whisper models for general transcription, Nemotron for speaker diarization, and MedASR for medical transcription. Speech-to-speech pipelines combine a speech-to-text model with a TTS model to enable voice conversion or multilingual translation. Music generation models like MusicGen are integrated for composition tasks.

Models are available in multiple quantization levels: 3-bit, 4-bit, 6-bit, 8-bit, and full precision formats. Smaller models like KittenTTS offer edge-friendly options for constrained environments, with variants like Nano, Micro, and Mini that reduce model size. The README provides links to model cards on Hugging Face but does not provide guidance on which model offers the best quality for a given task, leaving model selection to the user's experimentation and testing.

Web interface with 3D visualization and OpenAI-compatible API

mlx-audio provides a web interface with interactive 3D audio visualization and an OpenAI-compatible REST API. The web interface allows you to select models, input text, choose voices, control speech speed, and generate or transcribe audio through a browser without command-line use. The interface is installed via pip install -e ".[dev,server]" and runs locally.

The API mirrors OpenAI's endpoints, so existing tools and applications built for OpenAI can be retargeted to mlx-audio. This compatibility simplifies integration into existing applications and reduces porting effort. The API server runs on localhost and supports batch requests and streaming audio output. The API is powered by FastAPI and Uvicorn, which are installed as server dependencies. This allows building custom applications that send requests to mlx-audio instead of calling cloud APIs.

Voice customization, cloning, and speed control

mlx-audio supports voice customization on models that include voice design features, such as Qwen3-TTS from Alibaba. For models with voice cloning, you provide a reference audio file and the system generates speech in that voice. OmniVoice and Higgs Audio v3 support zero-shot voice cloning without training: the system learns the voice characteristics from one or a few audio samples.

The library also supports adjustable speech speed control, allowing you to slow down or speed up generated audio. The exact mechanisms for voice cloning and speed control are not detailed in the README; they depend on the underlying model's architecture. The pyproject.toml shows optional dependencies for STT (sentencepiece, zstandard) and TTS (mistral-common, sentencepiece).

Quantization trade-offs for device compatibility

Models can be loaded in different quantizations: 3-bit, 4-bit, 6-bit, 8-bit, or full precision. Lower quantization reduces model size and memory requirements, enabling inference on constrained hardware like older M1 chips, iPad Pros, and iPhones. The trade-off is audio quality and generation speed; lower bit-depth models may sound slightly degraded compared to full-precision versions.

The README does not provide benchmarks comparing audio quality or speed across quantization levels. Users must experiment to find the balance suited to their hardware and quality requirements. The library supports three-bit quantization (the lowest provided), which can enable very small models on M1 Macs, but quality loss at this level is not documented.

When mlx-audio reaches its limits

mlx-audio is limited to Apple Silicon hardware. Users on Intel Macs, Linux, or Windows cannot use it. The library depends on Apple's MLX framework, which has no native equivalent on other platforms.

The breadth of supported models creates maintenance and quality assurance challenges. Each new model requires integration, testing, and documentation. The README lists over 15 TTS models but does not indicate which are well-tested, which have known issues, or which have quality limitations. Voice quality varies significantly by model, and the README does not provide guidance on which models sound natural, which are production-ready, or which work well for specific languages. Users must experiment with multiple models to find one suitable for their use case.

Quantized models have reduced audio quality. The README does not specify the quality loss at each quantization level (3-bit, 4-bit, 6-bit, 8-bit) or recommend minimum bit depths for specific use cases. A user deploying on M1 hardware with very limited RAM might need 3-bit quantization but would not know whether the resulting audio quality is acceptable without testing.

Editorial conclusion

Use mlx-audio if you own an Apple Silicon Mac and want to run speech processing models locally without cloud dependencies. Do not use it if you need GPU support beyond Apple Metal, if you are on Intel Macs or Linux, or if you need production-grade audio quality comparisons across models. Before deploying, download and test your target models (Kokoro, Qwen3-TTS, OmniVoice, etc.) on your hardware to ensure audio quality meets your needs, since the README does not provide quality comparisons or production-readiness recommendations.

Frequently asked questions

What is mlx-audio?

mlx-audio is a Python library for text-to-speech, speech-to-text, speech-to-speech, and music generation on Apple Silicon Macs. It bundles multiple open-source models and provides command-line tools, a web interface with 3D visualization, and an OpenAI-compatible REST API.

Is MLX made by Apple?

MLX is a machine learning framework developed by Apple for efficient inference on Apple Silicon. mlx-audio uses MLX as its foundation to optimize speech model inference on M-series chips.

Is MLX open source?

Yes. mlx-audio is MIT licensed and available on GitHub. MLX, the underlying framework, is also open source.

Official sources

  1. Blaizzy/mlx-audio on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/blaizzy-mlx-audio.svg)](https://hysenlabs.com/projects/blaizzy-mlx-audio)