# Vocello: on-device text-to-speech for Apple Silicon, built in Swift and MLX

> Vocello is an MIT-licensed voice studio for Mac and iPhone that runs speech generation locally on Apple Silicon through a first-party Swift and MLX runtime. The interesting part is the long-form streaming design and the seed pinning; the awkward part is macOS 26 and the iOS export purchase.

**PowerBeef/Vocello** — Vocello: a local, private voice studio for Apple Silicon. Write a script, pick or describe a voice, and generate speech on-device, faster than realtime on an 8 GB M2 Mac mini. Native Swift + MLX, no Python. Mac app out now, iPhone beta on TestFlight. (Formerly QwenVoice.)

- Repository: https://github.com/PowerBeef/Vocello
- Website: https://vocello.vercel.app
- Stars: 371 · Forks: 33
- Language: Swift
- License: MIT
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/powerbeef-vocello

## What Vocello solves, and who it is actually for

Most text-to-speech tooling assumes a server. You send text to an endpoint, wait, and get audio back, and the text you sent has left your machine. Vocello takes the opposite position. The README describes generation that runs locally after models are installed, with scripts, recorded references, transcripts, saved voices, and history staying in local app storage unless you export them. There is no account and no credit system.

The audience follows from that. The README targets Apple Silicon specifically, and states that generation is faster than realtime on an 8 GB M2 Mac mini. That is a low bar for hardware, which matters: an 8 GB machine is the floor of the current lineup, not a workstation. If you are producing narration, prototyping voice lines, or building a spoken-audio feature and you cannot or will not send source text to a third party, this is the shape of tool you want.

It is a Mac and iPhone product, not a library. The README points at a DMG for macOS 26 or newer and a TestFlight beta for iPhone. If your workflow is a Python script that shells out to a TTS binary, the repository does contain a CLI, but the README frames the CLI as a caller that can pass an explicit seed for reproducible evidence, alongside benchmark callers. Treat the app as the primary surface.

## How the Swift and MLX runtime works, and where long-form fits

The README is explicit that Vocello is not a wrapper around a Python server. Generation runs through a first-party Swift runtime on MLX, and the repository layout backs that up: Sources/ and Packages/ sit next to QwenVoice.xcodeproj and a project.yml, with a pytest.ini and a benchmarks/ directory alongside them. The Python that remains appears to be test and benchmark scaffolding rather than the inference path.

Models are not bundled. The README says model installation downloads pinned model artifacts from Hugging Face, and the settings screen is where you install and manage the model package for each voice mode. That is a deliberate split: the app ships small, and the weights arrive on first use. It also means the first run is a network operation even though generation is not.

The most distinctive mechanism is long-form handling. Scripts past 900 characters become long-form projects. The README describes planned segments that stream one after another while you listen along, then join into a single finished file with a per-segment map in History. That is a real engineering choice rather than a UI nicety. A single long synthesis call holds the whole utterance in memory and gives you no feedback until it finishes; segmenting and streaming gives you audible progress and a per-segment record you can inspect when one segment comes out wrong.

Three voice modes sit on top of that pipeline. Built-in Voice uses one of nine Qwen3 speakers. Voice Design generates a voice from a plain-language brief. Voice Cloning captures a reference, requires a visible consent acknowledgment in Settings, and saves the result to a voice library. Ten languages are listed with automatic detection: Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian.

## Installing Vocello on macOS and generating a first take

The README gives a direct download rather than a package manager or a build-from-source instruction for end users. The release asset is named for the OS floor, which is a useful reminder that macOS 26 or newer is a hard requirement, not a suggestion. The README links the asset directly from the v2.4.0 release:

```bash
https://github.com/PowerBeef/Vocello/releases/download/v2.4.0/Vocello-macos26.dmg
```

After installing, the first real step is not writing a script. It is fetching a model. Open Settings and install the model package for the voice mode you intend to use. The model downloads are pinned artifacts fetched from Hugging Face, so this step needs network access. Generation itself does not.

With a model in place, the shortest path to a result is Built-in Voice. Choose one of the nine Qwen3 speakers, set the language, and pick a delivery. The README is unusually candid about the delivery presets: four of them (Neutral, Calm, Whisper, Sad) come through reliably, while the other four are directional hints that shape energy and pace without guaranteeing the named emotion on every take. If your first attempt sounds off, the preset you chose is a plausible cause before the model is.

For a reference-based workflow, Voice Cloning on the Mac accepts WAV, MP3, AIFF, M4A, FLAC, OGG, or WebM imports, and recording in the app works on both platforms. A transcript improves conditioning but is optional. Generation is gated behind the consent acknowledgment in Settings, so expect to confirm that before anything runs.

Reproducibility is handled through seeds. Every finished take records the seed it used. Pinning that seed from History makes later generations reproduce the take exactly with the same settings; leaving it unpinned produces a fresh take. The Expressive, Balanced, and Consistent variation settings control how much take-to-take variety you get, and multi-line batches share one seed so their lines form a consistent performance.

## Where Vocello gets in your way

The platform floor is the first constraint. macOS 26 or newer and Apple Silicon are both required, and the README states this plainly. Anyone on an Intel Mac, or on a macOS version below 26, is out entirely. There is no Linux or Windows path, and the Swift and MLX runtime means one is not going to appear as a side effect of packaging.

The second constraint is the voice inventory. Nine built-in speakers is a small cast for anything with more than a couple of characters, and the README's own honesty about delivery presets tells you the expressive range is uneven. Four presets land; four are directional. If your project needs a specific emotional register on every take, you will be doing retakes or building a curated emotion reference bank, which the README mentions as the mechanism that lets a cloned persona offer a delivery choice backed by its own verified reference clip.

The third is the iOS export boundary. The README states that for the upcoming official iOS 3.0 release, Design & Clone Export is a one-time purchase with a US base price of $19.99, while generation, listening, internal History, voice enrollment, and Built-in audio exports remain free, and macOS and CLI exports remain unrestricted. The same paragraph says the iOS product is configured but that metadata, purchase testing and App Store review are still pending, and explicitly calls this not an availability claim. Read that as: the iOS purchase exists in the code and in the pricing plan, not as a shipped storefront. The README also notes that the purchase supports official iOS distribution rather than granting an exclusive licence, and that modified builds are not the official App Store app and create no Apple entitlement.

Finally, there is a privacy boundary worth naming precisely. Generation is local, but model installation downloads artifacts from Hugging Face. If your threat model forbids any outbound request, the install step is still an outbound request.

## Vocello against a server-side TTS API

The obvious alternative is a hosted text-to-speech API from a cloud provider. The difference is not quality, it is where the text goes and what you pay for. A hosted API takes your script over the network, bills per character or per request, and returns audio. Vocello takes your script nowhere, bills nothing for generation, and spends your own compute instead.

That trade flips in both directions. A hosted API runs on any machine with a network connection, scales horizontally, and needs no model download. Vocello requires Apple Silicon, requires macOS 26 or newer, and requires you to fetch model weights before the first run. A hosted API also gives you a stable endpoint that many services can call. Vocello is an app with a CLI, and the README frames the CLI around benchmark and reproducible-evidence callers rather than as a general server interface.

The closer comparison is a self-hosted Python TTS stack. That gives you a server, runs on Linux, and lets you wire the model into your own pipeline. It also means managing Python environments and GPU drivers, and it is exactly the shape the README positions Vocello against when it says Vocello is not a wrapper around a Python server. If you already run a Python inference stack and it works, Vocello is not obviously better; it is a different bet on a native runtime and a single-vendor hardware target.

## Licence, the export purchase, and what upgrades cost you

The repository is MIT-licensed, and the README states that the source remains MIT-licensed even as the iOS export purchase is introduced. The README's own framing of that decision is worth quoting in substance: the purchase supports official iOS distribution rather than granting an exclusive licence to the source, and people may build and modify the MIT-licensed code, including its export checks, subject to the licence terms. The stated tradeoff is that the project accepts this instead of adding a licensing server, obfuscation, or restrictions on forks.

Two consequences follow. First, the export gate is a business boundary, not a technical one, and the README says so. Second, the code licence does not clear content rights: third-party models and assets retain their applicable terms. Voice cloning in particular is about the rights to a person's voice, and the app's consent acknowledgment is a gate in the UI, not a legal opinion. This is not legal advice; if you plan to ship generated audio commercially, the model terms and the voice rights are separate questions you need answered.

Upgrade cost is where the repository is thinner. The releases list shows v2.4.0 on 2026-08-01, v2.3.0 on 2026-07-31, and v2.2.2 on 2026-07-25, and the last push to the default branch was on 2026-09-15. The README links a per-release changelog under docs/releases/, but it does not document a migration path between versions, and it does not document rollback. Model packages are pinned artifacts, which suggests model updates are explicit rather than silent, but the README does not describe how a pinned model is replaced or how much disk a full set of voice-mode packages takes. Check the model download settings screen on your own machine before you commit to a long-form project.

## Conclusion

Adopt Vocello if you are on an Apple Silicon Mac running macOS 26 or newer, want speech generation that never leaves the machine, and can live with a nine-voice built-in set plus design and cloning. Do not adopt it if you need Windows or Linux, if you are below macOS 26, or if you need a server-side API that many processes can call. Before building anything on it, verify three things: that the model package for the voice mode you need is downloadable from Settings on your machine, that the licence position on the voice you intend to clone is one you can defend, and that the iOS export purchase boundary in SECURITY.md matches how you plan to distribute.

## FAQ

### Does Vocello run entirely offline?

Speech generation runs locally after models are installed, and generated audio, recorded references, transcripts, saved voices, and history stay in local app storage unless you export them. The install step is the exception: the README states that model installation downloads pinned model artifacts from Hugging Face.

### What hardware and macOS version does Vocello need?

The README requires Apple Silicon and macOS 26 or newer, and the release asset is named Vocello-macos26.dmg. The README states that generation is faster than realtime on an 8 GB M2 Mac mini.

### Is Vocello a wrapper around a Python text-to-speech server?

No. The README states that Vocello is not a wrapper around a Python server and that generation runs through a first-party Swift runtime on MLX. The repository does contain Python files for tests and benchmarks, but the README places generation in the Swift runtime.

## Sources

- [License: MIT](https://github.com/PowerBeef/Vocello/blob/main/LICENSE)
- [PowerBeef/Vocello on GitHub](https://github.com/PowerBeef/Vocello)
- [Project website](https://vocello.vercel.app)
- [README](https://github.com/PowerBeef/Vocello/blob/main/README.md)
- [Releases](https://github.com/PowerBeef/Vocello/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/powerbeef-vocello
