Open-source project
BoltzmannEntropy/MimikaStudio avatar
BoltzmannEntropy/MimikaStudio

MimikaStudio: local voice cloning and audiobook generation on Apple Silicon

MimikaStudio - A local-first application for macOS (Apple Silicon) + Agentic MCP Support

741 stars99 forksDartGPL-3.0

At a glance

What is it?
MimikaStudio is a Flutter desktop app for macOS on Apple Silicon that wraps several TTS and voice-cloning engines behind one UI and a local MCP server. It is local-first, but the licence is not open source in the usual sense.
Who is it for?
Adopt MimikaStudio if you work on an Apple Silicon Mac and want voice cloning, TTS and document-to-audiobook conversion on one machine without sending audio to a hosted API. Skip it if you need Linux or Windows binaries today, or if you need an OSI-approved licence for redistribution, because the source is under BSL-1.1 and binaries under a separate distribution licence.
Can I use it commercially?
Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly Dart, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What MimikaStudio does that a plain TTS script does not

Running a text-to-speech model locally is easy. Keeping four of them, a document parser, a job queue and a voice-prompt library in a state where you can actually use them is not. MimikaStudio is aimed at that second problem. The README describes four integrated capabilities: voice cloning from as little as 3 seconds of reference audio, text-to-speech with fast and expressive model families, a Read Aloud document reader with sentence-level highlighting across PDF, DOCX, EPUB, Markdown and TXT, and document-to-audiobook conversion with queueable chapter generation and reusable voice presets.

The audience is narrow and specific: people on an Apple Silicon Mac who want generated speech to stay on the machine. The README states the application runs fully on-device and includes first-run model download management. That last phrase is doing real work. A first run that pulls several hundred megabytes to a few gigabytes of weights is the step where most local TTS setups fail, and the project treats it as a product surface rather than a footnote.

It also runs as an agentic voice cloning server with a jobs queue for TTS, cloning and audiobook pipelines, exposed through both UI and API paths. That dual exposure is why the MCP topic appears on the repository. If you only ever want to convert one paragraph to speech, this is more machinery than you need. If you want an agent to call a local speech service, the machinery is the point.

The engine matrix, and why the model list is the real feature

The supported model table is where the design becomes legible. Kokoro-82M handles fast TTS in English with British RP and American voices. Qwen3-TTS appears in 0.6B and 1.7B sizes, each in a Base variant for voice cloning across 10 languages, a CustomVoice variant for preset speakers in English, Chinese, Japanese and Korean, and 8-bit versions of both. Chatterbox Multilingual covers voice cloning across 23 languages. Supertonic-2 is a multilingual ONNX model covering English, Korean, Spanish, Portuguese and French.

That is a deliberate refusal to pick a winner. Kokoro is the cheap path when you want speed and do not need cloning. Qwen3-TTS is the cloning path with the widest language coverage. Chatterbox covers the most languages for cloning. Supertonic-2 is the odd one out because it runs through ONNX rather than the MLX path the README advertises for the rest of the app.

The 8-bit variants matter more than they look. A 1.7B parameter model in full precision is a different memory proposition from the same model quantised, and on a laptop with unified memory the difference decides whether the app is usable alongside a browser. The README does not publish a memory table for each variant, so the honest answer is that you should measure the variant you intend to run on your own machine before standardising on it.

One inconsistency worth flagging: the repository's requirements.txt lists onnxruntime with a comment tying it to a standalone CosyVoice3 runtime, but CosyVoice3 does not appear in the supported model table in the README. The dependency list is broader than the documented feature set, which is common in a project that has shipped several releases in quick succession, but it means you should treat requirements.txt as the build input and the model table as the supported surface.

Installing on macOS and running a first clone

The repository ships install.sh at the top level alongside install.bat, a backend directory, a flutter_app directory and a requirements.txt. The Python dependency file states its own install command and notes that chatterbox is installed separately, which install.sh handles.

bash
pip install -r requirements.txt

The file also documents that IndexTTS-2 is not on PyPI and is installed from git inside install.bat and install.sh with --no-deps, which is why its runtime dependencies such as sentencepiece and matplotlib are listed explicitly in requirements.txt. If you install the Python dependencies by hand rather than through the script, you will need to reproduce that step yourself.

bash
pip install --no-deps git+https://github.com/index-tts/index-tts.git

The README points to the project site for getting started and states that macOS binaries are provided. The latest release, v2026.04.1, is described as adding in-app PDF page preview for audiobook source documents, disabling the old 7-day expiration and the Polar and LemonSqueezy purchase flow in the Pro UI, and removing pricing and buying paths from the website in favour of direct GitHub release downloads. So the intended first-run path for most users is now a direct download of a release binary, not a source build.

After launch, the AI Models screen is where first-run model downloads happen. Pick the smallest engine that matches the task: Kokoro-82M for plain English TTS, a Qwen3-TTS Base variant if you need cloning. Voice prompts are managed on their own screen, and the Jobs screen shows queue state for TTS, cloning and audiobook pipelines.

The README does not document a CLI entry point for the backend, and it does not document rollback if a model download is interrupted. Treat the first run as something to do on a stable connection.

Where MimikaStudio is the wrong tool

The platform constraint is the first hard boundary. The README says the codebase is cross-platform but that macOS binaries are the only ones currently provided, and that Windows support is planned for a future release. Linux is not mentioned at all. If your inference host is a Linux box with an NVIDIA card, this project is not a candidate, regardless of how well the model list matches your needs.

The Apple Silicon requirement is not incidental. The README states the app is optimized for Apple Silicon with native Metal acceleration via MLX, and the badge line reads macOS (Apple Silicon) with MLX-Audio. That means an Intel Mac is out of scope by design.

The licence is the second boundary and it is the one people miss. The repository topics and the project description point at GPL-3.0, but the README states plainly that source code is licensed under Business Source License 1.1 and that binary distributions fall under the MimikaStudio Binary Distribution License, with LICENSE, BINARY-LICENSE.txt and the website licence page as the references. Those two things cannot both be true of the same release. Anyone planning to redistribute the app, embed it in a product, or fork it for commercial use needs to read the actual licence files before writing code, and should not assume the topic tag is authoritative.

Third, the audiobook path is only as good as its parsers. The README lists PDF, DOCX, EPUB, Markdown and TXT as supported inputs, and v2026.04.1 added in-app PDF page preview, which suggests PDF handling is the area that needed the most iteration. Scanned PDFs without a text layer are not addressed anywhere in the README.

Finally, a local-first design is not a privacy guarantee by itself. The app runs on-device, but voice cloning of a real person's voice carries legal exposure in many jurisdictions regardless of where the computation happens. The repository does not provide legal guidance, and it should not be read as doing so.

How it compares with a single-engine pipeline

The obvious alternative is to skip the app and drive one model directly. Kokoro has its own Python package, listed in requirements.txt as kokoro>=0.9.4, and a short script can load it and synthesise a paragraph in a few lines. That approach gives you exactly one voice family, no UI, no queue, no document parsing, and no model download manager. It is faster to set up and it will not break when an unrelated engine's dependency is upgraded.

The second alternative is a hosted TTS API. You trade the Apple Silicon requirement for a network dependency, per-character billing and the fact that your reference audio leaves the machine. The README's central claim is that MimikaStudio avoids that last part, and for anyone working with sensitive recordings that is the deciding difference rather than a nice-to-have.

The third comparison is against the Python-only local stacks, the kind of thing the README gestures at with its reference to ebook2audiobook-style demo blocks. Those give you document-to-audiobook conversion from a command line, and they run on Linux. What they generally do not give you is a persistent job queue with a UI, a voice-prompt library, or an MCP endpoint that an agent can call. MimikaStudio's bet is that the orchestration layer is the product, not the models, which are all public anyway.

That bet has a cost. Every engine you add is another set of pinned dependencies. The requirements.txt pins transformers to 4.57.3 with the comment that qwen-tts requires it, and numpy is capped below 2.0.0. Those pins are the price of hosting several engines in one environment, and they will conflict with whatever else you have in that Python installation. Use a virtual environment.

Maintenance, releases and what the licence costs you

The repository is not archived, and the last push was on 2026-04-01. The release cadence is dense: v2026.03.10 on 2026-03-22, v2026.03.11 on 2026-03-24, and v2026.04.1 on 2026-04-01. Three releases in ten days suggests a period of active iteration rather than a settled codebase, and the v2026.04.1 notes describe removing a purchase flow and a 7-day expiration from the Pro UI, which is a product-direction change rather than a bug fix.

That matters for upgrade cost. The release note explicitly disables the old expiration and the Polar and LemonSqueezy purchase flow, and moves distribution to direct GitHub release downloads. If you built anything around the earlier Pro gating, it no longer applies. The README does not document a migration path for that change, and it does not document rollback between releases.

The dependency surface is the ongoing cost. requirements.txt pins transformers and caps numpy, and it notes that IndexTTS-2 must be installed from git with --no-deps, which means its transitive dependencies are your responsibility. A git-installed package has no version pin, so a rebuild months from now may pull different code.

On licensing, the practical point is this: BSL-1.1 is a source-available licence with usage restrictions that typically convert to an open source licence after a change date, and the README does not state a change date. Binary distributions are covered by a separate licence file. If your organisation has a policy against source-available licences, or you intend to redistribute the app, that policy question comes before any technical evaluation. This is not legal advice; read LICENSE, BINARY-LICENSE.txt and the website licence page, and route the question to whoever handles licensing where you work.

Editorial conclusion

Adopt MimikaStudio if you work on an Apple Silicon Mac and want voice cloning, TTS and document-to-audiobook conversion on one machine without sending audio to a hosted API. Skip it if you need Linux or Windows binaries today, or if you need an OSI-approved licence for redistribution, because the source is under BSL-1.1 and binaries under a separate distribution licence. Before committing, check the model download sizes for the Qwen3-TTS variant you intend to run, and read LICENSE and BINARY-LICENSE.txt rather than assuming GPL-style terms.

Frequently asked questions

Is voice cloning illegal?

The repository does not provide legal guidance, and running the models on-device does not change the legal position. MimikaStudio's README describes cloning from as little as 3 seconds of reference audio and says nothing about consent, jurisdiction or permitted use, so that question has to be answered outside the project.

Is there an AI that can mimic someone's voice?

Yes, and MimikaStudio ships several. The supported model table lists Qwen3-TTS 0.6B and 1.7B Base for voice cloning across 10 languages, Chatterbox Multilingual for cloning across 23 languages, and 8-bit variants of the Qwen3-TTS models.

How to protect yourself from voice cloning?

The README does not cover defensive measures against voice cloning. Its focus is the generation side: local cloning, TTS, Read Aloud and audiobook creation, with the app running fully on-device so reference audio is not sent to a hosted service.

How much does AI voice cloning cost?

MimikaStudio itself is downloaded from GitHub releases, and the v2026.04.1 release notes state that the old 7-day expiration and the Polar and LemonSqueezy purchase flow were disabled and that pricing and buying paths were removed from the website. The README does not state a price. The real cost is local: model downloads on first run and the Apple Silicon Mac required to run them.

Official sources

  1. BoltzmannEntropy/MimikaStudio on GitHub
  2. License: GPL-3.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/boltzmannentropy-mimikastudio.svg)](https://hysenlabs.com/projects/boltzmannentropy-mimikastudio)