Moonshine Voice: on-device speech to text, intent recognition and text to speech
Very low latency speech to text, intent recognition, and text to speech, for building voice agents and interfaces
At a glance
- What is it?
- Moonshine Voice is an MIT-licensed C++ toolkit with Python, JavaScript/WASM, iOS, Android, macOS, Linux, Windows and Raspberry Pi bindings for building real-time voice agents without an account or API key. The README is thin on failure modes, and the models are a separate download from the library.
- Who is it for?
- Adopt Moonshine Voice if you are building a voice interface that must run locally, on a phone, a Raspberry Pi or a browser tab, and you can accept that the library and the models are two separate things to manage. Do not adopt it if you need a hosted service with a managed model catalog, or if you need the legacy non-English non-streaming models under a permissive licence.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 29 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 20, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What problem Moonshine Voice is aimed at
Most speech to text stacks assume a server. You record audio, ship it to an endpoint, wait for a transcript, then act on it. That round trip sets a floor on latency, and it means the audio leaves the device. Moonshine Voice is built for the opposite arrangement: the README states that everything runs on-device, with no account and no API keys, and that the library is optimized for live streaming by doing work while the user is still talking.
The intended audience is narrow and specific. Developers building voice agents and voice interfaces, on hardware that ranges from a phone to a Raspberry Pi. The repository ships example directories for Android, C++, iOS, macOS, Python, Raspberry Pi, web and Windows, which tells you the maintainers expect the same core to be embedded in very different runtimes rather than called over HTTP.
The project covers three jobs in one library: speech to text, intent recognition, and text to speech. Intent recognition is the piece that distinguishes it from a plain transcription library. If you already have a transcript pipeline and only need words on a screen, you are using more surface area than you need.
How the library is put together
The repository layout separates the engine from its consumers. `core/` holds the C++ implementation, `language-bindings/` holds the per-language wrappers, and `examples/` holds runnable samples per platform. The Python package and the JavaScript/WASM build are bindings over that core rather than reimplementations, which is why the README can claim one library across all those targets.
The streaming design is the part worth understanding before you adopt it. The README says the models are optimized for live streaming and that the system does work while the user is still talking. That is a different contract from a batch transcriber: you feed audio continuously and consume partial results, rather than submitting a finished file. Anything you build on top has to tolerate revision, because a streaming recognizer can change its output as more audio arrives.
The models are trained from scratch rather than fine-tuned from an existing checkpoint. The README points to a Hugging Face open ASR leaderboard space for the accuracy claim and to `micro/README.md` for the small end of the range, which it describes as tiny 1MB models. That spread matters: the same API can back a desktop agent and a constrained embedded target, but you are choosing a point on that curve yourself.
Intent recognition and text to speech sit alongside transcription in the same library, so a voice agent can be assembled without a second runtime. The README does not describe how intent recognition is implemented, only that it is part of the toolkit. Treat that as undocumented until you read the API reference.
Installing moonshine-voice and running the mic demo
The README's quickstart is two commands. The first installs the Python package from PyPI, the second starts live microphone transcription in English. Note that the package name on PyPI is `moonshine-voice`, not `moonshine`.
pip install moonshine-voice
moonshine-voice mic --language enAfter the second command you should see transcription output as you speak. The `--language en` flag is the only option shown in the README, and it is the only one you should assume works without checking the API reference.
For platforms other than Python, the README does not put install steps in the repository root. It directs you to the Quickstart page on the documentation site for per-platform installation, and to the Examples page for runnable samples. The `examples/` directory in the repository mirrors that list: `examples/python/`, `examples/web/`, `examples/android/`, `examples/ios/`, `examples/macos/`, `examples/c++/`, `examples/windows/`, `examples/raspberry-pi/`.
If you are driving the library from a coding agent, the repository ships a skill definition rather than relying on the agent's training data. The README gives this command:
npx skills add moonshine-ai/moonshine --skill moonshine-voiceThe stated reason is that the skill teaches the current API shape so the agent does not reach for Whisper or the old `DialogFlow` names. That is a useful signal about how much the API has moved: if your agent's priors are stale, it will generate code against an interface that no longer exists.
Where Moonshine Voice is the wrong choice
The licence carve-out is the first real constraint, and it is easy to miss because the headline says MIT. The README states that the models are MIT by default in every language and at every size, with one exception: the legacy non-streaming models for languages other than English stay under the non-commercial Moonshine Community License, enumerated in `LICENSE`. If your product needs a permissive licence for a non-English language and you were planning to use a legacy non-streaming model, the library being MIT does not help you. You need to read the enumerated list.
On-device execution is also a resource commitment, not a free win. The README advertises models down to roughly 1MB, which implies the larger, more accurate models need meaningfully more memory and compute than that. The README does not publish a memory or CPU budget per model, so you cannot size your target from the repository alone. You have to run the mic demo on the actual device.
Streaming changes your application design. If your downstream logic assumes a final transcript arrives once, a recognizer that emits and revises partial results will break it. The README does not document a rollback or finalization contract for partial results, so that behaviour is something to establish empirically before you build on it.
Finally, the README is a pointer document. Installation details, model selection, accuracy numbers and the API surface all live on the documentation site or in the API reference. If you need a single self-contained repository you can read end to end before adopting, this is not that.
How it differs from Whisper-based pipelines
The obvious comparison is Whisper, and the README makes it explicitly: it claims the models are trained from scratch and range from higher accuracy than Whisper Large V3 down to tiny 1MB models, citing the Hugging Face open ASR leaderboard space. It also warns that a coding agent may default to Whisper if you do not install the skill, which tells you the maintainers see Whisper as the incumbent.
The difference in approach is architectural, not just numeric. A typical Whisper deployment receives a complete audio buffer and returns a transcript; the model is not designed around emitting results mid-utterance. Moonshine Voice is built the other way round, with streaming as the primary mode and latency reduction coming from processing audio while the user is still speaking. If your workload is transcribing uploaded files after the fact, that streaming design buys you nothing and you are paying for a capability you will not use.
The second difference is packaging. Whisper is a model plus whatever runtime you assemble around it. Moonshine Voice is a library with bindings for Python, JavaScript/WASM, iOS, Android, macOS, Linux, Windows and Raspberry Pi, plus intent recognition and text to speech in the same package. If you need a voice agent rather than a transcript, that breadth replaces a fair amount of glue code. If you need only transcription on a server, it replaces nothing.
Maintenance, releases and upgrade cost
The repository is not archived, and the last push was on 2026-08-31. The release history is short and recent: v0.1.5 on 2026-08-24, v0.1.3 on 2026-08-18, and v0.1.2 on 2026-08-13. Every release in that list is a 0.1.x version, which is the honest signal here: the project is pre-1.0 and the API is still moving.
The README gives direct evidence of that churn. The skill instructions exist specifically because the agent should not reach for "the old `DialogFlow` names", meaning a previous naming scheme was replaced. Combined with three patch releases inside a month, the practical upgrade cost is that you should expect to re-read the API reference when you bump versions, and you should pin the version you develop against rather than tracking the default branch.
The repository keeps a `CHANGELOGS.md` at the top level, which is where the actual per-release changes are recorded. The README does not document a deprecation policy or a support window, so there is no stated guarantee about how long a given API shape will survive.
On licensing, the library is MIT per the README, and the models are MIT by default with the enumerated non-commercial exception for legacy non-streaming non-English models. That exception is a licensing question, not a purely technical one, and the enumerated list in `LICENSE` is the document that decides it. Read it before you ship.
Editorial conclusion
Adopt Moonshine Voice if you are building a voice interface that must run locally, on a phone, a Raspberry Pi or a browser tab, and you can accept that the library and the models are two separate things to manage. Do not adopt it if you need a hosted service with a managed model catalog, or if you need the legacy non-English non-streaming models under a permissive licence. Before committing, run the mic demo on your target hardware, confirm which model files your platform binding expects, and read LICENSE for the exact list of models that fall under the Moonshine Community License rather than MIT.
Frequently asked questions
What is Moonshine Voice?
It is an open source AI toolkit for developers building real-time voice agents and applications, covering speech to text, intent recognition and text to speech. According to the README, everything runs on-device with no account or API keys, and it is optimized for live streaming.
How do I install moonshine-voice?
The README quickstart installs the Python package with pip and then starts live microphone transcription. Other platforms are covered on the Quickstart page of the documentation site rather than in the repository root.
Which platforms does Moonshine Voice support?
The README lists Python, JavaScript/WASM, iOS, Android, macOS, Linux, Windows and Raspberry Pi, with runnable samples in the Examples section of the documentation and matching directories under examples/ in the repository.
Is Moonshine Voice licensed under MIT?
The library is licensed under the MIT License per the README. The models are MIT by default in every language and at every size, except the legacy non-streaming models for languages other than English, which remain under the non-commercial Moonshine Community License and are enumerated in LICENSE.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/moonshine-ai-moonshine)