# FluidAudio: on-device speech-to-text, diarization and TTS on the Apple Neural Engine

> FluidAudio is a Swift SDK that runs Parakeet, Kokoro, Silero and Sortformer CoreML models locally on Apple hardware. It is a good fit for always-on dictation and meeting apps on iOS and macOS, and the wrong fit for anything that needs to leave Apple's ecosystem.

**FluidInference/FluidAudio** — Frontier CoreML audio models in your apps - text-to-speech, speech-to-text, voice activity detection, and speaker diarization. In Swift, powered by SOTA open source.

- Repository: https://github.com/FluidInference/FluidAudio
- Website: https://docs.fluidinference.com/introduction
- Stars: 2,892 · Forks: 433
- Language: Swift
- License: Apache-2.0
- Published: 2026-08-08 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/fluidinference-fluidaudio

## What FluidAudio solves, and for whom

The README frames FluidAudio as a Swift SDK for fully local, low-latency audio AI on Apple devices, with inference offloaded to the Apple Neural Engine. That single design decision drives the whole feature set. Because the models run on the ANE and the README states the SDK avoids GPU and MPS entirely, the SDK targets background processing, ambient computing and always-on workloads, where a GPU-bound pipeline would compete with the rest of the app and drain the battery.

The intended user is an Apple-platform developer who wants speech recognition, speaker separation, voice activity detection or synthesis inside an app without sending audio to a server. The README lists shipped apps that fit that profile: dictation tools such as Voice Ink and Spokenly, meeting assistants such as Slipbox, and transcription utilities such as Whisper Mate. That list is a better signal of the target audience than any feature table, because every one of those apps has a privacy or latency reason to keep audio on the device.

It is not a general audio toolkit. There is no music analysis, no mixing, no acoustic measurement. If you arrived here looking for studio monitors or an audio interface, you are on the wrong project; the related-search terms around Fluid Audio monitors and FX80 refer to a different company.

## How the CoreML pipeline is put together

FluidAudio is a Swift package that wraps converted CoreML models. The repository has a Sources directory for the Swift code, a Tests directory, and a Documentation folder that holds the model catalog. The heavy lifting lives in the models, not in the Swift layer: the README states that all models are publicly available on HuggingFace, converted and optimized by the Fluid Inference team, and carry permissive MIT or Apache 2.0 licenses.

The feature set maps to distinct model families. Automatic speech recognition uses Parakeet TDT v3 at 0.6b parameters for batch transcription, with TDT and CTC variants covering 25 European languages plus Japanese, and SenseVoice and Paraformer for Mandarin Chinese. Streaming ASR uses Parakeet EOU at 120m parameters with end-of-utterance detection, which the README scopes to English only. Text-to-speech uses Kokoro at 82m parameters for parallel synthesis with SSML and pronunciation control across nine languages, and PocketTTS for streaming synthesis with voice cloning in six languages across 6L and 24L variants. Voice activity detection uses Silero models. Speaker diarization runs as either a streaming pipeline for real-time processing or an offline batch pipeline with clustering, and the SDK also exposes speaker embedding extraction for voice comparison and identification.

Two details in that list matter more than the rest. First, the streaming and batch paths use different models with different language coverage, so a design that assumes one ASR model can serve both real-time and file transcription will hit a language wall. Second, end-of-utterance detection is a model-level capability, not a heuristic layered on top, which is why it is limited to the languages that model was trained on.

If you already have a model you want to run, the README points to a separate project called möbius for converting your own model to the required format.

## Installing FluidAudio and running a first transcription

FluidAudio ships as a Swift package, with a FluidAudio.podspec also present at the repository root for CocoaPods users. The Package.swift file is the entry point for Swift Package Manager. The README does not print a step-by-step install block, so the exact dependency declaration is something you should read from Package.swift in the repository rather than copy from a blog post.

Once the package is added, the README states that the included models can be integrated with just a few lines of code. The first real check is whether the ANE path is actually being used on your target device, since that is the entire reason to choose this SDK over a CPU-based one. The most reliable early test is a short batch transcription of a known audio file, compared against a transcript you wrote yourself.

One practical constraint to plan for: models are distributed through HuggingFace, which means the first run involves a download step rather than shipping weights inside your binary. The README does not document an offline bundling workflow, so if your app must work on first launch without network access, treat that as an open question to resolve before you commit.

For a Python integration rather than a Swift one, the README points to Senko, a speaker diarization pipeline that the project cites as a good example of integrating FluidAudio into a Python app.

## Where the Apple-only bet becomes a problem

The ANE dependency is not an implementation detail you can swap out. The README says inference is offloaded to the Apple Neural Engine and that the SDK avoids GPU and MPS entirely, which means the performance story and the platform story are the same story. On Apple silicon you get the design the team optimized for. Anywhere else, there is nothing to fall back on.

That has concrete consequences. A cross-platform product cannot share its audio layer between an iOS client and a Windows or Android client; the related search interest in FluidVoice for Windows and Voiceink for Android reflects demand the SDK itself does not serve. A server-side transcription service is also out of scope, since the point is to keep audio on the device.

Language coverage is the second constraint, and it is narrower than the feature list suggests. Batch ASR covers 25 European languages plus Japanese, with Mandarin handled by separate SenseVoice and Paraformer models. Streaming ASR with end-of-utterance detection is English only. TTS covers nine languages for Kokoro and six for PocketTTS. If your product needs one language across every feature, the intersection is smaller than any individual row.

The third limitation is documentation depth rather than capability. The README is a feature and showcase document. It does not publish per-device latency figures, memory ceilings, or a rollback path for model updates. Those are the numbers you would normally use to size a decision, and their absence means you have to measure them yourself on your own hardware.

## FluidAudio against Whisper-based local transcription

The most common comparison for a local speech-to-text library is Whisper, and the related searches show people asking exactly that. The difference is architectural rather than a matter of accuracy claims. Whisper-derived local stacks typically run a single general-purpose model through a CPU or GPU inference runtime and expose one transcription path. FluidAudio splits the job across purpose-built models: Parakeet TDT v3 for batch, Parakeet EOU for streaming with end-of-utterance detection, Silero for voice activity, Sortformer for diarization with overlapping speech, Kokoro and PocketTTS for synthesis.

That split buys latency and power characteristics a single large model cannot match on a phone, and it is why the README can describe background and always-on workloads as the target. It costs flexibility. With Whisper-style tooling you can often swap in a different checkpoint or a fine-tune and keep the same pipeline. With FluidAudio you get the models the team converted, and the README directs anyone who wants a different model to möbius rather than offering an in-SDK conversion path.

A second real alternative is to skip the SDK and call the models directly through CoreML. That removes the dependency on FluidAudio's release cadence and its Swift API surface, at the cost of writing the preprocessing, chunking, streaming and clustering logic yourself. FluidAudio's value is concentrated in exactly that plumbing plus the model conversions, so the build-versus-buy line falls there.

## Release cadence, licensing and what an upgrade costs you

The repository is not archived, and the last push was on 2026-08-19, which matches the v0.15.6 release on the same date. The two releases before it were v0.15.5 on 2026-07-07 and v0.15.4, tagged japanese kokoro, on 2026-06-16. That is a steady minor-version cadence over the summer, with the version numbers staying in the 0.15.x line, so the project has not declared a stable 1.0 API.

For an app developer the practical cost of that cadence is model churn, not code churn. The v0.15.4 tag shows new language support arriving as a release event, which means a model file your app downloaded can be superseded. The README does not document a rollback path or a model-version pinning strategy, so if your app caches weights locally you should decide your own pinning policy rather than assume the SDK handles it.

On licensing, the SDK itself is Apache-2.0, and the README states the bundled models carry permissive MIT or Apache 2.0 licenses. The repository also contains a ThirdPartyLicenses directory, which is where the per-model attribution lives. That directory is worth reading before shipping, because the SDK license and the model licenses are separate obligations and the README's summary is not a substitute for the actual files. Nothing here is legal advice; if your product has compliance requirements, have counsel read ThirdPartyLicenses and the model cards on HuggingFace.

## Conclusion

Adopt FluidAudio if you are shipping a macOS or iOS app that needs local transcription, speaker separation or speech synthesis and you are willing to add Apple-only dependencies. Do not adopt it if you need Linux, Windows or Android support, or if your transcription must run server-side. Before committing, verify two things yourself: that the specific ASR model you need covers your language, since the README scopes Parakeet TDT v3 to 25 European languages plus Japanese, and that your target device generation runs the ANE path fast enough for your latency budget, because neither the README nor the release notes publish per-device timings.

## FAQ

### What is FluidAudio?

It is a Swift SDK for fully local audio AI on Apple devices, covering speech-to-text, text-to-speech, voice activity detection and speaker diarization. The README states that inference runs on the Apple Neural Engine using CoreML models converted by the Fluid Inference team.

### How does FluidAudio compare with Whisper for local transcription?

FluidAudio uses purpose-built models rather than one general-purpose checkpoint: Parakeet TDT v3 for batch transcription, Parakeet EOU for streaming with end-of-utterance detection, and Silero for voice activity. The trade-off is that you work with the models the project has converted, and the README points to möbius if you want to convert your own.

### Is FluidAudio good for music production?

No. FluidAudio is a speech SDK covering ASR, TTS, VAD and speaker diarization, and the README describes no music analysis, mixing or audio measurement features. Searches for Fluid Audio monitors and studio hardware refer to a different company.

### Is there a free API for converting audio to text with FluidAudio?

FluidAudio is not an API service. It is a Swift package that runs models locally on the device, and the README states the models carry permissive MIT or Apache 2.0 licenses. There is no hosted endpoint described in the repository.

## Sources

- [Official documentation](https://docs.fluidinference.com/introduction)
- [Official README](https://github.com/FluidInference/FluidAudio#readme)
- [Project repository](https://github.com/FluidInference/FluidAudio)
- [Release notes](https://github.com/FluidInference/FluidAudio/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/fluidinference-fluidaudio
