whisper.rn: Running Whisper and Parakeet ASR Inside a React Native App
React Native binding of whisper.cpp.
At a glance
- What is it?
- whisper.rn is a TypeScript binding that puts whisper.cpp inference on the device, with Whisper and NVIDIA Parakeet TDT contexts, Silero VAD, and a realtime transcriber. Here is what the API actually does, what it refuses to decode, and who should stay away.
- Who is it for?
- Adopt whisper.rn if your React Native app must transcribe audio without sending it to a server, and you accept that the model file ships or downloads separately and that Parakeet input must be 16-bit PCM WAV. Do not adopt it if you need compressed audio decoded for you, or a managed transcription service.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 2 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What whisper.rn actually binds, and the app it fits
The package is a React Native binding of whisper.cpp, which is itself a C++ implementation of OpenAI's Whisper automatic speech recognition model. The README describes it as "React Native bindings for Whisper and NVIDIA Parakeet ASR through whisper.cpp." That sentence defines the scope: the heavy lifting happens in native code, and whisper.rn exposes it to JavaScript and TypeScript through a small number of context objects.
The problem it solves is deployment. A React Native app that wants speech recognition normally calls a cloud API, which means audio leaves the device, latency depends on a network round trip, and per-minute costs accumulate. whisper.rn moves inference into the app process. The repository layout supports this reading: there are android/, ios/, cpp/, cmake/, and vendor/ directories alongside src/, and the package.json files field ships src, lib, android, cpp, ios, cmake, and vendor while excluding build outputs. That is the shape of a project that compiles native code during the host app's build, not one that calls out to a server.
The audience is narrower than "React Native developers." It is developers who have already decided that on-device recognition is a requirement, usually because of privacy constraints, offline operation, or cost. If that decision is not already made, the cloud route is less work.
One thing the README does not do is give a supported-platform matrix. The screenshots mention testing on an iPhone 13 Pro Max and a Pixel 6, but the README does not state a minimum iOS or Android version. Treat that as unknown until you check the repository yourself.
The context object model: init, transcribe, release
The API is built around contexts. You initialize one, call transcribe on it, and release it when done. The README's first example shows the Whisper path: initWhisper takes a filePath pointing at a ggml model file, and the returned context exposes transcribe(audioPath, options), which returns an object with both a stop function and a promise. Awaiting the promise yields a result containing the inference text. The stop function exists so a caller can cancel an in-flight transcription, which matters on mobile where a user may navigate away mid-request.
The Parakeet path mirrors this. initParakeet takes a filePath and a useGpu flag, and its transcribe returns stop and promise, but the resolved value is richer: result, segments, and isAborted. The isAborted flag is how you distinguish a cancelled transcription from a completed one. The example also shows an explicit await parakeetContext.release() call, and the VAD section shows releaseAllWhisperVad() for releasing every VAD context at once. Memory is the reason these exist: a loaded model occupies real RAM, and nothing in the README suggests the binding frees it for you.
The data flow is consistent across all three contexts. A model file is loaded into native memory once. Audio is passed in as a file path, a require() asset, an HTTP URL, or a base64 string depending on the API. Native code runs inference and returns text plus timing information. The JavaScript layer never touches audio samples directly except when you hand it raw PCM.
That last detail is where the design gets opinionated, and the next section covers it.
Installing whisper.rn and getting one transcription out of it
Installation starts with npm. The README gives a single command for the JavaScript side.
npm install whisper.rnOn iOS, the README says to re-run npx pod-install. By default the package uses a pre-built rnwhisper.xcframework; if you want to compile from source instead, set RNWHISPER_BUILD_FROM_SOURCE to 1 in your Podfile. The README also notes that medium or large models on iOS make the Extended Virtual Addressing capability worth enabling, which is a hint about address space limits rather than memory bandwidth.
On Android, if ProGuard is enabled, the README asks for a keep rule in android/app/proguard-rules.pro.
# whisper.rn
-keep class com.rnwhisper.** { *; }Without that rule, the native classes the binding depends on can be stripped. The README also recommends ndkVersion = "24.0.8215888" or above for Apple Silicon Macs, and points to a troubleshooting document for the arm64 host CPU error that appears otherwise. For Expo, the README states that you must prebuild the project before using the library, and links to Expo's guide on using libraries in an Expo project.
Once installed, the shortest working path is the README's own example. You need a ggml model file on disk or in your bundle first.
import { initWhisper } from 'whisper.rn'
const whisperContext = await initWhisper({
filePath: 'file://.../ggml-tiny.en.bin',
})
const sampleFilePath = 'file://.../sample.wav'
const options = { language: 'en' }
const { stop, promise } = whisperContext.transcribe(sampleFilePath, options)
const { result } = await promiseWhat you should see is result holding the transcribed text. The filePath values are placeholders in the README itself; you supply a real path to a model you have downloaded. The README does not document where to obtain Whisper ggml models, only where to get the Parakeet GGUF files.
Parakeet TDT, and the audio format constraint that will bite you
Parakeet is the second engine, and it is not a Whisper variant. The README says ParakeetContext runs NVIDIA's Parakeet TDT 0.6B v3 model through the Parakeet API included in whisper.cpp, and that v3 supports English plus 24 other European languages. Models are GGUF files from ggml-org/parakeet-GGUF on Hugging Face, and the README lists four quantizations with approximate sizes: q4_0 at 356 MB, q4_k at 416 MB, q8_0 at 669 MB, and f16 at 1.26 GB. The example app downloads them at runtime rather than bundling them, which is the sensible choice given those numbers: a 1.26 GB asset inside an app binary is not something most teams will ship.
The constraint is in the input handling. The README states plainly that Parakeet file and base64 inputs must be WAV containing 16-bit PCM audio, and that transcribeData() accepts raw signed 16-bit PCM as a base64 string or an ArrayBuffer, with raw audio required to be mono at 16 kHz. Then comes the sentence that decides whether this engine is usable for you: "Compressed formats such as MP3, AAC, and FLAC are not decoded."
That is a hard boundary, not a caveat. If your app records from the device microphone, you control the format and can produce 16 kHz mono PCM. If your app receives user-uploaded MP3s or pulls audio from a media library, you need a decoding step somewhere before whisper.rn sees the file, and the README does not provide one for the Parakeet path. The VAD section does say detectSpeech supports the same formats as transcribe, and lists file paths, HTTP URLs, base64 WAV, and require() assets, but that is the VAD API, not Parakeet.
Parakeet also exposes maxThreads as a transcribe option in the example, which suggests the binding lets you tune CPU parallelism per call. The README does not explain how to choose a value.
VAD and the realtime transcriber: the parts that are documented least
Silero VAD is initialized with initWhisperVad, taking a model filePath, a useGpu flag that the README marks as iOS only, and nThreads. The returned context exposes detectSpeech, which accepts an audio source plus options, and detectSpeechData, which takes base64-encoded float32 PCM. Note the type mismatch: Parakeet's raw input is signed 16-bit PCM, while VAD's raw input is float32. If you are piping one into the other, you are converting.
The detectSpeech options are the most concretely documented part of the project. threshold is a speech probability from 0.0 to 1.0, minSpeechDurationMs defaults conceptually to 250 in the example, minSilenceDurationMs to 100, maxSpeechDurationS to 30, speechPadMs to 30, and samplesOverlap to 0.1. The example prints each segment's t0 and t1 in seconds. These are the knobs that determine whether a short interjection gets treated as speech or swallowed as silence, and the README presents them without guidance on tuning.
The RealtimeTranscriber section is where the documentation thins out. The README introduces it as providing "enhanced realtime transcription with features like Voice Activity Detection (VAD), auto-slicing, and memory management," and then the excerpt ends mid-sentence at a code comment. The package.json does show a docgen script that runs typedoc over src/index.ts and src/realtime-transcription/index.ts into docs/API, so generated API documentation exists in the repository even though the README does not carry it. If realtime is your use case, the API docs under docs/API are where to look, not the README.
Where whisper.rn is the wrong choice
The clearest failure mode is audio format. Anything that is not WAV with 16-bit PCM, for Parakeet, is out of scope. A podcast transcription feature that ingests MP3 uploads cannot use this binding without adding a decoder, and the README does not point at one.
The second is model size against device memory. The README recommends Extended Virtual Addressing on iOS specifically for medium and large Whisper models, which tells you that large models push against platform limits. It does not publish memory figures for any model, and it does not state which devices can run which sizes. The screenshots show tiny.en on an iPhone 13 Pro Max and a Pixel 6. That is the only device evidence in the README, and it is for the smallest model.
The third is build complexity. This package compiles native code. It has a ProGuard rule, an NDK version recommendation, a pod install step, a from-source toggle, and an Expo prebuild requirement. If your team does not already own an iOS and Android native build, every one of those is a place to get stuck, and the README defers several of them to a separate troubleshooting document.
Finally, if your audio is short, occasional, and already on a network, a cloud transcription API will be less work than this. The binding earns its complexity when inference has to happen locally, repeatedly, or on data that cannot leave the device.
Alternatives, and the real difference in approach
The obvious alternative is calling a hosted speech-to-text API from the same React Native app. The difference is not quality, it is where the model lives. A hosted API keeps the model on the provider's hardware, so the app binary stays small and there is nothing to compile, but every transcription is a network request and the audio leaves the device. whisper.rn inverts all three: the model is a file you supply, the binary carries native code, and the audio never leaves the process. For a note-taking app on a plane, that inversion is the whole point. For a call-center analytics tool with reliable connectivity and a data processing agreement, it is unnecessary work.
The second alternative is a different on-device runtime rather than a different binding. whisper.cpp itself is the engine underneath, and it can be embedded in a native app without React Native. Choosing whisper.rn over that is choosing to stay in the JavaScript layer: you get initWhisper, initParakeet, and initWhisperVad as TypeScript functions with typed options rather than JNI or Objective-C bridging code you write yourself. The cost is that you inherit this project's release cadence and its build integration. The repository has v0.7.4, v0.7.3, and v0.7.2 as recent releases, with the last push on 2026-09-14, so the project is being worked on, but you are still one layer removed from the upstream C++ library.
If neither trade appeals, the honest answer is that whisper.rn is not the tool. It is for teams that have already accepted on-device inference as a constraint.
Licence, maintenance and what an upgrade costs
whisper.rn is MIT licensed, and the repository carries a LICENSE file at the top level. MIT is permissive, so the binding itself imposes few obligations beyond preserving the notice. The part that needs your own attention is the model. The README points Parakeet users at ggml-org/parakeet-GGUF on Hugging Face and Whisper users at ggml model files, and the licences attached to those weights are separate from the binding's MIT licence. This is not legal advice; check the model card for whatever you ship.
On maintenance, the facts are narrow. The repository is not archived. The last push was on 2026-09-14, and the most recent release in the list is v0.7.4 from 2026-08-27. The version in package.json is 0.7.4, matching the release. The repository also contains AGENTS.md and CLAUDE.md at the top level, and .agents/ and .claude/ directories, which indicates agent tooling is part of the development setup rather than something shipped to consumers.
Upgrade cost is dominated by native rebuilds, not by API churn. Moving to a new version means re-running pod-install on iOS, rebuilding Android with a compatible NDK, and re-checking the ProGuard rule if the native package name changed. The exports map in package.json is fine-grained, with react-native, types, import, and require conditions plus per-path variants, so TypeScript consumers resolving through the types condition should get declarations from lib/typescript. The README does not document a migration guide or a breaking-change policy, so pinning a version and reading release notes before moving is the practical approach.
Editorial conclusion
Adopt whisper.rn if your React Native app must transcribe audio without sending it to a server, and you accept that the model file ships or downloads separately and that Parakeet input must be 16-bit PCM WAV. Do not adopt it if you need compressed audio decoded for you, or a managed transcription service. Before committing, verify on your own device that the chosen model size fits available memory, and check whether Expo's prebuild step and the iOS pod install fit your build pipeline.
Frequently asked questions
Is Whisper AI free to use?
The whisper.rn binding is MIT licensed and the repository includes a LICENSE file. The model weights are separate: Parakeet models come from ggml-org/parakeet-GGUF, and the README does not describe their terms.
Is there a real-time voice to text transcription service available?
The README introduces a RealtimeTranscriber with VAD, auto-slicing, and memory management, but the excerpt cuts off mid-example. The package.json includes a docgen script that generates API documentation into docs/API, which is where the full realtime API is documented.
What is Whisper talk?
The README describes whisper.rn as React Native bindings for Whisper and NVIDIA Parakeet ASR through whisper.cpp, which is a high-performance inference implementation of OpenAI's Whisper automatic speech recognition model. Transcription happens on the device through a context object created by initWhisper or initParakeet.
Is Whisper an AI?
The README calls Whisper an automatic speech recognition (ASR) model from OpenAI, and whisper.rn runs it on the device through whisper.cpp. Parakeet TDT 0.6B v3 is offered as a second model alongside it.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/mybigday-whisper-rn)