Model or dataset
vilassn/whisper_android avatar
vilassn/whisper_android

whisper_android: Offline Speech Recognition with Whisper and TensorFlow Lite for Android

Offline Speech Recognition with OpenAI Whisper and TensorFlow Lite for Android

692 stars112 forksC++MIT

At a glance

What is it?
whisper_android provides two Android application projects that run OpenAI's Whisper speech recognition model locally using TensorFlow Lite, with no internet connection required. One app uses the TFLite Java API for straightforward integration, and the other uses the TFLite Native API in C++ for better inference performance.
Who is it for?
whisper_android is the right starting point for Android developers who need offline, on-device speech recognition using OpenAI Whisper and want a working TFLite integration they can adapt. The last push was on 2026-03-18, so teams adopting it should treat it as a reference implementation rather than an actively maintained library.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Offline Speech Recognition on Android Without Cloud Dependency

whisper_android addresses the constraint that standard Android speech recognition sends audio to Google servers for processing. That approach requires a network connection and sends user audio to a cloud service, which is unsuitable for offline use or for applications that must keep audio data on the device.

This repository provides an alternative: OpenAI Whisper converted to the TensorFlow Lite model format, running entirely on the Android device. The conversion preserves the multilingual speech recognition capabilities of Whisper while making it portable enough to run on mobile hardware without a GPU. The repository ships pre-built APKs in demo_and_apk/ for immediate testing and a Python script in models_and_scripts/ to generate TFLite models from Whisper model weights.

The two included apps demonstrate the same underlying approach through two different Android integration styles, allowing developers to choose the one that matches their project's existing architecture.

Two Apps: Java API vs Native C++ API

The whisper_java/ directory contains an Android app that performs model inference using the TensorFlow Lite Java API. This is the more accessible path for developers working in Java or Kotlin who want to integrate Whisper without writing C++ code. The README describes it as ideal for Java developers integrating TensorFlow Lite.

The whisper_native/ directory contains an Android app that uses the TensorFlow Lite Native API, which calls the TFLite runtime through C++ code via the Android NDK. The README states this offers optimized performance compared to the Java path.

The performance difference between the two depends on the device and the model size used. The Java API adds a layer of JNI bridging to the native runtime, while the Native API calls the runtime directly. For most use cases, the Java API is simpler to maintain and integrate. The Native API is appropriate when inference latency is a hard constraint.

Both apps share the same model file format: the .tflite files and the filters_vocab_multilingual.bin vocabulary file used for multilingual transcription.

Opening in Android Studio and Integrating the Whisper Class

There is no command-line build path for this repository. The README's quickstart is: navigate to the whisper_java/ or whisper_native/ folder, open the project in Android Studio, then build and run on an Android device or emulator.

Once built, the Whisper class is the main interface. Initialize it, load the model and vocabulary files, then configure it for transcription:

java
Whisper mWhisper = new Whisper(this);
String modelPath = "path/to/whisper-tiny.tflite";
String vocabPath = "path/to/filters_vocab_multilingual.bin";
mWhisper.loadModel(modelPath, vocabPath, true);

The third argument to loadModel sets multilingual mode. To transcribe a pre-recorded audio file, provide the path to a WAV file and start the transcription. The README specifies that the audio format must be 16 kHz, mono, 16 bits:

java
String waveFilePath = "path/to/your_audio_file.wav";
mWhisper.setFilePath(waveFilePath);
mWhisper.setAction(Whisper.ACTION_TRANSCRIBE);
mWhisper.start();

Results come back asynchronously via an IWhisperListener interface, which provides onUpdateReceived for status messages and onResultReceived for the transcription text. The Recorder class handles live audio capture, storing audio at 16 kHz mono and forwarding samples to the Whisper instance via writeBuffer() for live transcription.

Generating Custom TFLite Models from Whisper Weights

The models_and_scripts/ directory contains generate_model.py, a Python script for converting Whisper model weights into the TFLite format used by the apps. The directory also ships pre-generated TFLite models so that developers do not need to run the conversion themselves unless they want to use a different Whisper model variant.

The pre-generated models are stored under models_and_scripts/generated_model/. The model name whisper-tiny.tflite appears in the README as the example path, indicating the pre-built models use the tiny variant of Whisper.

The conversion step requires Python and a compatible Whisper installation. The README does not document the full conversion command, directing users to the script directly. Running the script produces the .tflite file and the vocabulary binary that the Android apps require.

Limitations and Maintenance Status

The last push to the repository was on 2026-03-18. The repository is not archived.

The repository has no published GitHub releases and no versioned artifacts. There is no CI pipeline documented in the README, no automated test suite, and no Gradle configuration shown other than what is inside the Android Studio project directories. Teams using this as a base for a production app will own the maintenance of TensorFlow Lite version compatibility and Whisper model compatibility going forward.

The Recorder class note in the README states that handling audio data and transcriptions requires careful synchronization and error handling, and flags this explicitly as a responsibility of the integrating developer rather than something the library handles.

The repository acknowledges that the TFLite implementation was originally authored by Niranjan Yadla (whisper.tflite project, 2022) and provides a BibTeX citation for that source.

Android SpeechRecognizer as the Standard Alternative

The standard Android SpeechRecognizer API, provided by the Android platform and backed by Google's speech recognition service, is the most commonly used alternative for Android speech-to-text. It supports multiple languages, handles diverse accents, and updates its model through Google's infrastructure. The fundamental difference from whisper_android is connectivity: SpeechRecognizer sends audio to Google servers and requires an active internet connection. Applications built on SpeechRecognizer cannot function offline, and audio is processed in the cloud rather than on the device.

Whisper_android's advantage is the on-device constraint: audio never leaves the device and the app works without network access. The trade-off is that the developer must bundle a TFLite model into the APK or download it on first launch, which adds to the app's storage footprint. The tiny Whisper model referenced in the README is the smallest variant and balances accuracy against size.

The MIT license permits commercial use without copyleft obligations, unlike some alternative on-device ASR implementations.

Editorial conclusion

whisper_android is the right starting point for Android developers who need offline, on-device speech recognition using OpenAI Whisper and want a working TFLite integration they can adapt. The last push was on 2026-03-18, so teams adopting it should treat it as a reference implementation rather than an actively maintained library. Before using it in a production app, verify that the pre-generated TFLite models in the repository still match the audio format expectations (16 kHz, mono, 16-bit) and test on your target Android API level.

Frequently asked questions

What audio format does whisper_android require for transcription?

The README specifies that audio files must be in 16 kHz, mono, 16-bit WAV format. The Recorder class captures audio in this format when recording on-device.

Does whisper_android require an internet connection?

No. The entire inference pipeline runs on-device using TensorFlow Lite. The Whisper model weights are converted to the TFLite format and bundled with or downloaded to the app, and no audio is sent to a remote server.

Which should I choose: the Java API app or the Native C++ app?

The README describes the Java API app as ideal for Java developers integrating TensorFlow Lite, and the Native API app as offering enhanced performance. Use the Java version for simpler integration and the Native version when minimizing inference latency is a priority.

Official sources

  1. Issues
  2. License: MIT
  3. README
  4. vilassn/whisper_android on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/vilassn-whisper-android.svg)](https://hysenlabs.com/projects/vilassn-whisper-android)