# expo-speech-recognition: One Speech Recognition API for iOS, Android and Web

> This Expo module wraps SFSpeechRecognizer, Android SpeechRecognizer and the Web Speech API behind a single TypeScript interface. It is a native module, so it needs a development build, not Expo Go.

**jamsch/expo-speech-recognition** — Speech Recognition for React Native Expo projects

- Repository: https://github.com/jamsch/expo-speech-recognition
- Stars: 689 · Forks: 54
- Language: TypeScript
- License: MIT
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/jamsch-expo-speech-recognition

## What expo-speech-recognition actually solves

React Native has no speech recognition primitive. On iOS you would write against SFSpeechRecognizer; on Android against android.speech.SpeechRecognizer; on web against the WICG SpeechRecognition interface. Three APIs, three permission models, three event shapes. expo-speech-recognition's stated goal is code reuse across web and mobile, and it does that by exposing one module and one set of events that map onto all three backends.

The audience is narrow and specific: teams already inside the Expo ecosystem who need dictation, voice commands or transcription in a shipped app. If your project is bare React Native without Expo, the module still installs as an Expo module, but you take on the Expo module toolchain for one capability. If you only need text to speech, this is the wrong package entirely; that is a different problem with a different API.

## How the native backends are wired together

The repository layout tells the story: src/ holds the TypeScript surface, android/ and ios/ hold the native implementations, and expo-module.config.json registers the module with Expo's autolinking. app.plugin.js is the config plugin entry point.

That plugin is doing real work, not decoration. The README states it updates the Android App Manifest to include package visibility filtering for com.google.android.googlequicksearchbox, Google's speech recognition package, plus the required permissions for both platforms. Android 11 and later hide installed packages from your app unless you declare them, so without that manifest entry the recognizer can be invisible to your process. The plugin accepts androidSpeechServicePackages when you need to query services beyond the ones already listed under forceQueryable, and the README points at adb shell dumpsys package queries as the way to discover what your device actually exposes.

Events flow one way: the native recognizer emits start, end, result and error, and the JavaScript layer re-emits them. useSpeechRecognitionEvent registers a listener per event name, which is why the hook is described as the easiest entry point. The module also exposes capability probes rather than assuming a uniform platform: isRecognitionAvailable(), supportsOnDeviceRecognition(), supportsRecording(), getSupportedLocales(), and Android-only calls such as getSpeechRecognitionServices() and getDefaultRecognitionService(). Those exist because the three backends genuinely differ, and the API admits it instead of pretending otherwise.

## Installing expo-speech-recognition and running a first transcription

Install from npm. The README also lists version tags for older SDKs, expo-speech-recognition@sdk-54 and expo-speech-recognition@sdk-53, for projects that have not moved to the current SDK.

```bash
npm install expo-speech-recognition
```

Next, register the config plugin. The README shows that it can be added with no configuration at all from v0.2.22 onward, or with an options object when you want custom permission strings or extra Android speech service packages.

```js
// app.json
{
  "expo": {
    "plugins": [
      "expo-speech-recognition",
      [
        "expo-speech-recognition",
        {
          "microphonePermission": "Allow $(PRODUCT_NAME) to use the microphone.",
          "speechRecognitionPermission": "Allow $(PRODUCT_NAME) to use speech recognition.",
          "androidSpeechServicePackages": ["com.google.android.googlequicksearchbox"]
        }
      ]
    ]
  }
}
```

This is the step where Expo Go users stop. The README has a note aimed at beginners: if you have only used the Expo SDK and have no ios/ or android/ directories, you need a development build. It gives the two commands that regenerate the native projects after the plugin is configured.

```bash
npx expo run:android
npx expo run:ios
```

The README repeats the Expo team's recommendation to keep ios/ and android/ in .gitignore and rely on continuous native generation, so the plugin stays the source of truth for the manifest changes.

With a development build in hand, the hook-based example is the shortest path to a working screen. It requests permissions, sets a recognizing flag on the start and end events, and writes the first result transcript into state.

```tsx
import {
  ExpoSpeechRecognitionModule,
  useSpeechRecognitionEvent,
} from "expo-speech-recognition";

function App() {
  const [recognizing, setRecognizing] = useState(false);
  const [transcript, setTranscript] = useState("");

  useSpeechRecognitionEvent("start", () => setRecognizing(true));
  useSpeechRecognitionEvent("end", () => setRecognizing(false));
  useSpeechRecognitionEvent("result", (event) => {
    setTranscript(event.results[0]?.transcript);
  });
  useSpeechRecognitionEvent("error", (event) => {
    console.log("error code:", event.error, "error message:", event.message);
  });

  const handleStart = async () => {
    const result = await ExpoSpeechRecognitionModule.requestPermissionsAsync();
    if (!result.granted) {
      console.warn("Permissions not granted", result);
      return;
    }
  };
}
```

What you should see: the start event fires when recognition begins, result events arrive with a transcript array, and end fires when the session closes. If permissions are denied, the granted flag is false and the README's example logs the result object rather than starting a session.

## Where the abstraction leaks

The README's own table of contents is the honest part of this project. It has separate troubleshooting sections for Android and iOS, a platform compatibility table, an Android-only mute-the-beep section, an Android-only offline model download call, and iOS-only audio session methods such as setCategoryIOS() and setAudioSessionActiveIOS(). A module that needed none of those would not document them.

The practical consequences are concrete. On-device recognition is a capability you must probe, not assume: supportsOnDeviceRecognition() exists precisely because it can be false. Offline Android use depends on androidTriggerOfflineModelDownload(), which means the model is not simply present. Recording audio alongside recognition depends on supportsRecording(). And the beep that Android plays on start is treated as a problem to be muted, which tells you the default behaviour is audible to users.

The README also carries a section on improving accuracy of single-word prompts, which is a hint about a real failure mode: short utterances are harder for these recognizers than sentences. If your interface is a one-word voice command, expect to tune.

The larger limitation is architectural. This is a native module. It cannot run in Expo Go, and the README spends an expandable note explaining that to people who have never run prebuild. If your team's workflow depends on Expo Go for iteration, adopting this module changes that workflow.

## When a cloud speech API is the better fit

The distinguishing choice here is that recognition happens through the operating system's own engine. iOS uses SFSpeechRecognizer, Android uses SpeechRecognizer, web uses the browser's SpeechRecognition. There is no vendor account, no API key, and no per-minute billing in the install path.

A cloud recognizer such as a hosted transcription service takes the opposite approach. You send audio to a server and get text back. That buys you consistent accuracy across devices, control over the model, and support for languages the device engine may not have. It costs you network dependency, audio leaving the device, and a per-request price. If your app must work in airplane mode on an arbitrary Android phone, the cloud route is worse. If your app must transcribe a language the device does not ship, the on-device route is worse.

There is a middle path worth noting because this project supports it: the README documents transcribing audio files and lists supported input audio formats. That means you can capture audio with this module and hand the file to whatever engine you prefer, rather than being locked into live recognition. The README does not describe a built-in cloud fallback, so if you want one, you build it.

## Maintenance, releases and the MIT licence

The last push to the repository was on 2026-08-30, and the most recent release, v57.0.0, was tagged the same day. Before that came v56.0.4 on 2026-08-28 and v56.0.3 on 2026-08-27. The cadence is tight, and the version numbers track the Expo SDK rather than the module's own history, which is why the README offers sdk-54 and sdk-53 install tags. Plan upgrades around Expo SDK upgrades, not around the module's changelog alone. The repository includes .changeset/, so release notes are generated from changeset entries in CHANGELOG.md.

The peer dependencies are expo, react and react-native, all wildcards. In practice that means the module expects the Expo version it was built against; the package.json devDependencies pin expo to ~56.0.12 while the published version is 57.0.0, so check the compatibility table in the README before pairing a module version with an SDK.

The licence is MIT. That permits commercial and closed-source use, and it comes with no warranty. This is not legal advice; if your organisation has rules about bundled native code or about shipping recognizers that route audio to a platform vendor, read the licence and the platform terms yourself.

## Conclusion

Adopt expo-speech-recognition if you already ship an Expo app with native directories or a development build and you need live dictation, file transcription or volume metering on iOS, Android and web from one API. Do not adopt it if you depend on Expo Go, if your build pipeline cannot run prebuild, or if you need a speech engine the platform does not provide. Before committing, verify three things on your own devices: that isRecognitionAvailable() returns true on your minimum Android version, that supportsOnDeviceRecognition() is true where you promise offline use, and that your chosen locales appear in getSupportedLocales().

## FAQ

### Is Expo or the React Native CLI better for a project like this?

The README does not compare the two. It does state that expo-speech-recognition is built as an Expo module with a config plugin, and that projects without ios/ or android/ directories need a development build created with npx expo run:android or npx expo run:ios.

### What is speech recognition and how does it work in expo-speech-recognition?

The module implements the platform recognizers rather than its own engine: iOS SFSpeechRecognizer, Android SpeechRecognizer and the web SpeechRecognition interface. The native side emits start, end, result and error events, which the JavaScript layer re-emits to listeners registered with useSpeechRecognitionEvent.

### What is the purpose of Expo in a project that uses expo-speech-recognition?

The README does not describe Expo's purpose in general terms. It does say the config plugin updates the Android App Manifest and adds permissions for Android and iOS, and that the module is consumed through the Expo module system, with expo listed as a peer dependency.

### Is Expo similar to React in a project like expo-speech-recognition?

The README does not compare Expo and React. It lists expo, react and react-native together as peer dependencies, and the usage example is a React component that calls useState alongside the useSpeechRecognitionEvent hook.

## Sources

- [Issues](https://github.com/jamsch/expo-speech-recognition/issues)
- [jamsch/expo-speech-recognition on GitHub](https://github.com/jamsch/expo-speech-recognition)
- [License: MIT](https://github.com/jamsch/expo-speech-recognition/blob/main/LICENSE)
- [README](https://github.com/jamsch/expo-speech-recognition/blob/main/README.md)
- [Releases](https://github.com/jamsch/expo-speech-recognition/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/jamsch-expo-speech-recognition
