# sherpa-onnx: Offline Speech Processing for Embedded Systems and Mobile Apps

> sherpa-onnx is a C++ library built on next-gen Kaldi and ONNX Runtime that runs speech-to-text, text-to-speech, speaker diarization, VAD, and source separation entirely offline. It targets embedded hardware, mobile platforms, and edge devices, and exposes bindings for 12 programming languages.

**k2-fsa/sherpa-onnx** — Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, RK NPU, Axera NPU, Ascend NPU, x86_64 servers, websocket server/client, support 12 programming languages.

- Repository: https://github.com/k2-fsa/sherpa-onnx
- Website: https://k2-fsa.github.io/sherpa/onnx/index.html
- Stars: 14,966 · Forks: 1,719
- Language: C++
- License: Apache-2.0
- Published: 2026-08-08 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/k2-fsa-sherpa-onnx

## What sherpa-onnx Is and What Problem It Solves

sherpa-onnx provides a set of offline speech processing functions that run locally without an internet connection. The README lists the following capabilities: speech-to-text (both streaming and non-streaming), text-to-speech, speaker diarization, speaker identification, speaker verification, spoken language identification, audio tagging, voice activity detection (VAD), speech enhancement, keyword spotting, and source separation.

The core distinction from cloud-based speech APIs is that all processing happens on the device. There is no data transmitted to a remote server. This is essential for applications with strict data privacy requirements, for hardware deployed in environments without reliable connectivity, and for reducing latency to the point where cloud round-trip time is unacceptable.

The library is written in C++ and built on the ONNX Runtime for inference. The ONNX format decouples the model from the training framework, which means models trained in PyTorch, TensorFlow, or other frameworks can be exported to ONNX and used in sherpa-onnx without any Python dependency at runtime. The project targets the next-gen Kaldi ecosystem.

## Platform and Hardware Support

sherpa-onnx supports a wide range of platforms: Linux, macOS, Windows, Android, iOS, HarmonyOS, and openKylin. The supported CPU architectures include x86_64, x86, 32-bit and 64-bit ARM (arm64, aarch64), and RISC-V (riscv64). The README lists specific single-board computers: Raspberry Pi, NVIDIA Jetson Orin NX, NVIDIA Jetson Nano B01, LicheePi4A, VisionFive 2, and several Rockchip-based boards.

NPU acceleration is available for five chip families: Rockchip RKNN, Qualcomm QNN, Ascend NPU, Axera NPU, and Intel OpenVINO. Separate build scripts exist for each platform in the repository root: `build-android-arm64-v8a.sh`, `build-ios.sh`, `build-ohos-arm64-v8a.sh`, `build-rknn-linux-aarch64.sh`, and others.

The repository also includes WebAssembly support. Hugging Face Spaces are available for testing speech recognition, TTS, speaker diarization, audio tagging, and source separation in a browser without installing anything. There are separate WASM build scripts for different function subsets: `build-wasm-simd-asr.sh`, `build-wasm-simd-tts.sh`, `build-wasm-simd-vad.sh`, and others.

## Installing and Using sherpa-onnx with Python

For Python, sherpa-onnx is distributed as the `sherpa_onnx` package. The setup.py defines `package_name = "sherpa_onnx"` and sets `python_requires=">=3.7"`. Install with:

```bash
pip install sherpa_onnx
```

The setup.py also supports a CUDA variant. Setting the `SHERPA_ONNX_CMAKE_ARGS` environment variable to `-DSHERPA_ONNX_ENABLE_GPU=ON` before building from source appends `+cuda` to the package version string, indicating a GPU-enabled build.

Build scripts for specific targets are shell scripts in the repository root. For example, the Android ARM64 build:

```bash
build-android-arm64-v8a.sh
```

For Linux ARM64 (for Raspberry Pi and similar boards):

```bash
build-aarch64-linux-gnu.sh
```

The Python package installs command-line binaries alongside the library. These are placed in `build/sherpa_onnx/bin/` during the build and copied into the package. The setup.py locates them through the `get_binaries()` function.

## The 12 Programming Language Bindings

sherpa-onnx provides bindings for 12 languages: C++, C, Python, JavaScript, Java, C#, Kotlin, Swift, Go, Dart, Rust, and Pascal. It also supports WebAssembly. Each language has its own examples directory in the repository: `c-api-examples/`, `cxx-api-examples/`, `go-api-examples/`, `java-api-examples/`, `dart-api-examples/`, `dotnet-examples/`, and so on.

The Flutter framework is supported across Android (x64, x86, arm64, arm32), iOS (arm64), Windows, macOS, Linux, and Web. Pre-built Flutter demo apps are available from the `flutter` tag in the GitHub releases. The Tauri framework is also supported across Android, iOS, Windows, macOS, and Linux.

This breadth of language support is a practical advantage for teams building mobile or desktop applications. An iOS app can call the Swift binding directly, an Android app can use the Kotlin binding, and a cross-platform Flutter app can use the Dart binding. The C API sits as a common foundation that the higher-level bindings call into.

## Limitations and Cases Where sherpa-onnx Is the Wrong Choice

sherpa-onnx does not include its own model files. Users must obtain ONNX model files from Hugging Face or from the project documentation separately. The Hugging Face Spaces demonstrate what is possible, but they do not automatically provide a download path into your application. Finding, downloading, and integrating the correct model for a target language and task is the main setup burden.

The build process for non-Python targets is nontrivial. Compiling for Android, iOS, or a specific Linux ARM variant requires the appropriate cross-compilation toolchain and build dependencies. The repository provides build scripts for common targets, but using them requires familiarity with cross-compilation.

For streaming ASR, latency depends on the model architecture and the hardware. The README does not provide latency benchmarks. Teams integrating sherpa-onnx into a real-time transcription pipeline need to run their own measurements on their target hardware.

The Python package covers Python 3.7 and above, but the CMake build is the primary build path. Teams that want to integrate the C or C++ API directly need to manage the CMake build configuration, which adds complexity compared to a simple pip install.

## Comparison with Cloud Speech APIs

sherpa-onnx occupies a different position than cloud speech APIs from providers like Google, Azure, or Amazon. Cloud APIs are managed services: no model management, no hardware provisioning, and per-minute billing. sherpa-onnx requires you to manage model files and run inference on your own hardware, but charges nothing per call and sends no audio data off the device.

Within offline speech toolkits, Vosk is another offline ASR library with a comparable model ecosystem. The README for sherpa-onnx does not describe Vosk directly, but the practical difference is that sherpa-onnx covers a much wider function set (TTS, diarization, VAD, source separation, keyword spotting) while Vosk focuses primarily on ASR. sherpa-onnx also exposes more target platforms explicitly, including Flutter, Tauri, and WebAssembly.

For TTS specifically, piper is a widely used offline neural TTS tool. The README for sherpa-onnx does not compare against piper; the two serve overlapping use cases, but sherpa-onnx integrates TTS as one capability among many rather than as the primary focus.

## Release Cadence, NPU Ecosystem, and License

sherpa-onnx has a rapid release cadence. The three most recent releases were v1.13.7 on 2026-09-01, v1.13.8 on 2026-09-10, and a Tauri-specific release on 2026-08-25. The last push was on 2026-09-22. The project is not archived.

The NPU support has grown notably: Rockchip RKNN, Qualcomm QNN, Ascend, Axera, and Intel OpenVINO NPUs all have dedicated documentation pages linked from the README. Each has its own build script in the repository root. This coverage makes sherpa-onnx a viable option for teams deploying on edge inference chips rather than general-purpose CPUs or GPUs.

The license is Apache-2.0. Commercial use, modification, and redistribution are permitted. Derivative works do not need to be released as open source, and there is no copyleft requirement. The NOTICE file must be included in distributions.

## Conclusion

sherpa-onnx is the right choice for teams that need to run speech processing entirely offline on constrained hardware: embedded systems, mobile apps, Raspberry Pi, or edge servers with no internet dependency. The 12-language binding surface and the NPU acceleration options for Rockchip, Qualcomm, Ascend, and Axera chips make it distinctly suited for hardware-specific deployments. It is not the right tool for a team that wants a simple cloud API wrapper: the project requires choosing and managing model files from Hugging Face or the documentation rather than calling a hosted endpoint. Before adopting it, verify that a model for your target language is available and that your hardware platform is listed in the supported matrix. The license is Apache-2.0, which permits commercial use and redistribution without disclosure requirements.

## FAQ

### What is sherpa-onnx?

sherpa-onnx is a C++ library built on next-gen Kaldi and ONNX Runtime that performs speech-to-text, text-to-speech, speaker diarization, voice activity detection, keyword spotting, and source separation entirely offline. It supports Android, iOS, Raspberry Pi, and server platforms, and exposes bindings for 12 programming languages.

### How do I use sherpa-onnx?

For Python, install the sherpa_onnx package with pip install sherpa_onnx. You then need model files in ONNX format for your target task; these are available from Hugging Face. For mobile and embedded targets, use the platform-specific build scripts in the repository root to compile the C++ library for your architecture.

### How do I install sherpa-onnx?

The Python package installs with pip install sherpa_onnx and requires Python 3.7 or newer. For other languages or platforms, clone the repository and run the build script for your target: build-android-arm64-v8a.sh for Android ARM64, build-ios.sh for iOS, or build-aarch64-linux-gnu.sh for ARM64 Linux.

### What TTS models does sherpa-onnx support?

The README lists text-to-speech as a supported function and the Hugging Face Spaces include a speech synthesis demo. The README does not enumerate specific TTS model names, but the project documentation at k2-fsa.github.io/sherpa/onnx/index.html provides the model list. Models must be obtained in ONNX format separately.

## Sources

- [Official documentation](https://k2-fsa.github.io/sherpa/onnx/index.html)
- [Official README](https://github.com/k2-fsa/sherpa-onnx#readme)
- [Project repository](https://github.com/k2-fsa/sherpa-onnx)
- [Release notes](https://github.com/k2-fsa/sherpa-onnx/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/k2-fsa-sherpa-onnx
