# Rhino: Picovoice's On-Device Speech-to-Intent Engine

> Rhino is an on-device Speech-to-Intent engine from Picovoice that infers structured intent data directly from spoken commands, without sending audio to a cloud service. It returns a JSON object with the detected intent and named slot values and runs on microcontrollers, mobile devices, browsers, and desktop platforms.

**Picovoice/rhino** — On-device Speech-to-Intent engine powered by deep learning

- Repository: https://github.com/Picovoice/rhino
- Website: https://picovoice.ai/
- Stars: 711 · Forks: 94
- Language: Python
- License: Apache-2.0
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/picovoice-rhino

## What Rhino Solves and Who It Is For

General-purpose speech recognition turns audio into text. Rhino skips that step and goes directly from audio to structured intent data. Given a spoken command like "Can I have a small double-shot espresso?", Rhino emits:

```json
{
  "isUnderstood": "true",
  "intent": "orderBeverage",
  "slots": {
    "beverage": "espresso",
    "size": "small",
    "numberOfShots": "2"
  }
}
```

This approach targets embedded devices, IoT products, and mobile apps where a small known set of voice commands must be recognized quickly, offline, and without audio leaving the device. The README describes Rhino as compact and computationally efficient and notes it is suited for IoT use. It is distinct from Picovoice's Porcupine engine, which handles wake-word detection, and from the full Picovoice platform, which combines wake-word and speech-to-intent in sequence for Alexa-style experiences.

## Contexts, Expressions, and Slots

Rhino works within a domain called a Context. A context is a file that defines the set of intents the engine will recognize, along with the phrases that map to each intent and the slot types that extract variable details from those phrases.

An expression is a phrase the user might speak. Slots capture variable values within an expression. The README shows an example:

```yaml
turnLightOff:
  - Turn off the lights in the $location:lightLocation.
```

Here `$location:lightLocation` tells Rhino to expect a variable of type `location` and capture its value in `lightLocation`. The slot type itself is defined as a list of allowed phrases:

```yaml
lightLocation:
  - "attic"
  - "balcony"
  - "basement"
  - "bathroom"
  - "bedroom"
  - "entrance"
  - "kitchen"
  - "living room"
```

Custom contexts are created and downloaded using Picovoice Console at console.picovoice.ai. The context file is a compiled binary that the Rhino engine loads at runtime, not a YAML file deployed directly. The README points to a Speech-to-Intent Syntax Cheat Sheet for the full expression grammar.

## Language Support and Platform Coverage

Version 4.1 supports nine languages: English, Chinese (Mandarin), French, German, Italian, Japanese, Korean, Portuguese, and Spanish. Support for additional languages is listed as available for commercial customers on a case-by-case basis.

Rhino runs across a wide range of hardware. The README lists ARM Cortex-M, STM32, and Arduino for microcontrollers; Raspberry Pi for single-board computers; Android and iOS for mobile; Chrome, Safari, Firefox, and Edge for browsers; and Linux (x86_64), macOS (x86_64 and arm64), and Windows (x86_64 and arm64) for desktops.

SDKs are provided for Python, .NET, Java, Flutter, React Native, Android, iOS, vanilla JavaScript and HTML, React, Node.js, and C. Each platform has a corresponding demo directory in the repository under `demo/`.

## Getting Started with the Python Demo

To run Rhino on a microphone in Python, clone the repository first:

```console
git clone --recurse-submodules https://github.com/Picovoice/rhino.git
```

The `--recurse-submodules` flag is required because the repository uses submodules for platform libraries.

Install the Python demo package:

```console
sudo pip3 install pvrhinodemo
```

Then run the microphone demo, replacing the two placeholders with your AccessKey and the path to a context file:

```
rhino_demo_mic --access_key ${ACCESS_KEY} --context_path ${CONTEXT_PATH}
```

The AccessKey is a free credential obtained from Picovoice Console. Context files are either downloaded from the Console after designing your domain or pulled from the pre-built examples in the `resources/` directory of the repository.

For the .NET SDK, the demo is run from `demo/dotnet/RhinoDemo` with `dotnet run -c MicDemo.Release`. For Node.js, a similar pattern applies in the `demo/nodejs/` directory.

## Licensing Model and the Role of AccessKey

Rhino is licensed under Apache-2.0 for the SDK code. However, using the engine requires an AccessKey obtained from Picovoice Console. The free tier provides an AccessKey for personal and development use; commercial deployment requires a separate agreement with Picovoice.

This two-layer structure means the source code is open but the inference capability is gated. The pre-built model files and context compiler are not included in the repository under open terms. Developers evaluating Rhino for a commercial product need to contact Picovoice to understand pricing before committing to the integration.

The most recent release is v4.1, published on 2026-09-09. The previous major version was v4.0 from December 2025, and v3.0 was released in October 2023, suggesting a roughly annual cadence for major versions.

## Where Rhino Fits Against General ASR and Cloud Voice APIs

Google Cloud Speech-to-Text and Amazon Transcribe are the most commonly used alternatives for voice input in applications. Both transcribe speech to text in the cloud with broad vocabulary coverage. They are the wrong choice when the application must work offline, when audio privacy is a requirement, or when the target device lacks reliable internet connectivity.

Rhino differs from those services in two ways. First, it runs on-device with no network requirement. Second, it never produces free-form text: if the spoken command does not match any expression in the loaded context, Rhino returns `isUnderstood: false`. This binary behavior is a design constraint, not a limitation: it eliminates an entire class of parsing and validation logic that developers would otherwise need to write on top of raw transcript text.

For applications where users need to speak arbitrary sentences and receive text output, a general ASR service is the correct tool. For applications with a fixed command vocabulary that must work offline, Rhino's structured output and embedded runtime are a genuine advantage.

## Conclusion

Rhino is the right engine when voice interactions are scoped to a fixed, known domain and you need them to run on the device without a network call. It fits IoT products, voice-controlled kiosks, and mobile apps where latency and privacy matter. It is the wrong choice when users need open-ended free-form dictation or when the domain of commands is open and constantly changing, because Rhino only recognizes what its context file defines. Before starting, create a context at Picovoice Console, obtain a free AccessKey, and confirm that your target platform and language combination is supported in the current v4.1 release.

## FAQ

### Does Rhino require an internet connection to process voice commands?

No. Rhino processes audio entirely on the device without any network calls. It is designed for offline use cases including IoT devices, embedded hardware, and mobile apps where network access cannot be guaranteed.

### How do I create a custom context for Rhino?

Custom contexts are built using Picovoice Console at console.picovoice.ai. You define intents, expressions, and slot types in the Console's interface, then download the compiled context file to use with the Rhino SDK.

### What is the difference between Rhino and Picovoice Porcupine?

Porcupine is a wake-word detection engine that listens for a trigger phrase such as "Hey Siri". Rhino is a Speech-to-Intent engine that infers structured command data from what the user says after a trigger. The full Picovoice platform chains both engines together.

## Sources

- [License: Apache-2.0](https://github.com/Picovoice/rhino/blob/master/LICENSE)
- [Picovoice/rhino on GitHub](https://github.com/Picovoice/rhino)
- [Project website](https://picovoice.ai/)
- [README](https://github.com/Picovoice/rhino/blob/master/README.md)
- [Releases](https://github.com/Picovoice/rhino/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/picovoice-rhino
