# Picovoice Cheetah: an on-device streaming speech-to-text engine with a server-side licence check

> Cheetah runs speech recognition locally across Python, C, Java, .NET, Node.js, mobile and web bindings, but every deployment needs an AccessKey validated online. Here is how the Python package works, where it stops being the right choice, and what to confirm before you commit.

**Picovoice/cheetah** — On-device streaming speech-to-text engine powered by deep learning 

- Repository: https://github.com/Picovoice/cheetah
- Website: https://picovoice.ai/
- Stars: 670 · Forks: 77
- Language: Python
- License: Apache-2.0
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/picovoice-cheetah

## What Cheetah solves, and who actually needs streaming instead of file transcription

Most speech-to-text libraries take a finished audio file and return a transcript. Cheetah is built for the other shape of problem: audio that keeps arriving. The README describes it as an on-device streaming speech-to-text engine and lists streaming-speech-to-text among its topics, so the unit of work is a live stream rather than a completed recording. That matters for captions that must appear while someone is still talking, for voice commands that need a partial result before the sentence ends, and for anything where shipping audio to a server is unacceptable.

The README makes the privacy claim directly: all voice processing runs locally. For teams working under data-handling constraints, that removes the audio-upload step entirely. The engine ships bindings for Linux, macOS, Windows, Android, iOS, Chrome, Safari, Firefox, Edge and Raspberry Pi 3, 4 and 5, which is a wider platform list than most single-language ASR projects attempt. The primary language of the repository is Python, and the PyPI package is pvcheetah.

Who it is for: application developers who want transcription embedded in a client rather than behind an HTTP endpoint, and who are willing to accept a licensing dependency to get it. It is not aimed at researchers training or fine-tuning models, and it is not a transcription API you call over the network.

## The AccessKey: local inference, online authorisation

This is the design decision that shapes every deployment, and the README is unusually blunt about it. An AccessKey is described as the authentication and authorisation token for deploying Picovoice SDKs, and the documentation states that you need internet connectivity to validate it with Picovoice licence servers even though voice recognition runs fully offline.

So the offline story has a boundary. Audio never leaves the device. The key does, at least for validation. The same paragraph says the AccessKey verifies that usage stays within your account limits, and points to the Picovoice Console profile page for usage limits and real-time usage. Getting, renewing or adjusting those limits means contacting the enterprise sales team or an existing Picovoice contact.

For a desktop tool that runs on a machine with connectivity, this is a non-issue. For an air-gapped appliance, a kiosk that boots without a network, or a first-run experience on a device that has never been online, it is a hard blocker that no amount of local processing removes. Read the AccessKey section before you evaluate anything else, because it determines whether the rest of the feature list is reachable for your product.

## Language coverage and the commercial path to more

Cheetah Streaming Speech-to-Text currently supports English, French, German, Italian, Portuguese and Spanish, according to the README. That is six languages, and the list is explicit rather than open-ended. The next line states that support for additional languages is available for commercial customers on a case-by-case basis, with a link to Picovoice consulting.

Two consequences follow. First, if your product ships in one of those six languages, the default model is likely enough to start. Second, if it ships in anything else, you are not looking at a configuration change; you are looking at a commercial conversation. The README does not describe a self-service route to training or downloading an additional language model, and it does not publish per-language accuracy figures in the repository text. The accuracy claim is a link to a Picovoice benchmark page, which means the numbers live outside the repository and should be read there rather than assumed.

## Installing the Python package and running the first transcription

The README's Python path is the shortest route to hearing the engine work. There are two packages: pvcheetah is the SDK, and pvcheetahdemo is the demo. Install the demo package first.

```bash
pip3 install pvcheetahdemo
```

Then run the microphone demo, substituting your own AccessKey. The README writes the placeholder as ${ACCESS_KEY}, obtained from Picovoice Console.

```bash
cheetah_demo_mic --access_key ${ACCESS_KEY}
```

On a working setup the demo captures from the default microphone and prints transcription as you speak. If the key is invalid, expired or missing, expect the run to fail at initialisation rather than mid-stream, because validation happens up front. The README does not document a specific error message or exit code for that case.

If you would rather build the C demo from source, the repository expects submodules. Clone with --recurse-submodules, then configure and build with CMake, then run the binary with an access key, a model path and a library path.

```bash
git clone --recurse-submodules https://github.com/Picovoice/cheetah.git
cmake -S demo/c/ -B demo/c/build && cmake --build demo/c/build
./demo/c/build/cheetah_demo_mic -a ${ACCESS_KEY} -m ${MODEL_PATH} -l ${LIBRARY_PATH}
```

The README says to replace ${LIBRARY_PATH} with the appropriate library under lib, and ${MODEL_PATH} with the default model file at lib/common/cheetah_params.pv or a custom one. Those two paths are where most first builds go wrong, so check them before debugging anything else.

## Where Cheetah is the wrong tool

The AccessKey requirement is the first limitation, and it is architectural rather than cosmetic. A device that must transcribe before it has ever reached the network cannot complete key validation, so the local-processing benefit does not rescue that scenario.

The second is language. Six documented languages is a real constraint for international products, and the README routes everything else through commercial channels. If you need a language outside that set and cannot enter a commercial agreement, Cheetah is simply not the engine for you, regardless of how well the streaming model performs.

The third is scope. Cheetah is a streaming engine, not a batch transcription pipeline. Nothing in the README presents it as a tool for transcribing an archive of recordings, diarising speakers, or producing timestamped word-level output for subtitle authoring. Those are different products. If your workload is files on disk rather than a live microphone or audio callback, the streaming design is a mismatch, and the six-language list applies just as much to offline batch work.

Finally, the repository does not document rollback, model versioning policy, or what happens to a deployed application when an AccessKey is revoked. Those gaps should be treated as unknowns, not as things that probably work out.

## How it compares with Whisper-style offline models

The obvious alternative for local transcription is an open model such as Whisper, which you download and run yourself with no vendor account and no licence server in the loop. The difference is not accuracy claims; it is the operating model. A self-hosted model gives you unlimited offline operation and the freedom to fine-tune, in exchange for managing model weights, runtime dependencies and your own streaming wrapper, since Whisper is designed around chunks of audio rather than a continuous stream.

Cheetah inverts that trade. You get a maintained streaming engine with official bindings for Python, C, iOS, Android, Flutter, React Native, Node.js, Java, .NET and the web, and you give up the fully self-contained deployment because of the AccessKey check. The README's own framing supports this reading: it presents privacy as a property of where the audio is processed, not as a promise that the application never touches the network.

There is a middle position worth naming. If your only requirement is that audio not be uploaded, a self-hosted model satisfies it with no vendor relationship at all. If your requirement is a supported streaming API across many platforms with a commercial support path, Cheetah is the more direct fit. Pick based on which of those two you actually need, not on benchmark pages.

## Licence, maintenance and upgrade cost

The repository is licensed under Apache-2.0, and the README carries the standard Apache licence badge. That covers the source in this repository. It does not by itself describe the terms attached to the AccessKey, the hosted model files, or the commercial language arrangements, which are governed by Picovoice's own agreements rather than by the Apache grant. Treat the licence question as two questions: what the repository code permits, and what the service terms permit. This is not legal advice; read the actual terms before shipping.

The last push to the repository was on 2026-09-09, and the most recent release listed is 4.1.0 from 2026-06-24, following v4.0 in March 2026 and 3.0.0 in December 2025. That is a steady release cadence across roughly three quarters, and the repository is not archived.

Upgrade cost is where the version history deserves attention. The jump from 3.0.0 to v4.0 is a major version, which normally signals breaking changes in the binding APIs, and the README does not publish a migration guide in the text available. If you pin a version, expect to re-read the binding docs for your platform when you move across a major boundary. The model file is a separate artefact from the package, so a package upgrade and a model upgrade are two events to track, not one.

## Conclusion

Adopt Cheetah when you need continuous streaming transcription inside a desktop, mobile or browser process and can accept an AccessKey check against Picovoice licence servers. Skip it if your product must work with no internet connectivity at first launch, if you need a language outside the six documented ones without a commercial agreement, or if you need a fully permissive stack with no vendor relationship. Before writing code, confirm three things: that the AccessKey is valid and within your account limits, that the platform you target appears in the supported list, and that the model file path you pass to the C demo resolves to the default model under lib/common/cheetah_params.pv or to a custom model you have generated.

## FAQ

### How do I use Picovoice Cheetah from Python?

Install the demo package with pip3 install pvcheetahdemo, then run cheetah_demo_mic --access_key ${ACCESS_KEY}, replacing the placeholder with a key from Picovoice Console. The SDK itself is distributed as pvcheetah.

### How do I install Picovoice Cheetah?

For Python, the README installs the demo with pip3 install pvcheetahdemo. For C, clone the repository with --recurse-submodules and build the demo with cmake -S demo/c/ -B demo/c/build && cmake --build demo/c/build.

### Does Picovoice Cheetah work without an internet connection?

Voice processing runs locally, but the README states that you need internet connectivity to validate your AccessKey with Picovoice licence servers. Audio stays on the device; the key does not.

### Which languages does Picovoice Cheetah support?

Cheetah Streaming Speech-to-Text currently supports English, French, German, Italian, Portuguese and Spanish. The README says additional languages are available for commercial customers on a case-by-case basis.

### Which platforms can run Picovoice Cheetah?

The README lists Linux (x86_64), macOS (x86_64, arm64), Windows (x86_64, arm64), Android, iOS, Chrome, Safari, Firefox, Edge, and Raspberry Pi 3, 4 and 5. Bindings exist for Python, C, iOS, Android, Flutter, React Native, Node.js, Java, .NET and the web.

## Sources

- [License: Apache-2.0](https://github.com/Picovoice/cheetah/blob/master/LICENSE)
- [Picovoice/cheetah on GitHub](https://github.com/Picovoice/cheetah)
- [Project website](https://picovoice.ai/)
- [README](https://github.com/Picovoice/cheetah/blob/master/README.md)
- [Releases](https://github.com/Picovoice/cheetah/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/picovoice-cheetah
