Picovoice Leopard: on-device speech-to-text without a server
On-device speech-to-text engine powered by deep learning
At a glance
- What is it?
- Leopard is a local speech-to-text engine with bindings for Python, Node, Java, .NET, Flutter and the web. It runs offline, but it needs an AccessKey validated over the network and a paid plan for real volume.
- Who is it for?
- Adopt Leopard when audio cannot leave the device and you accept a commercial licence tied to a Picovoice AccessKey. Skip it if you need an open model you can retrain, or if your budget is zero: the trial ends and usage limits are enforced server-side.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 21 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Leopard solves, and who is buying it
Leopard targets a specific constraint: audio that must not leave the machine. The README states that all voice processing runs locally, so a transcription job does not ship audio to a cloud API. That matters for medical notes, legal recordings, field devices and anything with a data residency rule. It also matters on hardware with no reliable uplink, such as a Raspberry Pi 3, 4 or 5, which the README lists as supported alongside Linux x86_64, macOS on x86_64 and arm64, Windows on x86_64 and arm64, Android, iOS, and Chrome, Safari, Firefox and Edge.
The intended user is an application developer, not a data scientist. Leopard is a finished engine with a fixed model, not a training framework. You pick a language, load a library, and call a transcribe function. The README lists English, French, German, Italian, Japanese, Korean, Portuguese and Spanish, and notes that additional languages are handled case by case for commercial customers. If your language is outside that set, the project is not a candidate until you talk to sales.
That framing also sets the price of entry. The AccessKey section is explicit: anyone using Picovoice needs a valid AccessKey, it must be kept secret, and internet connectivity is required to validate it against Picovoice license servers even though recognition itself runs offline. The same section says the AccessKey verifies that usage stays within your account limits, which you can see on the Picovoice Console profile, and that continuing after the trial means contacting the Enterprise Sales team. So the honest description is offline inference, online authorization.
The engine, the model file, and the AccessKey handshake
The repository layout shows a thin binding layer over a native core. Top-level entries include include/, lib/ and resources/, and the C demo takes a library path plus a model file path, pointing at lib/common/leopard_params.pv as the default model. The bindings under binding/ wrap that core for Python, Node.js, Java, .NET, iOS, Android, Flutter, React Native and the web. The Python package is published as pvleopard, the Node packages as @picovoice/leopard-node, leopard-react, leopard-react-native and leopard-web, the Android and Java artifacts as ai.picovoice:leopard-android and ai.picovoice:leopard-java, the .NET package as Leopard, iOS as the Leopard-iOS CocoaPod, and Flutter as leopard_flutter.
The data flow is short. You hand the engine an audio file path and an AccessKey. The binding loads the native library and the parameters file, the engine decodes the audio and returns a transcript, and the AccessKey is checked against Picovoice servers. There is no queue, no daemon and no streaming socket in the documented interface; the demos are file-oriented, which is why the sample apps include python-subtitle and python-youtube variants that feed media files in.
The model file is the interesting constraint. Because the parameters live in a file you pass explicitly, the C demo can be pointed at a custom model rather than the default. The README does not explain how such a model is produced or what format it uses. Treat that as a commercial engagement detail rather than a documented workflow.
Install and transcribe a file with the Python demo
The fastest way to confirm Leopard works on your machine is the Python demo package, which the README installs with pip3. It pulls the demo entry point and the underlying engine, so you do not need to build anything from source.
pip3 install pvleoparddemoNext, run the file demo. The README shows the command with two placeholders: your AccessKey from the Picovoice Console and the path to the audio you want transcribed. The flags are --access_key and --audio_paths, and the latter accepts a path, so you can pass several files in one call.
leopard_demo_file --access_key ${ACCESS_KEY} --audio_paths ${AUDIO_FILE_PATH}Replace ${ACCESS_KEY} with the key from your console account and ${AUDIO_FILE_PATH} with a real file. On a successful run the demo prints the transcript for each file. The first invocation needs network access so the key can be validated; after that the recognition step itself is local. The README does not document accepted sample rates or channel layouts for the demo input, so if a file produces nothing useful, test with a plain mono recording before assuming the engine is at fault.
For a Node.js equivalent, the README installs @picovoice/leopard-node-demo globally with yarn and runs leopard-file-demo with --access_key and --input_audio_file_path. The C route is heavier: cmake -S demo/c/ -B demo/c/build builds it, and the binary takes -a, -l and -m for access key, library path and model path, followed by the audio file.
Where Leopard is the wrong choice
The AccessKey requirement cuts against the privacy pitch in one specific way. Recognition is local, but authorization is not. An air-gapped machine, a factory floor with no egress, or a device behind a restrictive firewall will fail the key check regardless of how good the model is. If your deployment cannot reach Picovoice license servers, Leopard is not usable, and no amount of local processing changes that.
The second limit is the model itself. Leopard ships a fixed, pre-trained engine. The README does not describe fine-tuning on your own audio, adapting to a domain vocabulary, or exporting the model to run under a different runtime. Teams with unusual accents, heavy jargon, or a need to own the weights will find the boundary quickly. The custom model path in the C demo exists, but the documentation stops at passing the file.
Third, language coverage is eight languages. Anything else is a commercial conversation. And the README does not document rollback, offline key caching duration, or what happens to a running deployment when a key expires or a usage limit is hit. Those are real operational questions, and the repository is silent on them.
How Leopard differs from Whisper-based stacks
The obvious alternative for local transcription is an OpenAI Whisper model run through a wrapper such as whisper.cpp or faster-whisper. The difference in approach is ownership. Whisper weights are downloadable and the runtime is yours to modify, quantize and embed; you can run it with no account, no key check and no network at any point. Leopard gives you a packaged engine with a support path and a commercial licence instead.
The trade-off is practical rather than ideological. A Whisper deployment puts the burden on you: choosing a model size, managing memory, and accepting whatever accuracy and latency your hardware delivers. Leopard hands you a single parameters file and a binding, and the README points at Picovoice's own speech-to-text benchmark for accuracy and real-time-factor figures rather than publishing numbers in the repository. If you want to measure before you commit, run both on your own audio; the benchmark page is a vendor claim, not a neutral test.
There is also a middle option worth naming. If your product is a wake word plus a command, not full transcription, Picovoice's own Porcupine and Rhino occupy that slot, and the related searches around Cheetah, Jaguar and other Picovoice engines suggest people conflate the family. Leopard is the batch transcription member of that set.
Licence, maintenance and the cost of upgrading
The repository is Apache-2.0, and the LICENSE file sits at the top level. That covers the source in this repository, but the Apache grant does not make the service free. The AccessKey section ties usage to account limits and directs anyone past the trial to Enterprise Sales. Read the licence and the commercial terms as two separate documents, because they are.
The code is current: the last push was on 2026-09-09, and v3.0 was released on 2025-12-18, after v2.0 in December 2023 and v1.2 in March 2023. That cadence is slow but not stalled, and the major version jump from 2.0 to 3.0 is the upgrade event to plan for. The README does not include a migration guide or changelog in the text available here, so pin your binding version and read the release notes before moving a production service from v2.0 to v3.0. Each binding ships on its own registry, which means a Python upgrade and a Node upgrade are separate decisions with separate version numbers.
Editorial conclusion
Adopt Leopard when audio cannot leave the device and you accept a commercial licence tied to a Picovoice AccessKey. Skip it if you need an open model you can retrain, or if your budget is zero: the trial ends and usage limits are enforced server-side. Before committing, verify the account limits on the Picovoice Console profile page, confirm that your target language is one of the eight supported, and check that your audio is a format the demos accept, since the README does not document a resampling step.
Frequently asked questions
How do I use Picovoice Leopard to transcribe an audio file?
Install the demo with pip3 install pvleoparddemo, then run leopard_demo_file with --access_key set to your Picovoice Console key and --audio_paths set to the file you want transcribed. The demo prints the transcript for each file passed.
How do I install Picovoice Leopard on macOS?
macOS on x86_64 and arm64 is a supported platform, and the Python route is pip3 install pvleoparddemo followed by the leopard_demo_file command. For a native build you would use the C demo, which the README builds with cmake -S demo/c/ -B demo/c/build, and the iOS demo uses pod install.
Does Picovoice Leopard need an internet connection?
Recognition runs fully offline, but the README states that internet connectivity is required to validate your AccessKey with Picovoice license servers. A device that cannot reach those servers cannot authorize the engine.
Which languages does Picovoice Leopard support?
The README lists English, French, German, Italian, Japanese, Korean, Portuguese and Spanish. Additional languages are handled on a case-by-case basis for commercial customers.
Is Picovoice Leopard free to use commercially?
The repository is Apache-2.0, but the README says anyone using Picovoice needs a valid AccessKey, usage is checked against your account limits, and continuing after the trial means contacting the Enterprise Sales team. The licence and the commercial terms are separate.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/picovoice-leopard)