Framework
Picovoice/speech-to-text-benchmark avatar
Picovoice/speech-to-text-benchmark

Picovoice's speech-to-text-benchmark: a Python harness for comparing STT engines on the same audio

speech to text benchmark framework

698 stars73 forksPythonApache-2.0

At a glance

What is it?
The repository is a Python framework that runs a fixed set of datasets through cloud and offline speech-to-text engines and reports WER, punctuation error rate, core-hours, latency and model size. It is a measurement tool, not an engine, and its value depends entirely on whether you trust the datasets and normalisation it ships.
Who is it for?
Adopt it if you need one harness that can put a cloud API and an offline model on the same audio and print the same metrics, and if you are willing to read normalizer.py and metric.py before trusting the numbers. Do not adopt it as a general STT leaderboard for models it does not wrap: the engine list is fixed in engine.py, and adding an engine means writing that adapter yourself.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 83 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What problem speech-to-text-benchmark actually solves

Engine vendors publish word error rates on their own terms. One reports on LibriSpeech test-clean, another on a private corpus, a third on a multilingual mix. The numbers land in marketing pages, not in a script you can rerun. This repository exists to remove that asymmetry: it fixes the datasets, fixes the metrics and fixes the invocation so that Amazon Transcribe, Azure Speech-to-Text, Google Speech-to-Text, IBM Watson, OpenAI Whisper, Whisper.cpp, Vosk, Moonshine, Picovoice Cheetah and Picovoice Leopard all get measured the same way.

The audience is narrow and specific. It is engineers choosing a recognition engine for a product, and researchers who need reproducible numbers rather than a vendor slide. It is not for someone who wants a hosted leaderboard they can query over HTTP. Everything here runs on your machine, against your copy of the audio, using your API keys.

The README describes it as "a minimalist and extensible framework for benchmarking different speech-to-text engines." Minimalist is accurate in both directions: the entry points are a handful of Python files, and the engine list is hard-coded, so extension means editing the source rather than registering a plugin.

Datasets, metrics and the invocation that ties them together

The data flow is linear. dataset.py loads a corpus, normalizer.py cleans both the reference transcript and the engine output, engine.py dispatches to the adapter for the chosen engine, metric.py computes the scores, and results.py writes them out. The supported datasets are named as constants: COMMON_VOICE, LIBRI_SPEECH_TEST_CLEAN, LIBRI_SPEECH_TEST_OTHER, TED_LIUM, MLS, VOX_POPULI and FLEURS. Languages are a separate fixed list: EN, FR, DE, ES, IT, PT_BR and PT_PT. IBM Watson is the exception, English only, which the README states outright.

Five metrics are reported. Word error rate is the edit distance between reference and hypothesis words divided by the number of reference words. Punctuation error rate follows Meister et al. and is reported for periods and question marks. Core-hour is CPU hours per hour of audio, and the README says it is omitted for cloud engines, which makes sense since you do not control their hardware. Word emission latency measures the delay between a word finishing and its transcription appearing, and is measured only for streaming engines. Model size is the combined acoustic and language model size in MB, omitted for cloud engines.

That split matters more than it first appears. A cloud engine and an offline engine cannot be compared on core-hour or model size, because the framework refuses to produce those numbers for the cloud side. If your decision hinges on on-device footprint, the comparison table will have empty cells for every API you test.

Installing it and running a first benchmark

The README states the benchmark was developed and tested on Ubuntu 22.04, and that FFmpeg must be installed first. Datasets are downloaded separately, and the FLEURS instructions live in script/README.md rather than the main file, so that one corpus needs a detour. Once the audio is on disk, the Python dependencies come from requirements.txt:

bash
pip3 install -r requirements.txt

That file pins cloud SDKs and offline runtimes side by side: amazon-transcribe, azure-cognitiveservices-speech, google-cloud-speech, ibm_watson, openai-whisper, pywhispercpp, vosk, moonshine-voice, pvcheetah and pvleopard, plus editdistance, inflect, matplotlib, numpy, psutil, pytube, requests and soundfile. Installing all of it pulls a large dependency tree even if you only intend to test one engine.

A minimal run against a local Whisper model looks like this, with the dataset folder pointing at your downloaded copy:

bash
python3 benchmark.py \
--engine WHISPER_TINY \
--dataset LIBRI_SPEECH_TEST_CLEAN \
--language EN \
--dataset-folder /path/to/librispeech

The engine value is the model selector here, not a separate flag. The README lists WHISPER_TINY, WHISPER_BASE, WHISPER_SMALL, WHISPER_MEDIUM, WHISPER_LARGE_V1, WHISPER_LARGE_V2, WHISPER_LARGE_V3 and WHISPER_LARGE_TURBO as valid values. For a cloud engine the shape changes, because credentials become required arguments:

bash
python3 benchmark.py \
--dataset COMMON_VOICE \
--dataset-folder /path/to/common_voice \
--language EN \
--engine AMAZON_TRANSCRIBE \
--aws-profile ${AWS_PROFILE} \
--aws-location ${AWS_LOCATION}

Expect the run to print per-utterance progress and finish with the aggregate metrics. If you want punctuation error rate, add --punctuation; the default punctuation set is .? and you can narrow or widen it with --punctuation-set, which accepts one or more of ., ? and ,. There is also a separate benchmark_latency.py entry point in the repository root, which is where the streaming latency measurement lives rather than in benchmark.py.

Where the harness stops being useful

The framework measures what it measures, and the gaps are real. It has no rollback or resume logic documented anywhere in the README, so a long Common Voice run that fails partway through is a run you restart. There is no caching layer described for engine output either, which means re-scoring the same audio against the same engine costs the same as the first time, and for a paid cloud API that is money spent twice.

Normalisation is the quiet variable. Word error rate is only as meaningful as the text preprocessing applied before the edit distance is computed, and that logic sits in normalizer.py, which the README does not explain. Two engines that differ in how they handle casing, numerals or filler words can produce a WER gap that has nothing to do with acoustic accuracy. Anyone quoting these numbers without reading that file is quoting a preprocessing choice as much as a model.

The engine list is hard-coded. If the engine you care about is not one of the ten, this repository cannot benchmark it without you writing an adapter in engine.py. And the dataset list is fixed too: your own production audio, with its accents, background noise and domain vocabulary, is not in it. A model that wins on LibriSpeech test-clean can lose badly on call-centre recordings, and this harness will not tell you that.

How it differs from hosted STT leaderboards

The obvious alternative is a hosted leaderboard such as the Hugging Face Open ASR Leaderboard, which the related searches suggest people look for. The difference in approach is who runs the audio. A hosted leaderboard evaluates submitted models on its own infrastructure against its own fixed test sets, which gives you broad model coverage and no setup cost, but it covers open models and not commercial APIs, and you cannot point it at your audio.

This repository inverts that. It runs on your machine, covers both cloud APIs and offline models in one table, and can be pointed at any of seven supported corpora. The cost is that you supply the compute, the credentials and the dataset downloads, and the model coverage is whatever engine.py happens to wrap. If you want to know how Whisper large-v3 compares to Amazon Transcribe on your own TED-LIUM subset, this is the tool. If you want a ranked list of forty open models, it is not.

Maintenance, licensing and the cost of upgrading

The repository is not archived, and its last push was on 2026-07-08. The Apache-2.0 licence covers the framework code in this repository. It does not cover what the framework pulls in: the cloud SDKs, the Whisper and Whisper.cpp weights, Vosk models, Moonshine, and the Picovoice Cheetah and Leopard packages all carry their own licences and terms, and the datasets have their own terms as well. Running a benchmark does not grant rights to redistribute the audio or the model weights, and nothing in the README addresses that. That is a question for whoever owns the data, not a legal opinion this article can give.

The upgrade cost is concentrated in requirements.txt. Every engine SDK is pinned to an exact version, so a fresh install is reproducible, but bumping one pin can change transcription output and therefore change your scores. If you compare a run from today against a run from six months ago, check the pins before concluding the model improved or regressed. The absence of published releases means there is no changelog to consult; the commit history on master is the record.

Editorial conclusion

Adopt it if you need one harness that can put a cloud API and an offline model on the same audio and print the same metrics, and if you are willing to read normalizer.py and metric.py before trusting the numbers. Do not adopt it as a general STT leaderboard for models it does not wrap: the engine list is fixed in engine.py, and adding an engine means writing that adapter yourself. Before you run anything, verify that your dataset layout matches what dataset.py expects, that the credentials for a cloud engine are reachable from the machine running benchmark.py, and whether the punctuation flag matters for your use case, because PER is only reported when you pass --punctuation.

Frequently asked questions

What is the best STT model in speech-to-text-benchmark?

The repository does not declare a winner. It reports word error rate, punctuation error rate, core-hour, word emission latency and model size per engine and per dataset, and the README's results section is where the numbers live. Which engine wins depends on the dataset and language you pick, and core-hour and model size are omitted for cloud engines, so no single ranking covers everything.

What is considered a good WER in speech-to-text-benchmark?

The README defines word error rate as the edit distance between reference and hypothesis words divided by the number of reference words, but it does not state a threshold for a good value. Any threshold depends on the dataset and on the normalisation applied in normalizer.py before the distance is computed.

Is ASR considered AI in the context of speech-to-text-benchmark?

The repository does not discuss that question. It benchmarks recognition engines and lists deep-learning and deep-neural-networks among its topics, but the README makes no claim about how ASR should be classified.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. Picovoice/speech-to-text-benchmark on GitHub
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/picovoice-speech-to-text-benchmark.svg)](https://hysenlabs.com/projects/picovoice-speech-to-text-benchmark)