Open-source project
nyrahealth/CrisperWhisper avatar
nyrahealth/CrisperWhisper

CrisperWhisper 2.0: a forked CTranslate2 backend with an install order that matters

Controllable Transcription. Verbatim ( every, filler, pause, stutter, vocal sound) , or intended ( what the speaker meant to say, optimized for readability) with word-level timestamps.

1,436 stars94 forksPythonNOASSERTION

At a glance

What is it?
CrisperWhisper turns the said versus meant choice into an explicit mode and reports word timings to a few tens of milliseconds. The cost is an installation with two backends, a one time conversion step, and a hard warning about faster-whisper sitting in the same environment.
Who is it for?
CrisperWhisper is worth evaluating on transcript quality rather than on the headline scores, and those scores come from the vendor's own benchmark.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 18 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

A core install pulls in no inference backend at all

The base package installs five libraries and nothing that can transcribe.

bash
# NVIDIA GPU (Linux): fastest, includes speculative decoding.
# An NVIDIA driver is all you need; CUDA libraries arrive via pip.
pip install "crisperwhisper[ct2]"

# Pure PyTorch: runs anywhere torch does (macOS, Windows, CPU)
pip install "crisperwhisper[transformers]"

The core dependency list is tokenizers, numpy, soundfile, soxr and huggingface_hub, and the project file states the intent plainly in a comment: the core install pulls in NO inference backend, and you choose one through an extra, ct2 for the fast CTranslate2 path, transformers for pure torch, or all for both. That split is why the two install lines carry different hardware stories. The ct2 route is documented for NVIDIA GPUs on Linux, with the CUDA libraries delivered through pip so that a driver is the only prerequisite, and it is the path that carries speculative decoding. The transformers route needs nothing beyond torch, so it runs on CPU, on macOS and on Windows. The package itself requires Python 3.10 or newer, while the published classifiers stop at 3.12, which leaves newer interpreters permitted by metadata and untested by the classifier list.

The ct2 extra pins a fork, and faster-whisper overwrites it

The fastest path does not use upstream CTranslate2, and the project file says so at length.

toml
    # which installs into the same site-packages/ctranslate2/ directory as the
    # fork and silently clobbers it (its FeatureExtractor is vendored in
    # crisperwhisper/features.py instead).

The comment above that block instructs anyone editing the extra not to add faster-whisper, and the reason is a filesystem collision rather than a version conflict: faster-whisper pulls upstream ctranslate2, both land in the same ctranslate2 directory under site-packages, and the upstream build overwrites the fork without an error. The mitigation is visible in the file layout, since the feature extractor is vendored into crisperwhisper/features.py instead of living inside the library. Two version facts follow from the same comments. The pin is at least 4.7.1.post2, described as the first fork build carrying generate_dual_greedy and transcribe_dual along with a native speculative loop, so an older cached copy of ctranslate2 in an environment will not do. And the fork wheels are compiled with CUDA dynamic loading on, dlopening libcublas.so.12 at runtime, so a separate nvidia cublas pip package is required and crisperwhisper/_nvidia_libs.py preloads that copy so GPU inference finds it. The visible copy of that comment ends mid sentence.

The first model load converts weights and needs two packages the fast path omits

Installing the fast path is not enough to run it once.

bash
pip install "crisperwhisper[ct2,convert]"   # or add [convert] later

On the CTranslate2 backend the first load of a model downloads it from HuggingFace and converts it once into a local cache, and that conversion step needs torch and transformers, which the lean ct2 extra deliberately does not install. The front page handles it as a documented gotcha rather than a failure: add the convert extra, either at install time or later. A CPU only torch is enough for the conversion, and both packages can be removed once the converted model is cached, which is the shape of a machine where you convert once and then serve from the local directory. Loading a directory that already contains a converted model, identified by a model.bin file, never needs them at all. The transcription API itself is short.

python
model = CrisperWhisperModel()          # nyralabs/CrisperWhisper2.0_large
# or pick a size: CrisperWhisperModel("turbo")  # turbo / medium / small

# Verbatim transcription (default): every filler, repetition, stutter,
# false start, and vocal event
result = model.transcribe("meeting.wav", language="en")
print(result.text)

Four size shorthands are offered, large by default, with turbo described as the fastest and recommended as the speculative draft, medium as the size and quality compromise, and small as the smallest.

The disfluency leaderboard mixes human labelled and synthetic sets

The headline ranking comes from a benchmark the model author publishes, and its own footnote splits the languages in two.

| # | System | Disfluency F1 | |--:|--------|--------------:| | 1 | CrisperWhisper 2.0 Pro | 93.5 | | 2 | CrisperWhisper 2.0 | 87.8 | | 3 | ElevenLabs Scribe v2 | 79.2 | | 4 | Microsoft MAI-Transcribe-1.5 | 77.5 | | 5 | CrisperWhisper 1.0* | 64.8 | | 6 | Inworld STT | 59.5 | | 7 | xAI Grok Speech-to-Text | 42.8 | | 8 | Deepgram Nova-3 | 37.8 | | 9 | Fish Audio ASR | 35.0 | | 10 | AssemblyAI Universal-3 Pro | 30.5 |

The metric counts fillers, repetitions, cut offs and vocal sounds as separate typed events, so a system is rewarded for writing down disfluencies that were spoken and penalised for inventing ones that were not. Two details in the footnote matter when reading the average. English and German are scored on human labelled evaluation sets, while the other eight languages use synthetic verbatim sets, so a ten language average mixes two kinds of ground truth. And the previous version of this model appears at number five with a stated caveat that it covers English and German only, so its average is computed over two languages rather than ten. The claim that this system leads on disfluency F1 across ten languages ahead of the closed source alternatives is made by the same party that runs the benchmark.

Boundary error is measured only where a system got the word right

The word timing table carries a scoring rule that changes how to read it.

| # | System | Boundary error | |--:|--------|---------------:| | 1 | CrisperWhisper 2.0 | 29.6 ms | | 2 | xAI Grok Speech-to-Text | 37.1 ms | | 3 | CTC-seg | 49.3 ms | | 4 | ElevenLabs Scribe v2 | 51.3 ms | | 5 | NeMo-FA | 60.0 ms | | 6 | Deepgram Nova-3 | 63.3 ms | | 7 | WhisperX | 64.8 ms | | 8 | Cartesia Ink-Whisper | 69.4 ms | | 9 | Canary | 85.5 ms |

The measurement is mean absolute word boundary error on read speech, with TIMIT named as the corpus, and the note under it states the rule: scored on exactly the words each system gets right. That restriction is the important part, because a system that misses or miswrites a word is not charged timing error for it, so the ranking compares alignment quality on each system's own correct output rather than end to end usability on shared audio. The timings themselves are said to come from supervised cross attention, and a separate post is given for how that extraction works and for conversational speech results, which are summarised in the opening as about 30 ms mean boundary error on read speech and 41 ms on conversational speech. Both the benchmark and the aligner write up are hosted by the same lab that builds the model.

Verbatimize reconstructs disfluencies from a clean transcript you already have

The most interesting entry in the box does not transcribe from scratch.

code
[um] so we we need to, to reschedule the th- thursday meeting to [uh] march third at nine thirty [laughter]

Verbatimize takes audio plus a transcript you already trust, reproduces your content word for word, and then inserts only the disfluencies and vocal events that are actually present in the audio. The stated effect is a jump in rare word recall from 6.8 percent to 96.1 percent compared with re-transcribing from audio, on the reasoning that the expensive clean transcripts already exist in quantity and only the disfluency layer is missing. The proposed uses are text to speech data, clinical speech analysis and dataset construction, all three of which need disfluency annotation that hand labelling does not scale to. The same design shows up in the two transcript modes. Verbatim is the default and keeps every filler, repetition, stutter, false start and vocal event in a consistent bracketed format, while intended returns what the speaker meant with numbers, dates and emails written the way a person would type them, so the spoken march third at nine thirty above comes back as March 3 at 9:30. Both modes are described as working across most languages Whisper supports.

Longform continues conditionally, and its link points at a different heading

Audio of any length is a stated feature, and the mechanism is described in one sentence.

Each window continues from the words already transcribed, which the page calls conditional continuation, so there are no duplicated or dropped words at the seams and no fragile timestamp token bookkeeping to maintain. The consequence is that the 30 second window that Whisper derived systems inherit is no longer a hard boundary: the front page says audio longer than 30 seconds is handled automatically. The same runtime also claims built in mitigation of Whisper's looping hallucination failure mode, where a decoder repeats itself after a pause, and that mitigation is tied to the CTranslate2 path rather than to the pure torch one. Two small documentation facts sit next to the feature list. The quickstart snippet ends on a lone f, the start of a word timestamp loop that never appears. And the longform link in that same paragraph targets an anchor named what else is in the box, so a section heading has been renamed since the reference was written. The visible release history gives v2.0.3 as a transformers only install fix that adds word timestamps for languages written without spaces, v2.0.1 as early end of sequence recovery, and v2.0.0 as the 2.0 release.

MIT in the package metadata, a research license on the weights

Three licensing statements sit next to each other and they do not agree.

The repository carries a LICENSE file at the root, and the project metadata declares license = {text = "MIT"} along with the OSI approved MIT classifier, so an installer reading the package sees an MIT library. The model weights are a separate artifact with their own license file on the model host, and the standard, non Pro checkpoints are released under a non-commercial research license, with commercial licensing offered separately. The Pro variants are described as the best models, with improved performance, hotword boosting and training on additional proprietary data. So the code path and the weights you download have different terms, and the proprietary data in the Pro line is a further layer on top. Worth noting for anyone automating installs: the license field reported for this repository at the hosting level is not asserted, which does not match the MIT declaration inside pyproject, so the two sources of licence information disagree before you even reach the weights. The front page sentence about where Pro models are available stops after the words are available, so the terms for that line are not stated on the page.

Editorial conclusion

CrisperWhisper is worth evaluating on transcript quality rather than on the headline scores, and those scores come from the vendor's own benchmark. Check three things first: whether you need the fast CTranslate2 path at all, since the pure PyTorch extra runs on CPU and macOS and needs no fork; whether your environment already contains faster-whisper, which the project warns will overwrite the pinned CTranslate2 directory; and which license applies to the weights, since the package metadata says MIT while the standard models are published under a non-commercial research license and the Pro weights are separate. Whoever reports timings from it should know the boundary error is measured only on the words each system transcribed correctly.

Frequently asked questions

Does CrisperWhisper need an NVIDIA GPU?

No. The pure PyTorch extra runs anywhere torch does, including macOS, Windows and CPU, while the ct2 extra is the fast path documented for NVIDIA GPUs on Linux, where the CUDA libraries arrive through pip and a driver is all you need. Speculative decoding is only mentioned on the ct2 path.

What is the difference between verbatim and intended mode in CrisperWhisper?

Verbatim is the default and keeps fillers, repetitions, stutters, false starts and vocal events in a bracketed format such as [um] and [laughter]. Intended returns the readable version with numbers, dates and emails written as a person would type them, so spoken march third at nine thirty comes back as March 3 at 9:30.

Can I use CrisperWhisper models commercially?

The code and the weights are licensed separately. The package metadata declares MIT and the repository carries a LICENSE file, the standard model weights are released under a non-commercial research license with commercial licensing offered separately, and the Pro models are described as trained on additional proprietary data with hotword boosting.

How long can audio be for CrisperWhisper?

Any length. Audio past 30 seconds is handled automatically, and the longform path uses conditional continuation, where each window continues from the words already transcribed, so seams do not duplicate or drop words and there is no timestamp token bookkeeping.

Why does my CrisperWhisper install break when faster-whisper is present?

The ct2 extra pins a fork of CTranslate2 rather than upstream, and the project file warns against adding faster-whisper, because that package installs upstream ctranslate2 into the same ctranslate2 directory under site-packages and overwrites the fork without an error. The feature extractor is vendored into crisperwhisper/features.py to avoid the clash.

Official sources

  1. Issues
  2. nyrahealth/CrisperWhisper on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/nyrahealth-crisperwhisper.svg)](https://hysenlabs.com/projects/nyrahealth-crisperwhisper)