Model or dataset
CrispStrobe/CrispASR avatar
CrispStrobe/CrispASR

CrispASR: one C++ binary for Whisper, Parakeet, Canary and 116 more speech backends

C++ ggml runtime hub for multilingual ASR and TTS models: Cohere Transcribe, Parakeet TDT, Voxtral, Canary 1B v2, etc, plus universal forced alignment, and more

723 stars114 forksC++MIT

At a glance

What is it?
CrispASR is a ggml-based C++ runtime that puts multilingual ASR, 62 TTS engines and forced alignment behind a single CLI and HTTP server. It is a good fit if you want no Python in the deployment path and a poor fit if you need a stable API surface.
Who is it for?
Adopt CrispASR when you want several ASR and TTS architectures behind one binary and can pin a release, and skip it when you need a frozen API or a small auditable dependency tree. Before committing, run crispasr --version against the exact tarball you plan to ship, confirm the health endpoint on your chosen port, and check whether your backend of choice is in the capability tiers that tools/test-all-backends.py reports.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem CrispASR targets, and who ends up using it

Running open-weights speech models usually means one runtime per architecture. Whisper has whisper.cpp, Parakeet and Canary have their own toolchains, Voxtral arrives through a Python stack, and every new model drags in a fresh set of wheels. If you ship an on-premise transcription service, that collection of runtimes is the deployment. CrispASR attacks that directly: one C++ binary named crispasr, one consistent CLI, and the backend chosen either by flag or auto-detected from the GGUF file you pass to -m. The README describes the project as having started as a fork of whisper.cpp and then extended into what it calls a unified speech engine.

The audience is narrower than the backend count suggests. This is for people who already decided they want ggml-based inference and are willing to track a fast-moving project. The README lists CLI, HTTP server, a C ABI, and bindings for Python, Rust, Dart, Go, Ruby, Java and JavaScript, so the intended consumers are application teams embedding speech rather than researchers training models. The absence of a homepage and the presence of a HuggingFace Space demo and a Flutter companion app called CrisperWeaver point the same way: distribution through releases and bindings, not through a hosted service.

How the backend dispatch actually works

The README gives the dispatch rule plainly: pick the backend at the command line or let CrispASR auto-detect it from the GGUF file. In practice you either name the model file, as in the Parakeet, Canary and Voxtral examples, or you pass --backend together with -m auto, which the README says downloads the model on first use and reuses it afterwards. The first auto-download example is roughly 135 MB.

The repository is layered rather than monolithic. Alongside the crispasr/ directory there are separate top-level directories for crisp_audio/, crisp_lid/ (language identification), crisp_punc/ (punctuation), crisp_truecase/, and glint/, plus a vendored ggml submodule and a crispasr-sys/ crate for the Rust side. docs/architecture.md is described as covering the layered layout and the src/core/ primitives. That layout matters for anyone evaluating the project: the ASR backends share a core, but language ID, punctuation and truecasing are distinct components with their own code paths, which is why the README can offer them as add-ons instead of claiming one model does everything.

The server side is configured through environment variables using what the README calls the CRISPASR_<BACKEND>_<FEATURE> convention. The .env.example file shows the practical set: CRISPASR_MODEL, CRISPASR_BACKEND, CRISPASR_LANGUAGE, CRISPASR_PORT, CRISPASR_AUTO_DOWNLOAD, CRISPASR_CACHE_DIR, CRISPASR_API_KEYS and CRISPASR_EXTRA_ARGS. Note that CRISPASR_AUTO_DOWNLOAD defaults to 0, so a container that expects to fetch its own weights will not do so unless you change it.

Install and a first real transcription

The README's own instruction is to download one file from the Releases page and unzip it, with separate assets for Windows CPU, Windows CUDA, macOS and Linux. The Windows CPU build is documented as needing AVX2, with a -cpu-legacy variant for older CPUs. The README is explicit that the -hip and -vulkan builds require the matching GPU driver and do not fall back to CPU, while the Linux -cuda tarballs do fall back since v0.8.30. That distinction is worth reading twice before you pick a download.

First check that the binary starts. The README says this should print a version banner and exit, and that CUDA builds additionally print cuda toolkit and cuda runtime ABI, which is how you tell the CUDA 12 and CUDA 13 packages apart.

bash
crispasr --version

Then transcribe a file. The README's examples use samples/jfk.wav and a plain Whisper GGUF, with -f for the audio input.

bash
crispasr -m ggml-base.en.bin -f samples/jfk.wav

To avoid hunting for a model, the auto path names the backend and lets CrispASR fetch weights. The README shows qwen3 as the backend value here.

bash
crispasr --backend qwen3 -m auto -f samples/jfk.wav

The same binary does TTS. The README's Kokoro example passes the text with --tts and the destination with --tts-output.

bash
crispasr --backend kokoro -m auto --tts "Hello world" --tts-output out.wav

For a server deployment, the repository ships docker-compose.yml, which builds from .devops/main.Dockerfile, publishes ${CRISPASR_PORT:-8080} and mounts ./models and a crispasr-cache volume. Its healthcheck curls /health on that port, so that endpoint is the one to watch after startup.

bash
docker compose up -d

Where CrispASR is the wrong tool

The release history is the clearest warning. v0.8.30 is titled "audio that was wrong on the way in, and binaries that could not start", which tells you that a release can break both the audio path and the startup path at once. Three releases landed in the first week of September 2026, and the project's last push was on 2026-09-15. That cadence is healthy for a project tracking new model architectures, and it is exactly the cadence that makes an unpinned dependency painful. If your product needs a transcription API whose behaviour does not move between minor versions, this is not the layer to build on without a lockfile and a regression suite of your own.

The breadth is also a cost. A project carrying 119 backends, 62 of them TTS, plus translation and post-processing models, cannot give every backend equal attention. The README points to docs/regression-matrix.md and tools/test-all-backends.py as reporting capability tiers, which is an admission that backends differ in how completely they are covered. Before you standardise on, say, Voxtral, check what tier that backend sits in rather than assuming parity with Whisper.

Finally, the GPU story has sharp edges. The -hip and -vulkan builds do not fall back to CPU, so a machine without the matching driver fails rather than degrading. The README also maintains a troubleshooting page whose topics include the binary printing a banner and stopping, reading the exit code, and using a --no-gpu bisect. Those are the failure modes of a fast-moving native binary, and they are the reason a container with a pinned image tag is the safer default.

CrispASR versus whisper.cpp

The comparison is not abstract, because CrispASR began as a fork of whisper.cpp and the README says so. whisper.cpp is the narrower tool: it runs Whisper models, and its scope is the thing that makes it predictable. CrispASR keeps the same ggml foundation but adds other architectures behind the same binary, which is why a Parakeet, Canary or Voxtral invocation looks almost identical to the Whisper one in the README's examples.

The real difference is what you inherit. Choosing whisper.cpp means one model family and a small surface to audit. Choosing CrispASR means accepting the whole hub, including the TTS engines, the language-ID and punctuation components, the translation path and the bindings, whether or not you use them. In exchange you get one build and one CLI across model families, and you can switch architecture without changing your deployment. The repository also ships a COMPARISON.md at the top level, which is where the maintainers argue this case themselves; read it with the awareness that it is the project's own framing.

If your only requirement is Whisper transcription on CPU, the extra surface buys you nothing and costs you upgrade attention. If your requirement is to try Parakeet or Canary next quarter without a second runtime in the image, the fork's premise is the whole point.

Licence, upgrades and the cost of keeping up

CrispASR is MIT licensed, and the repository carries a THIRD_PARTY_NOTICES.txt alongside the LICENSE file. MIT is permissive, but the binary bundles ggml and, per the README, full runtimes for many model architectures, so the notices file is the place to look for what else is inside. Model weights are a separate question from the code: the README's auto-download path fetches GGUF files, and their licences come from whoever published the weights, not from this repository. That is a distinction worth resolving before shipping, and it is not legal advice, just the file to read.

The upgrade cost is real and it is mostly about verification. With releases arriving days apart, the practical discipline is to pin a version, keep the tarball or image tag you validated, and re-run your own audio through it after each bump. The README documents a benchmarking page that is careful to say it measures transcribe time and not cold start, and it mentions phase-timing environment variables plus proof-of-work rules. That framing is useful: if you benchmark CrispASR without separating model load from inference, you will misread your own numbers.

There is also a compliance angle the project has taken on itself. docs/eu-ai-act.md is described as covering synthetic-audio marking through watermark, C2PA and a spoken disclaimer, what counts as a voice clone, the speaker-biometrics boundary, why there is no emotion recognition, and what remains the deployer's duty. For anyone shipping the TTS side into the EU, that document is the starting point rather than a formality.

Editorial conclusion

Adopt CrispASR when you want several ASR and TTS architectures behind one binary and can pin a release, and skip it when you need a frozen API or a small auditable dependency tree. Before committing, run crispasr --version against the exact tarball you plan to ship, confirm the health endpoint on your chosen port, and check whether your backend of choice is in the capability tiers that tools/test-all-backends.py reports.

Frequently asked questions

How does CrispASR compare with whisper.cpp?

CrispASR started as a fork of whisper.cpp and extends that base into what the README calls a unified speech engine, so the same binary also runs Parakeet, Canary, Voxtral and others, plus 62 TTS engines. whisper.cpp stays scoped to Whisper models, which means a smaller surface to audit and fewer moving parts to upgrade.

Do I need Python or PyTorch to run CrispASR?

No. The README states there are no Python dependencies, no PyTorch and no pip install, just one C++ binary and a GGUF file. Language bindings for Python, Rust, Dart, Go, Ruby, Java and JavaScript are offered separately.

How do I install CrispASR on Linux or Windows?

Download one file from the Releases page and unzip it. The README lists crispasr-windows-x86_64-cpu.zip, crispasr-windows-x86_64-cuda.zip, crispasr-macos.tar.gz and crispasr-linux-x86_64.tar.gz, with -cuda and -vulkan Linux variants. The Windows CPU build needs AVX2, with a -cpu-legacy zip for older CPUs.

Does CrispASR need a GPU, and do the GPU builds fall back to CPU?

It runs on CPU, and the macOS build has Metal support built in. The README warns that the -hip and -vulkan builds require the matching GPU driver and do not fall back to CPU, while the Linux -cuda tarballs do fall back since v0.8.30.

Which port does the CrispASR server listen on?

Port 8080 by default. The docker-compose.yml file publishes ${CRISPASR_PORT:-8080} and its healthcheck curls /health on that port, and .env.example sets CRISPASR_PORT=8080.

Official sources

  1. CrispStrobe/CrispASR on GitHub
  2. Issues
  3. License: MIT
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/crispstrobe-crispasr.svg)](https://hysenlabs.com/projects/crispstrobe-crispasr)