parakeet.cpp: NVIDIA Parakeet speech recognition in C++, faster than the PyTorch original
Parakeet implementation in C++ with ggml
At a glance
- What is it?
- The LocalAI team ported every Parakeet ASR family to ggml, validated all checkpoints at WER 0 against NeMo, added cache-aware streaming with end-of-utterance detection, and shipped binaries that beat whisper.cpp on speed.
- Who is it for?
- parakeet.cpp is what a good model port looks like: pick a family with strong accuracy, port it to ggml, hold yourself to byte-identical transcripts against the reference, publish the benchmark methodology, and ship pre-built binaries.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
A C++17 port with a strict parity bar
parakeet.cpp comes from the LocalAI team, the group behind the open-source inference engine that runs LLMs, vision, voice, and image models on consumer hardware. The project is a C++17 implementation of NVIDIA's NeMo Parakeet speech recognition models built on ggml, the same backend family that powers llama.cpp and whisper.cpp. The pitch is concrete: fast, dependency-light automatic speech recognition on CPU or GPU with no Python runtime at inference time.
What makes the project credible rather than just another port is the parity bar it sets for itself. Every published checkpoint is validated at WER 0 against NeMo, meaning the transcripts come out byte for byte identical to the original PyTorch runtime. Coverage spans all the offline Parakeet families, CTC, RNNT, TDT, and the hybrid TDT-CTC, in 0.6B, 1.1B, and 110M sizes, English plus a multilingual v3, and a full parity matrix per model lives in docs/parity.md.
Streaming with end-of-utterance detection
Offline transcription is the easy half of ASR. The harder half is streaming, and parakeet.cpp covers it in two models. The parakeet_realtime_eou_120m-v1 checkpoint implements cache-aware streaming with end-of-utterance detection through a --stream flag, and the streaming transcript matches NeMo's own cache-aware streaming byte for byte.
The second streaming model is nvidia/nemotron-3.5-asr-streaming-0.6b, which is multilingual and prompt-conditioned, covering more than 40 locales. A target language is passed with a --lang flag that defaults to auto, and both the offline and cache-aware streaming transcripts match NeMo per language at WER 0. For anyone building voice agents or dictation tools, a streaming model that behaves identically to the reference implementation while running as a single native binary is the difference between a demo and a product.
The supported model lineup
Every model in the table is published as GGUF in the single collection repo mudler/parakeet-cpp-gguf, in f16, q8_0, q6_k, q5_k, and q4_k variants. The list covers parakeet-tdt_ctc-110m as the small English anchor, the 0.6B family in CTC, RNNT, and TDT flavors including the multilingual v3 TDT with 25 European languages, the 1.1B family across the same architectures, the 120M realtime EOU streaming model, and the 0.6B multilingual nemotron streaming model released under OpenMDW-1.1.
All of them can be converted yourself with the scripts/convert_parakeet_to_gguf.py script if you prefer not to trust the published quantizations. That combination, one collection repo for everything plus a converter script, keeps the setup story simple: pick a model, download the GGUF, run the CLI.
Measured speed, not marketing speed
The performance claims come with methodology and plots in benchmarks/BENCHMARK.md. On a 20-core x86 CPU against NeMo PyTorch on LibriSpeech test-clean with 8 threads, f32 runs 1.11 to 1.69 times faster with a median of 1.40x while staying byte-identical. Quantization pushes further: f16 reaches up to 1.70x at 57 percent of the f32 size, q8_0 up to 1.86x at 37 percent, and q4_k reaches 26 percent of the size with a small monotonic WER cost. Peak RAM also lands roughly 2 times lower than NeMo, and lower still once quantized.
On GPU, tested on an NVIDIA GB10 Grace-Blackwell system against NeMo in its own container, parakeet.cpp wins on all 10 models with a median of 1.25x and up to 4.3x on the large TDT and hybrid models. The README attributes most of the gap to NeMo's TDT greedy decoding lacking CUDA-graph acceleration while this implementation uses a lean C++ loop, and notes the log-mel front end runs on the GPU through a ggml DFT-matmul graph with the CPU path unchanged. There is even a recorded race against NeMo on the same clip, slowed down so the sub-100ms finish is watchable.
Against whisper.cpp on the same audio
The most useful comparison for most users is whisper.cpp, the incumbent local ASR binary. parakeet.cpp runs circles around it on speed: the 110M Parakeet is faster than whisper base.en and far faster than large-v3-turbo, while the larger Parakeets match or beat whisper's accuracy. The published races put the gap at about 12 times faster against whisper.cpp turbo on GPU and about 27 times faster on CPU, at the same accuracy.
That framing matters when choosing an ASR backend. Whisper remains the broader multilingual choice and has its own ecosystem, but for English and the covered European languages, a Parakeet checkpoint in quantized GGUF gives you the accuracy tier you need at a fraction of the compute, which matters for always-on listening, batch transcription of archives, and edge devices where every watt counts.
Pre-built binaries and where it fits
Every release ships pre-built parakeet-cli bundles, so compilation is optional. The platform table covers Linux x64 in cpu, vulkan, and cuda variants, Linux arm64 in cpu, and macOS, with the quantization format options carried through the GGUF collection. Because inference needs no Python runtime, deployment targets include servers, desktops, and anything that can run a static binary.
The natural fits are the places where whisper.cpp has been the default answer: local dictation, meeting transcription, subtitle generation, voice pipelines inside larger applications, and embedding ASR into products without shipping a Python environment. The LocalAI connection points the same direction, since the team's broader goal is running models locally on any hardware, and a byte-identical, faster-than-PyTorch ASR port is a clean addition to that stack.
Editorial conclusion
parakeet.cpp is what a good model port looks like: pick a family with strong accuracy, port it to ggml, hold yourself to byte-identical transcripts against the reference, publish the benchmark methodology, and ship pre-built binaries. With all Parakeet families covered, cache-aware streaming with end-of-utterance detection, multilingual support past 40 locales, and measured speedups from 1.1x to 27x depending on the baseline, it gives local ASR a new default for anyone who does not need whisper's full language breadth.
Frequently asked questions
What is Whisper CPP used for?
whisper.cpp runs OpenAI Whisper speech recognition locally in C++ via ggml, commonly for offline transcription, subtitles, dictation, and voice pipelines on CPUs and GPUs without Python. parakeet.cpp targets the same use cases and, per its benchmarks, runs the same audio about 12x faster on GPU and 27x faster on CPU against whisper.cpp turbo at matching accuracy for covered languages.
What are the key differences between Nemotron 3.5 ASR and Whisper?
In the parakeet.cpp context, nvidia/nemotron-3.5-asr-streaming-0.6b is a multilingual RNNT streaming model covering 40 plus locales with prompt-conditioned language selection via --lang, while Whisper models are encoder-decoder transcribers typically run offline per chunk. Nemotron 3.5 supports cache-aware streaming that matches NeMo byte for byte, and it runs as a small 0.6B quantized GGUF with no Python runtime.
Does parakeet.cpp support quantized models and how much accuracy is lost?
Yes, every checkpoint is published as GGUF in f16, q8_0, q6_k, q5_k, and q4_k. f16 and q8_0 are near-lossless while running up to 1.86x faster than NeMo on CPU, and q4_k cuts the model to 26 percent of the f32 size with a small monotonic WER cost, so you trade a little accuracy for a lot of footprint.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/mudler-parakeet-cpp)