# parakeet.cpp runs NVIDIA Parakeet speech models in C++ with Metal acceleration and no ONNX or Python runtime

> parakeet.cpp is a pure C++ implementation of NVIDIA's Parakeet speech recognition models, built on the author's axiom tensor library for automatic Metal acceleration on Apple Silicon. It covers offline and streaming models, diarization, beam search with an ARPA language model, and a flat C API, but the documented install path points at GitHub releases and the repository currently has none.

**Frikallo/parakeet.cpp** — Ultra fast and portable Parakeet implementation for on-device inference in C++ using Axiom with MPS+Unified Memory

- Repository: https://github.com/Frikallo/parakeet.cpp
- Stars: 306 · Forks: 15
- Language: C++
- License: MIT
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/frikallo-parakeet-cpp

## The install path points at GitHub releases the repository has not published

Installation is documented as downloading a prebuilt tarball attached to a GitHub release, for macOS arm64, macOS x86_64, or Linux x86_64, and unpacking it where you like.

```bash
tar -xzf parakeet-v0.1.0-macos-arm64.tar.gz
cd parakeet-v0.1.0-macos-arm64
# On macOS, clear the Gatekeeper quarantine attribute first:
xattr -dr com.apple.quarantine .
./bin/parakeet --help
```

The layout is designed for relocation. The archive ships a self-contained `bin/parakeet` and `bin/example-server` plus `lib/libaxiom` with `@rpath` and `$ORIGIN` set, so the binaries resolve their own dependencies relative to the install directory instead of a system path. C API headers arrive under `include/parakeet/` for embedders.

The gap is that the project currently has no GitHub releases at all, and no versioned tag to point a download at. That means the documented quick route is not available to a reader today, and the `xattr` step above only matters once a tarball exists to quarantine.

## Building on macOS requires full Xcode because axiom compiles Metal shaders

The source build is short, but one requirement is easy to get wrong.

```bash
git clone --recursive https://github.com/frikallo/parakeet.cpp
cd parakeet.cpp
make build
make test
```

The clone is recursive because axiom is pulled in as a submodule, and the compiler requirements are C++20 with Clang 14 or GCC 12 or newer, CMake 3.20 or newer, and macOS 13 or newer for the Metal GPU path.

The macOS caveat is specific: building needs the full Xcode install rather than Command Line Tools, because axiom compiles its Metal shaders with `xcrun metal` and `xcrun metallib`, and those tools ship only with Xcode. The documented way around it is to run the prebuilt tarball instead of compiling, since the `.metallib` is embedded into the shipped `libaxiom.dylib` and the end user needs no Xcode or CLT at all. So the build tooling cost sits on whoever compiles it, not on whoever runs it.

## Every ASR model shares one audio path, fixed at 16 kHz mono into an 80-bin Mel spectrogram

All ASR models run through the same front end: 16kHz mono WAV becomes an 80-bin Mel spectrogram and then a FastConformer encoder. That shared pipeline is what makes the decoder the only thing you really choose, and it also means the format handling is handled for you rather than by you.

Input files can be WAV, FLAC, MP3, or OGG with automatic resampling, so a 44.1 kHz stereo MP3 does not need converting before it is passed to a transcriber. Timestamps are per word with start and end times plus confidence scores on all decoders.

Silence is handled before recognition rather than after. Silero VAD strips silence ahead of ASR, and the timestamps are remapped afterwards so the word times still refer to the original recording rather than to the trimmed buffer. Silero weights are produced by their own converter, separate from the Parakeet weight conversion, which matters if you want voice activity detection without shipping a recogniser.

## Six model classes split along offline and streaming lines

The supported set is organised by whether audio arrives in one piece or continuously, and by whether the output is text alone or text with speaker labels.

| Model | Class | Size | Type | Description |
|-------|-------|------|------|-------------|
| `tdt-ctc-110m` | `ParakeetTDTCTC` | 110M | Offline | English, dual CTC/TDT decoder heads |
| `tdt-600m` | `ParakeetTDT` | 600M | Offline | Multilingual, TDT decoder |
| `eou-120m` | `ParakeetEOU` | 120M | Streaming | English, RNNT with end-of-utterance detection |
| `nemotron-600m` | `ParakeetNemotron` | 600M | Streaming | Multilingual, configurable latency (80ms to 1120ms) |
| `sortformer` | `Sortformer` | 117M | Streaming | Speaker diarization (up to 4 speakers) |
| `diarized` | `DiarizedTranscriber` | 110M+117M | Offline | ASR plus diarization to speaker-attributed words |

Two things stand out. Nemotron trades accuracy against responsiveness with a configurable latency window from 80ms to 1120ms, which is the knob for a live captioning use case. And the diarized class is a composite: 110M of recogniser plus 117M of Sortformer, so speaker attribution costs roughly double the model weight of transcription alone and is capped at four speakers.

## Decoder and search strategy are chosen at the call site

Loading a model does not commit you to a decoding strategy. Multiple decoders are available and switch at the call site: CTC greedy, TDT greedy, CTC beam search, and TDT beam search. The `tdt-ctc-110m` model carries both CTC and TDT heads precisely so that decision can be deferred, while `ParakeetTDTCTC` exposes both, `ParakeetCTC` uses greedy argmax or beam search with an optional language model, `ParakeetRNNT` is autoregressive LSTM, and `ParakeetTDT` adds duration prediction on top of the LSTM.

Beam search can fuse an ARPA n-gram language model, which is the usual answer to a proper-noun problem in medical or product vocabulary. A second mechanism targets the same problem differently: phrase boosting biases decoding through a token-level trie, so a domain vocabulary can be pushed without retraining the model. Word timestamps with confidence scores are available on every decoder, not only on the accurate one.

The streaming path is separate underneath. Streaming models use a cache-aware FastConformer encoder with causal convolutions and bounded-context attention, and take chunked audio input, which is why EOU and Nemotron behave differently from the offline models over the same audio.

## Weight conversion is a Python step that a C++ runtime cannot skip

The C++ side loads safetensors and a vocabulary file, not the NVIDIA checkpoints, so a conversion pass sits between the download and the first transcription. That pass is Python and needs torch.

```bash
# Download from HuggingFace
huggingface-cli download nvidia/parakeet-tdt_ctc-110m --include "*.nemo" --local-dir .

# Convert to safetensors
pip install safetensors torch
python scripts/convert_nemo.py parakeet-tdt_ctc-110m.nemo -o model.safetensors
```

One converter covers every model type. `110m-tdt-ctc` is the default, and the others are selected with the model flag, for example `--model 600m-tdt` for the multilingual 600M TDT checkpoint, with `eou-120m`, `nemotron-600m`, and `sortformer` also supported. Silero VAD weights come from `scripts/convert_silero_vad.py -o silero_vad_v5.safetensors`.

The practical consequence is that the promise of no Python runtime applies to inference, not to model preparation. Shipping an app means shipping safetensors you converted ahead of time.

## The Makefile switches to Ninja when it finds it, and gates the optional binaries

Build control sits in a Makefile that probes for Ninja and falls back to Unix Makefiles with a job count from `hw.ncpu` or `nproc`, defaulting to 4. `make build` configures and compiles in one step, `make test` runs the test binary, `make debug` switches the build type, and `make bench` runs the benchmark, with `ARGS` passed through.

Optional targets are switched with variables rather than a separate script. `make build CLI=OFF` disables the command line binary, and `make build SERVER=ON` enables the example server, which map onto CMake flags for the CLI and the server example. `make install` runs `cmake --install`, and there are `format` and `format-check` targets that run clang-format across src, include, and examples, with `format-check` using `--dry-run --Werror`.

Three consumption paths are supported once installed: CMake `find_package(Parakeet)` with `target_link_libraries(myapp PRIVATE Parakeet::parakeet)`, `add_subdirectory` for a vendored copy, and pkg-config for a plain compiler invocation. That breadth is what makes it usable both as a system dependency and as a third_party directory.

## Embedding means a flat extern C surface rather than a C++ ABI

There is a C API described as a flat `extern "C"` FFI, intended for Python, Swift, Go, Rust, and other languages, with the headers shipped under `include/parakeet/` and a pure C99 example that demonstrates calling it. The word flat matters: no C++ types cross the boundary, so a binding written in a language without a C++ runtime does not need one at runtime.

The examples directory doubles as the feature documentation, since each feature has its own small program: basic for the shortest transcription path, timestamps, beam-search, phrase-boost, batch, vad, gpu for Metal plus FP16 with a timing comparison, stream, nemotron, diarize, diarized-transcription, c-api, cli for the full command line with all options, and server.

At the top level the C++ entry is intentionally small. A transcriber takes a model path and a vocabulary path, optionally moves to the GPU with `to_gpu()` and to half precision with `to_half()`, and returns a result whose text field is what you print. The project quotes about 27ms encoder inference for 10 seconds of audio with the 110M model on an Apple Silicon GPU, a 96x speedup over CPU, and roughly 2x memory reduction from FP16, which is the reason the Metal path exists at all.

## Conclusion

parakeet.cpp fits anyone shipping on-device transcription in an app that cannot carry an ONNX or Python runtime, particularly on Apple Silicon where the Metal path is the reason to pick this over a Python binding. It does not fit a team that needs a stable published artifact today, since the install instructions depend on release tarballs the repository does not yet have, and building from source on macOS costs a full Xcode install rather than Command Line Tools. Before committing, confirm you can run `xcrun metal`, decide whether you need offline or streaming models, and check that your decoder choice is available at the call site rather than fixed at load time.

## FAQ

### Which speech models does parakeet.cpp support?

Six classes: tdt-ctc-110m and tdt-600m for offline transcription, eou-120m for English streaming with end-of-utterance detection, nemotron-600m for multilingual streaming with latency configurable from 80ms to 1120ms, sortformer for diarization of up to 4 speakers, and a diarized class that combines ASR with diarization for speaker-attributed words.

### Does building parakeet.cpp on macOS require Xcode?

Yes, the full Xcode install rather than Command Line Tools, because axiom compiles Metal shaders with xcrun metal and xcrun metallib. Running the prebuilt tarball needs no Xcode at all, since the metallib is embedded in the shipped libaxiom.dylib.

### How do I get model weights for parakeet.cpp?

Download the .nemo checkpoint from HuggingFace with huggingface-cli, install safetensors and torch, then run scripts/convert_nemo.py to produce model.safetensors. One converter covers 110m-tdt-ctc as the default plus 600m-tdt, eou-120m, nemotron-600m, and sortformer selected with the --model flag.

### Can parakeet.cpp handle streaming audio?

Yes. The eou-120m and nemotron-600m models take chunked audio input through a cache-aware streaming FastConformer encoder with causal convolutions and bounded-context attention, and Nemotron exposes configurable latency from 80ms to 1120ms.

### What are the build requirements for parakeet.cpp?

C++20 with Clang 14 or newer or GCC 12 or newer, CMake 3.20 or newer, and macOS 13 or newer for the Metal GPU path. The documented build is git clone --recursive, cd parakeet.cpp, make build, then make test.

## Sources

- [Frikallo/parakeet.cpp on GitHub](https://github.com/Frikallo/parakeet.cpp)
- [Issues](https://github.com/Frikallo/parakeet.cpp/issues)
- [License: MIT](https://github.com/Frikallo/parakeet.cpp/blob/main/LICENSE)
- [README](https://github.com/Frikallo/parakeet.cpp/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/frikallo-parakeet-cpp
