# parakeet-rs is one crate, six Parakeet models and a pinned runtime API

> parakeet-rs wraps NVIDIA's Parakeet family in Rust over ONNX Runtime, covering offline transcription, cache-aware streaming, speaker diarization and multi-speaker attribution behind one API and a set of feature flags. The friction is not the models, it is the execution provider and the ONNX Runtime API level you have to match.

**altunenes/parakeet-rs** — very fast speech-to-text, diarization, streaming (even in CPU) with NVIDIA Parakeet in Rust

- Repository: https://github.com/altunenes/parakeet-rs
- Website: https://huggingface.co/altunenes/parakeet-rs
- Stars: 404 · Forks: 65
- Language: Rust
- License: MIT
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/altunenes-parakeet-rs

## The api-24 through api-28 flags pin an ONNX Runtime ABI

The feature list in `Cargo.toml` is where the real constraints live. The default is `["cpu", "ort-defaults", "api-28"]`, and the `api-24` through `api-28` features each forward to the matching ONNX Runtime API in the `ort` crate.

That means the default build expects a runtime at API level 28. Every other entry in the feature list is an execution provider: `cuda`, `tensorrt`, `coreml`, `directml`, `migraphx`, `openvino`, `webgpu`, `nnapi`, plus `load-dynamic` and `preload-dylibs`.

Two loading modes sit alongside the providers. `load-dynamic` defers to a runtime found on the system instead of a linked one, and `preload-dylibs` handles the shared libraries those providers depend on. On Linux with a system-installed runtime, `load-dynamic` is usually the path you want.

The gap in the documentation is what happens on a mismatch. The project does not spell out the error you get when the runtime is older than the pinned API level, so treat that combination as something to verify locally rather than something to assume. The runtime dependency is `ort` 2.0.0-rc.13, a release candidate, with default features switched off and only `std` and `ndarray` enabled.

## CoreML is called unstable, so Apple silicon is steered to WebGPU

The platform note at the top is blunt: CoreML is unstable with this model. For Apple hardware the recommendation is the WebGPU execution provider, which uses Metal underneath despite the name, or plain CPU.

The comparison offered is a single data point on a single machine, an M3 with 16GB, where the author reports CPU alone beating Whisper on Metal. It is one machine and one Whisper configuration, not a benchmark, but it is enough to say the CPU path is worth measuring before you buy a GPU runtime.

That leaves `webgpu` as the Apple entry to try first, `cpu` as the always-available fallback, and `coreml` as something the project itself flags. Everything else in the feature list, cuda and tensorrt and openvino and directml, is aimed at other hardware.

Audio handling is self-contained: `hound` reads WAV files, `realfft` does the transform work, and `tokenizers` with the `onig` feature loads the vocabulary. Error handling uses `eyre`, and `serde` plus `serde_json` carry the model metadata.

## Streaming is three chunk sizes, and each model family gets its own

The streaming examples make the latency knob explicit by fixing the chunk size in samples. End-of-utterance streaming processes 2560 samples, commented as 160ms at 16kHz, so text appears in small pieces and the model decides where an utterance ends. The Nemotron cache-aware model uses 8960 samples, 560ms at 16kHz, and adds punctuation. Multi-speaker attribution uses 17920 samples, about 1.12s at 16kHz, which is a longer window because speaker identity has to stay stable across it.

Audio everywhere is `Vec<f32>` at 16kHz mono and normalized, and the same `transcribe` shape appears across families, so switching models is mostly a change of constructor and chunk size.

The Cohere Transcribe addition is the outlier: it is offline and multilingual across 14 languages, takes its arguments as `lang`, `pnc` and `itn` for punctuation and inverse text normalization, and supports long-form audio. It sits behind the `cohere` feature and a runnable demo in `examples/cohere.rs`.

The crate also ships `examples/unified.rs`, `examples/shared_model.rs` and `examples/streaming.rs`, so the repository is where the argument shapes are actually documented.

## Nemotron auto-detects the variant, so set_target_lang can silently do nothing

The Nemotron streaming loader points `from_pretrained` at a directory of ONNX files and detects which variant it got. Two share the API and differ in what they emit.

The English-only 0.6B model is described as verbatim, preserving disfluencies such as `um` and `uh`, which is what you want when every spoken word carries meaning. The Multilingual 3.5 0.6B model covers 40 language-locales across three tiers, 19 transcription-ready, 13 broad coverage, and 8 that NVIDIA says need fine-tuning before they reach production quality. It produces polished output with proper casing and punctuation and drops disfluencies, at the same speed and size as the English-only model.

That auto-detection is convenient and it hides a trap. The language hint only applies to the multilingual variant:

```rust
use parakeet_rs::{Nemotron, NemotronMode};

let mut model = Nemotron::from_pretrained(path, None)?;
```

`set_target_lang` is documented as a no-op when an English-only model is loaded, so a mistyped hint produces no error and no effect. Check `model.mode()` against `NemotronMode::Multilingual` before you trust a language setting.

## Diarization latency presets are NVIDIA's numbers, up to eight speakers

Speaker diarization is a separate feature, `sortformer`, wrapping Nemotron-3 diarization in Sortformer v3, and it handles up to 8 speakers. The constructor takes a model file path, `diarize` takes audio with its sample rate and channel count, and each segment carries a speaker id plus start and end values that the example divides by 16,000 to print as seconds.

The latency presets come straight from NVIDIA's model card and are exposed as constructor-style calls: `offline()` at 30.4 seconds is the default, `low_latency()` at 1.04 seconds, and `very_low_latency()` at 0.64 seconds, with the example runner accepting `low`, `very-low` or `ultra`. Those figures are the model's published chunk lengths, not measurements from this project.

For real-time use, `diarize_chunk()` and `feed()` preserve state across calls, and `examples/streaming_diarization.rs` demonstrates the `feed` and `flush` pair.

If you need your own streaming parameters, `scripts/export_diar_sortformer.py` exports the ONNX at dual resolution with self-describing metadata. The `multitalker` feature sits on top of `sortformer` and adds speaker-attributed text per chunk.

## The crate licence is MIT OR Apache-2.0, and the weights are not

`Cargo.toml` declares `license = "MIT OR Apache-2.0"`, so the Rust code is permissively licensed in either choice. The model weights are a separate matter and each download carries its own terms.

Parakeet Ultra, described as Moondream's post-trained parakeet-tdt-0.6b-v3 with better accuracy, is CC-BY-4.0. Orukeet, a pinned INT8 export that works with the same `ParakeetTDT` runtime, is CC BY-SA 4.0, which is copyleft and about 672 MB of weights. The base CTC and TDT exports come from the community and upstream HuggingFace repositories.

That distinction matters for anyone shipping a product. The crate's licence tells you nothing about the artefact you actually execute, and the strongest copyleft term in the set is attached to the fastest path.

The model weights are also downloaded by hand rather than pulled by the build. CTC needs `model.onnx`, `model.onnx_data` and `tokenizer.json`. TDT needs `encoder-model.onnx`, `encoder-model.onnx.data`, `decoder_joint-model.onnx` and `vocab.txt`. The README does not document a cache directory convention or a first-run download.

## Orukeet's downloader verifies hashes and licences, then works offline

Orukeet has the only scripted download in the repository, and it is the most careful part of the setup documentation. The sequence is:

```bash
python3 -m pip install huggingface-hub
model_dir=$(python3 scripts/download_orukeet.py)
cargo run --release --example orukeet -- "$model_dir" audio.wav
# After installation, no network is needed:
model_dir=$(python3 scripts/download_orukeet.py --offline)
```

The downloader is described as verifying a release manifest and all required file hashes, licences included, and then printing the cached model directory so you can pass it to the example. The `--offline` flag is the point: once the cache is populated, the example runs with no network at all.

It also adds nothing to the runtime. The text is explicit that this uses the existing local TDT runtime, adds no streaming support and changes no defaults, so adopting Orukeet is a model swap rather than a new code path.

Parakeet Ultra is handled differently. It is the same file set as TDT, loads through `ParakeetTDT`, and you can regenerate it yourself with `scripts/export_parakeet_ultra.py`.

## The CTC example passes 1600 where the TDT example passes 16000

One detail does not line up. The CTC snippet calls `transcribe_samples(audio, 1600, 1, Some(TimestampMode::Words))` while the TDT snippet calls the same method with `16000`. Those two arguments are the sample rate and the channel count, and every other snippet in the readme, including the streaming examples and the Cohere call, works at 16kHz mono.

A 1600 Hz interpretation of 16kHz audio would play back at a tenth of the speed, so treat that value as a documentation error and use the rate your audio actually has. If you are copying from the readme rather than from `examples/raw.rs`, that is where the mistake will land.

The timestamp modes are the other per-model difference. CTC is shown with `TimestampMode::Words`, TDT with `TimestampMode::Sentences`, and both then iterate `result.tokens` printing each token's start, end and text. So the granularity of what you can align to your audio is chosen per model family, not per call site.

The repository's own example list is longer than the readme shows: `raw.rs`, `streaming.rs`, `unified.rs`, `shared_model.rs`, `diarization.rs`, `streaming_diarization.rs`, `multitalker.rs`, `cohere.rs` and `orukeet.rs`. Reading those is more reliable than reading the snippets.

## Conclusion

parakeet-rs suits Rust projects that need local transcription without a Python process, especially streaming and speaker-attributed work where the chunk sizes and latency presets are documented in the examples. Before you build on it, check three things: that your execution provider is one the crate actually enables, because CoreML is described as unstable with these models and Apple users are pointed at WebGPU or CPU, that your ONNX Runtime matches the `api-28` default or that you pin a lower one, and that your weights licence fits your use, since the crate is MIT OR Apache-2.0 while Orukeet is CC BY-SA 4.0. Version 0.3.8 shipped 2026-09-23 and still builds on `ort` 2.0.0-rc.13.

## FAQ

### Which execution provider should parakeet-rs use on Apple silicon?

The project states that CoreML is unstable with this model and points Apple users at the WebGPU execution provider, which uses Metal underneath, or plain CPU. The author reports CPU alone beating Whisper on Metal on an M3 with 16GB, which is one machine rather than a benchmark.

### What does the api-28 feature in parakeet-rs do?

The api-24 through api-28 features forward to the matching ONNX Runtime API in the ort crate, and api-28 is part of the default feature set alongside cpu and ort-defaults. The default build therefore expects a runtime at API level 28.

### How do I get streaming transcription out of parakeet-rs?

The end-of-utterance model is fed 2560-sample chunks, which the example comments as 160ms at 16kHz, while the Nemotron cache-aware model uses 8960 samples, 560ms at 16kHz, and adds punctuation. Audio is f32, 16kHz mono, normalized.

### Can parakeet-rs tell different speakers apart?

Yes, behind the sortformer feature for streaming diarization up to 8 speakers, and behind the multitalker feature for speaker-attributed text per chunk. Latency presets come from NVIDIA's model card, with offline at 30.4 seconds as the default and lower latency options at 1.04 and 0.64 seconds.

### What licence applies to the parakeet-rs models?

The crate is MIT OR Apache-2.0, but the weights carry their own terms. Parakeet Ultra is CC-BY-4.0 and Orukeet is CC BY-SA 4.0 at about 672 MB, so read the licence of the specific model you download rather than the licence of the crate.

## Sources

- [altunenes/parakeet-rs on GitHub](https://github.com/altunenes/parakeet-rs)
- [License: MIT](https://github.com/altunenes/parakeet-rs/blob/master/LICENSE)
- [Project website](https://huggingface.co/altunenes/parakeet-rs)
- [README](https://github.com/altunenes/parakeet-rs/blob/master/README.md)
- [Releases](https://github.com/altunenes/parakeet-rs/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/altunenes-parakeet-rs
