OpenASR: a local-first speech-to-text CLI with a signed model catalog
Local-first speech-to-text: no cloud, no telemetry, fail-closed by design. One CLI, seven model families, signed model catalog, OpenAI-compatible local API.
At a glance
- What is it?
- OpenASR is an Apache-2.0 Rust CLI and local HTTP server that runs 30 speech models across 16 families on your own machine. The interesting part is not the model list, it is the fail-closed design: no telemetry, no silent uploads, and a signed catalog check before any pack runs.
- Who is it for?
- Adopt OpenASR if audio cannot leave your machine, you want an OpenAI-compatible endpoint at 127.0.0.1:8080, and you accept pre-v1 churn in CLI flags and pack format. Skip it if you need a stable API contract today, or if your only target is a GPU stack other than CUDA and Metal, since Vulkan and ROCm ship as release archives rather than images.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 10 days ago.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem OpenASR picks: transcription that cannot phone home
Most speech-to-text tooling assumes a network call. You send audio to an endpoint, get text back, and trust the vendor's retention policy. OpenASR inverts that default. The README states that in the default local mode audio stays on the machine, and that remote compute is available only when you explicitly pair and enable it. The engine is described as producing a real transcript or telling you why it cannot, rather than falling back to a remote service.
The audience is narrow and specific. It is engineers who have audio they are not allowed to ship to a third party: legal recordings, medical dictation, internal meeting capture, anything under a data processing agreement that names a jurisdiction. It is also people who want an OpenAI-shaped endpoint without an OpenAI bill, which is why the server exposes /v1/audio/transcriptions and the README says existing SDKs work against base_url="http://127.0.0.1:8080/v1".
The project is pre-v1 and says so in the README: CLI flags, API surface, and pack format may change between 0.x releases. That is the honest framing, and it should shape how you adopt it. Pin a version, read the changelog before bumping, and do not build a product surface on flags you have not checked against the release you installed.
How the engine, the registry and the server fit together
The repository layout tells you more than the marketing copy. Cargo.toml lists a workspace of seven members: openasr-cli, openasr-client, openasr-core, openasr-ffi, openasr-server, openasr-system-audio, and xtask. The inference backend is ggml, MIT-licensed per the README, and the crate list shows metal and matrixmultiply alongside realfft, cpal for capture, and axum with multipart and ws features for the HTTP layer. CPU and Apple Metal are the two execution targets named in the README.
Model distribution is the part worth understanding before you install anything. A model-registry directory sits at the repository root and is copied into the runtime image by the Dockerfile. The README says every model download is verified against a signed catalog before it runs, and the workspace pulls in ed25519-dalek, which is the signature primitive you would expect for that check. Packs are not part of the image: the Docker section states the published images carry binary plus model-registry metadata only, and that you pull models at runtime into a volume mounted at /data. The HTTP server never auto-downloads a pack.
Concurrency has a documented shape too. Offline native requests are serial by default. The --max-native-sessions-per-model N flag sets both an admission limit and, for eligible direct-GPU Cohere, Moonshine, Qwen, and Whisper jobs, a batch width capped at 8. CPU, scheduler, adapter, realtime, FireRed-AED, and FireRed2 paths stay serial, and translations follow the offline policy. If you were hoping to saturate a machine with parallel transcriptions, that paragraph is the constraint you need to read twice.
Installing OpenASR and getting a first transcript
Three install routes are documented: a Homebrew tap, a one-line installer script, and prebuilt binaries from Releases. The Homebrew and installer paths are for macOS and Linux.
brew install quintinshaw/tap/openasrOr, if you prefer the script, the README gives this exact command:
curl -fsSL https://dl.openasr.org/install.sh | shPiping a remote script into a shell is a decision, not a formality. The project publishes checksums and a signed catalog for models, but the install script itself is fetched over HTTPS with no verification step shown in the README. If that matters to your environment, take the Releases binary instead and verify it against the release notes.
The first transcription is deliberately interactive. The README notes that the first run offers to download a model and that you confirm first.
openasr transcribe recording.wavTo see what is available before committing to a download, browse the catalog:
openasr search
openasr pull whisper-smallSubtitles with speaker labels use two flags together, and the README gives this example:
openasr transcribe meeting.wav -f srt --diarizeFor the server path, start it and post a file to the OpenAI-compatible route. The model name in the form field is the registry name, not a filesystem path.
openasr serve
curl http://127.0.0.1:8080/v1/audio/transcriptions \
-F [email protected] -F model=qwen3-asr-0.6bIf you would rather not install anything on the host, the Docker route mounts a data volume and pulls the pack inside the container:
docker pull quintinshaw/openasr:latest
docker run --rm -d --name openasr \
-p 8080:8080 -v openasr-data:/data quintinshaw/openasr:latest
docker exec openasr openasr pull whisper-small --yesThe container defaults to binding 0.0.0.0:8080 with a self-signed certificate and device pairing. On first start, if OPENASR_PAIRING_ADMIN_TOKEN is unset, serve generates a random token, writes it owner-only to /data/pairing-admin-token, and prints it to stdout. Read that token before you expose the port.
The fail-closed choices, and where they cost you
Fail-closed is a real design decision with a real bill attached. The CUDA image is the clearest example. The compose file states that the entrypoint runs openasr doctor before serving and refuses to start if no GPU device is visible inside the container, and that it will not silently fall back to CPU. If startup fails with "no GPU device reported", the deploy.resources.reservations.devices block did not actually grant a GPU. That is a deliberate trade: a misconfigured host fails loudly instead of quietly running ten times slower.
The same instinct shows up in the TLS handling. Behind a TLS-terminating reverse proxy you may drop --tls-self-signed and set OPENASR_ALLOW_INSECURE_NON_LOOPBACK=1, and the Dockerfile comment is explicit that this env only waives TLS, never pairing. So there is no configuration in which the server becomes an open transcription endpoint on a non-loopback address. If you wanted a bare HTTP service on a trusted internal network with no pairing step, this is the wrong tool, and no flag will make it that tool.
The serial-by-default execution model is the other cost. Batching is available only for eligible direct-GPU Cohere, Moonshine, Qwen, and Whisper jobs, capped at 8, and gated behind --max-native-sessions-per-model. Everything else, including CPU, realtime, and the FireRed paths, stays serial. A team expecting to serve many concurrent users from one box should size the hardware for that, not for the batch width.
The README also points at docs/KNOWN_LIMITATIONS.md for what does not work yet, and that file is not reproduced here. Treat it as required reading before you promise a capability to anyone.
OpenASR compared with whisper.cpp and faster-whisper
The obvious alternative for local transcription is whisper.cpp, and the difference is scope rather than speed. whisper.cpp is an inference engine for Whisper, with bindings and a small CLI; you supply the model file and wire up the rest yourself. OpenASR is a distribution and orchestration layer around a ggml engine: a signed model catalog, a pull command, a registry that records each pack's upstream license, a server with pairing, and a desktop app. You are trading control for an opinionated pipeline.
faster-whisper sits on the other side of that line. It is a Python library built on CTranslate2, so it fits naturally into a Python service where you already manage model files and concurrency yourself, and it gives you programmatic control that a CLI does not. OpenASR gives you a binary, a local HTTP endpoint, and no Python dependency, at the cost of the pre-v1 flag churn the README warns about.
The honest summary: if you want one model family and full control over the inference loop, whisper.cpp or faster-whisper will be less machinery. If you want seven-plus model families behind one binary, a verified download path, and an OpenAI-shaped endpoint that never leaves the host, OpenASR is doing work those two do not attempt.
Licensing, releases and what an upgrade actually costs
The engine is Apache-2.0, and the ggml backend is MIT. That part is simple. The model packs are not. The README states that each pack ships under its own upstream license as recorded in the registry and pack metadata, and that packs may use Apache-2.0, MIT, CC-BY, FunASR, or other upstream terms. The README itself adds that this is not an exhaustive license guarantee. If you are shipping a product, the per-pack license in model-registry is the line you need to read, not the repository LICENSE file. The name, logo, and official app icons are reserved separately, and Apache-2.0 covers the code, not the branding.
Upgrade cost is shaped by the release cadence. The workspace version in Cargo.toml is 0.1.41, while the most recent listed releases are v0.1.36 and desktop-v0.1.22, both dated 2026-08-22, with v0.1.35 the day before. That is a fast 0.x line. The README's own warning that CLI flags, API surface, and pack format may change between 0.x releases is the practical consequence: a bump can require re-checking your flags and your pinned model names.
The workspace has one deliberate mitigation. Inter-crate dependencies are path-only with no version pin, and every member crate inherits its version from [workspace.package], so a release bump touches one field. That keeps internal version skew from becoming a problem. It does not protect your scripts. Pin the version you install, keep the model pack files you pulled, and read CHANGELOG.md before moving.
What the documentation does not settle
Several things a reader will want are not addressed anywhere in the README or the repository files. There is no rollback procedure documented for a model pack once it is pulled, and no stated mechanism for pinning a pack to a specific revision. The README does not describe what happens to a running server when the registry metadata changes under it.
Performance claims are pointed at rather than made. The README references a committed performance baseline in perf/PERFORMANCE.md and says benchmarks live there; it does not quote numbers in the body, and the repository description's claim of models that run faster than real-time is not substantiated in the text above. If throughput decides your adoption, that file is where to look, and you should reproduce it on your own hardware rather than trusting a table.
The desktop app is described as wrapping the same engine in a native GUI with no hidden network calls, distributed for macOS on Apple Silicon and Windows x64 on Windows 10+, with Linux desktop listed as coming soon. The relationship between the desktop release train (desktop-v0.1.22) and the core release train (v0.1.36) is not spelled out, so do not assume a desktop build tracks the newest core version.
Editorial conclusion
Adopt OpenASR if audio cannot leave your machine, you want an OpenAI-compatible endpoint at 127.0.0.1:8080, and you accept pre-v1 churn in CLI flags and pack format. Skip it if you need a stable API contract today, or if your only target is a GPU stack other than CUDA and Metal, since Vulkan and ROCm ship as release archives rather than images. Before committing, run openasr doctor on the exact host, pull the one model family you need with openasr pull, and read docs/KNOWN_LIMITATIONS.md plus the license line for that pack in model-registry.
Frequently asked questions
What is OpenASR?
OpenASR is a local-first speech-to-text tool: a Rust CLI, a local OpenAI-compatible HTTP API, and a ggml inference engine, released as the Apache-2.0 open core behind the OpenASR desktop app. It runs 30 models across 16 families on CPU and Apple Metal, with each model download verified against a signed catalog before it runs.
Is OpenASR considered AI?
The README does not discuss OpenASR in terms of AI categories. It describes a ggml inference engine that runs speech models locally, so the question of whether that counts as AI is outside what the documentation settles.
How do I install OpenASR and pull a model?
On macOS or Linux you can use the Homebrew tap quintinshaw/tap/openasr, the installer at https://dl.openasr.org/install.sh, or a prebuilt binary from Releases. Models are separate: openasr search browses the catalog and openasr pull whisper-small installs one, and the first transcribe run offers to download a model with your confirmation.
Does OpenASR send my audio anywhere?
In the default local mode, the README states audio stays on your machine. Remote compute is available only when you explicitly pair and enable it, and the project documents no telemetry and no silent network fallback.
Can OpenASR run as a Docker container with a GPU?
Yes, via the cuda-latest image or the gpu compose profile, which requires the NVIDIA Container Toolkit and a driver compatible with CUDA 13.2. The entrypoint runs openasr doctor before serving and refuses to start if no GPU device is visible, rather than falling back to CPU.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/quintinshaw-openasr)