hayamimi routes each utterance to a specialist model and never touches a GPU
早耳 - Real-time multilingual speech-to-text on CPU only. Live subtitles, browser dashboard, speaker labels, translation. No GPU, no cloud.
At a glance
- What is it?
- A CPU-only streaming transcription stack that scores 3.8% CER on real broadcast Japanese audio against whisper-large-v3-turbo's 13.8% on the same clips, by picking a dedicated model per language instead of one general-purpose one. It adds two-pass refinement, speaker labels, live translation and an OBS overlay, all under a 2GB resident model cap.
- Who is it for?
- Adopt hayamimi when latency on a machine without a GPU matters more than breadth, when your audio is Japanese, Chinese, Cantonese, Korean or one of the two dozen European routes it names, and when a 100ms final line plus a 0.5s partial is the shape of latency you need. Stay away if you depend on hotword biasing on the Japanese tier, where the feature currently does nothing, or if your domain has sparse punctuation, where the opt-in 4-class model underperforms the default.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 16 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Five routes pick a specialist model, and roughly 1600 languages fall through to Omnilingual ASR
The central bet is that a single general-purpose model is the wrong tool for streaming. Instead of defaulting to Whisper and accepting its ceiling, hayamimi routes each utterance to whichever specialist model is best for that language, and the catalog has five routes: Japanese, Chinese, Korean, Cantonese, and English plus two dozen European languages each get a dedicated best-in-class model. Everything else falls back to Meta's Omnilingual ASR.
The routing is only affordable because of the runtime underneath. Every model runs quantized to INT8 ONNX through sherpa-onnx, with no PyTorch and no CUDA in the stack, which is what lets a six-core desktop CPU transcribe at 10 to 50 times realtime.
The accuracy claim is narrow and specific, and worth reading with the same care. On real broadcast Japanese audio, measured on 2026-09-01 with a head-dropout fix and CJK number normalization in the pipeline, the routing reaches 3.8% CER against whisper-large-v3-turbo's 13.8% on the same clips. A fuller comparison against Whisper variants, cloud STT APIs and other local models sits in docs/results/comparison.md, and that document is where the languages hayamimi loses are recorded. The README points at it rather than claiming uniform superiority.
That asymmetry is the honest shape of this project: a routing table in front of several strong models, with the wins concentrated where a specialist exists.
Partial text refreshes every 0.5s and the final line lands about 100ms after you stop
Early ear is the name for someone who picks things up fast, and the latency targets are the design goal rather than a side effect. In-progress draft text updates roughly every half second while you are still speaking, so a partial is on screen before the sentence ends. A finalized line typically lands about 100 milliseconds after you stop talking.
That figure is a Japanese number, and the documentation says so: docs/design/goals.md holds the targets for the other languages rather than implying 100ms across the catalog. Treat it as the ja route's number and read the goals file before quoting it for anything else.
The partial/final split is what the OBS overlay is built around. At http://localhost:8833/ the confirmed line and the in-progress line are separate rows, and appending ?show=final or ?show=partial to that URL renders only one of them, so a streamer can place and style each as its own OBS browser source instead of trying to crop one caption box into two behaviours.
Everything above runs without a cloud call. No audio leaves the machine, which is the other reason the project is framed around CPU-only operation rather than around a small local fallback mode.
After 2s of silence recent utterances get re-decoded: 15.5% CER down to 12.0%
Two-pass refinement trades a little latency for accuracy on the lines that matter. Once two seconds of silence have passed, recent utterances are batch re-decoded, producing a higher-accuracy clean transcript alongside the live one. On real broadcast Japanese audio that takes CER from 15.5% to 12.0%.
The mechanism is batch re-decoding rather than a better model, which is the interesting part. The live path is optimised for a line that lands in 100 milliseconds; the refine pass is free to spend compute on a few seconds of audio at once, and it is that batch context that produces the improvement. The dashboard keeps the two apart, with a second column filling in refined text as it lands, so a viewer watching live sees the fast transcript and a reader later sees the better one.
The same re-decode step carries a second job, described in the speaker labelling section below, where it is what turns rough live labels into stable ones.
Documents cited for these numbers are docs/results/scorecard.md for the main comparison and docs/design/goals.md for the per-language latency targets.
Speaker labels are remapped after refinement, which is where DER 25.7% drops to 13.9%
Diarization here runs twice, and the second pass rewrites the first pass's answers rather than extending them. With --speakers, each utterance is tagged S1, S2 and onward live, using CAM++ nearest-centroid matching. That live pass is coarse: mean DER of 25.7% across five AMI meetings.
When the refine pass runs, it re-diarizes each speaker group with pyannote segmentation-3.0 and remaps the resulting clusters onto the same S{n} labels already shown on screen. That is what takes mean DER to 13.9% on the same five meetings.
Keeping the label set stable across both passes is the design constraint. If the refine pass invented its own numbering, every transcript a viewer had already read would point at the wrong speaker. The labels are therefore an interface, and remapping is the mechanism that keeps it one.
Two limits are worth stating. DER near 14% is not a small number in a meeting transcript, so speaker labels are an aid rather than a transcript. And the measured figures come from five AMI meetings, which is a narrow base for a claim about meetings generally.
--serve puts a dashboard, an OBS overlay and a transcript history on port 8833
The serving mode starts a local HTTP server with three views. The dashboard at /dashboard is the dense one: a partial-text strip for in-progress speech, a finals feed with language badges, speaker chips, per-line latency, inline translations under each line, and a second column for refined text as it arrives. The root path is the minimal OBS overlay described earlier, and /transcript is plain scrolling history.
Behind those pages is an EventHub that every stage of the pipeline publishes to, covering finals, translations, model loads, warnings and session summaries. An embedding application can listen to it directly, so the dashboard is one consumer of the event stream rather than the owner of it.
The same mode exposes runtime control over HTTP: GET and POST on /config change language, translation and VAD settings, POST /reset clears a session, and /replacements and /itn_overrides swap the find/replace and ITN dictionaries. In-process the same work is done by RoutedASR.set_replacements() and set_itn_overrides(). None of it requires restarting the process, which is what makes dictionary fixes practical mid-session rather than a restart ritual.
The dictionaries themselves are worth understanding: --replace is post-hoc find and replace and works everywhere, while --hotwords biases decoding toward proper nouns and currently has no effect on the Japanese tier.
The WebSocket ingest wants 16 kHz mono PCM, and only one client may produce audio
Network input exists so the microphone does not have to sit next to the machine doing the work. With --input ws, a phone or a stackchan-class ESP32 board streams mic audio over the LAN into the same pipeline, including the dashboard and overlay.
.venv/Scripts/python scripts/realtime_transcribe.py --input ws --serveThe ingest endpoint is ws://<host>:8766/ingest, and the handshake is explicit: connect to /ingest, send one JSON text frame describing the stream, then send raw audio as binary frames.
{"sr": 16000, "format": "pcm_s16le", "channels": 1}The server resamples anything that is not already 16 kHz, and replies with the same partial, final, translation and refine JSON events the dashboard's SSE stream carries, so a client can render its own subtitles rather than scraping the page. Only one audio-producing client is accepted at a time, and scripts/ws_mic_client.py is the dependency provided for sending from a machine.
Local capture has its own branch. --input speaker transcribes whatever the PC is playing, using WASAPI loopback on Windows only with no Stereo Mix configuration needed, and --input mix sums that with the microphone into one stream, which is how a meeting gets onto a single transcript.
The opt-in tiers cost resident RAM, and the pins were verified on Windows 11 with Python 3.11
Two optional tiers trade memory for accuracy, and both are off by default. --en-tier v2 swaps the default multilingual Parakeet v3 for an English-only Parakeet TDT v2, which needs download_models.py --en-parakeet-v2 and adds roughly 660MB resident, taking measured FLEURS English WER from 10.0% to 6.6%. --punct-model 4class swaps the default BERT restorer for one model predicting 、, 。, ? and ! directly, which adds real question and exclamation support the default lacks at about 37MB int8 and 4.6ms per line, but it is trained on dense web text and underperforms the default on sparse-punctuation domains such as TV captions.
Memory is bounded rather than hoped for. LRU model eviction keeps resident models under a configurable cap, with a default of under 2GB total, which is what lets a five-route catalog and a BERT restorer coexist on a modest desktop.
The dependency pins are exact rather than ranged, and the file says where they came from: numpy 2.4.6, sherpa-onnx 1.13.6, onnxruntime 1.29.0, ctranslate2 4.8.1, sentencepiece 0.2.2, fugashi 1.5.2, unidic-lite 1.0.8, soundcard 0.4.6 on Windows for the loopback path, and huggingface_hub 1.28.0, all verified on Windows 11 with Python 3.11 on CPU. kiwipiepy is listed as optional because it fixes token-spaced Korean output from SenseVoice and is LGPL-2.1-or-later, with a note that it can be commented out if that matters for your distribution; licence details for each package are in THIRD_PARTY_NOTICES.md. Installation steps are not shown in the portion of the README available here, so requirements.txt and that README are where to start.
Editorial conclusion
Adopt hayamimi when latency on a machine without a GPU matters more than breadth, when your audio is Japanese, Chinese, Cantonese, Korean or one of the two dozen European routes it names, and when a 100ms final line plus a 0.5s partial is the shape of latency you need. Stay away if you depend on hotword biasing on the Japanese tier, where the feature currently does nothing, or if your domain has sparse punctuation, where the opt-in 4-class model underperforms the default. Verify three things first: which route your language actually takes, including the Omnilingual ASR fallback covering the other roughly 1600, whether the Windows 11 pins in requirements.txt hold on your platform, and what THIRD_PARTY_NOTICES.md says about the packages it pulls in, including the optional LGPL kiwipiepy.
Frequently asked questions
Does hayamimi need a GPU or a cloud speech API?
No. Every model runs as a quantized INT8 ONNX model through sherpa-onnx, with no PyTorch and no CUDA in the stack, and no audio leaves the machine. Resident models are kept under 2GB by default through LRU eviction.
How accurate is hayamimi on Japanese broadcast audio?
It reaches 3.8% CER on real broadcast Japanese audio, remeasured on 2026-09-01 with a head-dropout fix and CJK number normalization, against 13.8% for whisper-large-v3-turbo on the same clips. The wider comparison, including languages hayamimi loses, is in docs/results/comparison.md.
What does hayamimi do for languages outside Japanese, Chinese, Korean, Cantonese and English?
The catalog has five routes, with Japanese, Chinese, Korean, Cantonese and English plus two dozen European languages each on a dedicated model. Everything else, roughly 1600 languages, falls back to Meta's Omnilingual ASR.
How does hayamimi label different speakers in one recording?
With --speakers it tags utterances S1, S2 and onward live using CAM++ nearest-centroid matching, then the refine pass re-diarizes each group with pyannote segmentation-3.0 and remaps the clusters onto the same labels, taking mean DER from 25.7% to 13.9% on five AMI meetings.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/oboroge0-hayamimi)