# hyprwhspr: system wide dictation for Linux desktops

> A Python dictation daemon that captures audio on a global hotkey, runs it through one of six transcription backends, and pastes the result into whatever window you were already typing in.

**goodroot/hyprwhspr** —  Native speech-to-text for Linux - Fast, accurate, private, and hackable system-wide dictation

- Repository: https://github.com/goodroot/hyprwhspr
- Website: https://hyprwhspr.com
- Stars: 1,234 · Forks: 106
- Language: Python
- License: MIT
- Published: 2026-10-06 · Updated: 2026-10-06 · Language: en
- Canonical page: https://hysenlabs.com/projects/goodroot-hyprwhspr

## What hyprwhspr actually is

hyprwhspr is a system wide dictation service for Linux desktops. It binds a global hotkey, captures a microphone stream, hands the audio to a speech to text backend, and injects the resulting text into whatever window currently holds focus. The name nods at Hyprland, but the project is not exclusive to it. The README names Sway, Niri, KDE Plasma, GNOME and plain X11 sessions as supported, and the injection path splits explicitly between wl-clipboard and wtype on Wayland and xclip, xdotool and xprop on X11.

The part that sets it apart from the usual tray application is the backend list. Cohere Transcribe, Parakeet TDT V3, Whisper, Qwen3-ASR, a generic REST endpoint and a realtime WebSocket mode all sit behind one interface. That is a wider menu than most dictation tools offer, and it is the reason the repository is worth reading even if you never move off Whisper.

The project is young but moving. At the time of writing it sits at 1205 stars and 99 forks with 3 open issues, and the last push to main landed on 2026-09-17. Release v1.45.2 was published 2026-09-16, a day earlier, which suggests tags track real work rather than ceremony. The licence is MIT, so the code is fair game for anything you want to build on top of it.

## Installing without hand assembling dependencies

Two supported paths cover most people. Arch users get an AUR package, and everyone else runs a shell installer that clones the tree into ~/.local/share/hyprwhspr/src, installs distro dependencies, and drops into the same interactive setup.

On Arch:

```bash
# Install for stable
yay -S hyprwhspr

# Or install for bleeding edge
yay -S hyprwhspr-git
```

On Ubuntu, Debian, Fedora or openSUSE:

```bash
curl -fsSL https://hyprwhspr.com/install.sh | bash
```

There is a manual route too, which is the one to read if you want to understand what the installer does on your machine:

```bash
git clone https://github.com/goodroot/hyprwhspr.git
cd hyprwhspr

# Install dependencies for your distro
./scripts/install-deps.sh

# Run interactive setup
./bin/hyprwhspr setup
```

Note the entry point is a shell script in bin/, not a console script from a Python package index. Updates for managed installs go through the CLI itself rather than a reinstall.

```bash
# Managed release installs
hyprwhspr update
```

That single command wraps verified release archives, atomic activation, a recovery journal and automatic rollback when activation fails. `hyprwhspr install status` inspects the state and `hyprwhspr install repair` fixes a bad one. It is a level of lifecycle care you rarely see in a desktop utility, and it arrived in v1.44.0 for the non Arch distributions.

## What the setup walkthrough actually configures

Setup is interactive and the README says so plainly. It walks through seven things: pick a transcription backend, download the models you chose, configure the themed visualizer if you want one, configure bar integration for Waybar or Noctalia, install systemd user services, set permissions, and finally validate the installation.

Two details there matter more than they look. The systemd user units mean dictation starts with your session rather than after you open a terminal. The permissions step exists because the hotkey reader talks to evdev directly, and on Wayland that generally needs an input group membership that only applies after you log out and back in.

Because permissions are involved, the first useful dictation attempt usually happens on your second login. The README is honest about this and puts the log out step first in its first use section.

Setup is also idempotent, which the update notes call out explicitly. Re-running it after a backend change does not duplicate services or re-download models you already have, so `hyprwhspr setup auto --backend qwen3-asr` is a safe way to switch engines.

## Backends, models, and where the audio goes

Here is where the repository gets unusually well organised. Instead of one fat requirements file, there is a requirements file per backend: requirements-cohere-transcribe.txt, requirements-faster-whisper.txt, requirements-faster-whisper-cuda.txt, requirements-onnx-asr.txt, requirements-onnx-asr-gpu.txt, requirements-qwen3-asr.txt, requirements-rest.txt, requirements-realtime.txt, requirements-realtime-elevenlabs.txt and requirements-pywhispercpp.txt, plus files for the CLI, the visualizer, optional extras and tests. The heavy inference stacks never install unless you pick the backend that needs them.

The shared core in requirements.txt is small and tells you what the application itself does:

```text
sounddevice>=0.5.0
numpy>=1.26.0
soxr>=0.5.0
soundfile>=0.12.1
evdev>=1.9.0
pyperclip>=1.11.0
pyudev>=0.24.4
pulsectl>=24.12.0
rich>=14.0.0
jsonschema>=4.18.0
```

sounddevice and soundfile for capture, soxr for resampling, evdev and pyudev for the hotkey and device discovery, pulsectl for the optional audio ducking that lowers other playback while you record, rich for CLI output and jsonschema for the read only config validation added in v1.44.0. Nothing there pulls in a toolkit or a model runtime.

Accuracy and speed map onto hardware. A recent NVIDIA card uses CUDA where the backend supports it, AMD and Intel parts can go through Vulkan, and onnx-asr covers machines with no usable GPU at all. The README advertises in memory models as the default so nothing is reloaded per utterance, and the visualizer memory can be capped separately so it does not fight your local models for VRAM. Models themselves are managed through `hyprwhspr model` with download, list, status, unload and reload, so you can free a card without stopping the daemon.

Cloud is optional, not implicit. Cohere and the ElevenLabs realtime mode are in the list, but transcription stays local unless you configure one of them. Qwen3-ASR, added in v1.45.0, is the multilingual pick for Chinese, Japanese and Korean and runs through a pinned llama.cpp sidecar with a 1.7B or 0.6B model, Vulkan on all three GPU vendors and a CPU fallback. Its own release notes still call it experimental, with no streaming, no timestamps and no prompt conditioning.

## Getting text into the focused window

The default gesture is deliberately plain: press `Super+Alt+D`, hear a beep, talk, press the same combination again, and the text lands wherever you were typing. Recording modes in the configuration guide cover toggle, push to talk, auto and long form, and the long form one supports pause, think, resume and submit for people who dictate long passages.

Injection is configurable. The default path copies to the clipboard and pastes into the focused buffer; an auto enter option can follow the paste when you want the cursor to land on a fresh line. On Wayland that means wl-clipboard and wtype, on X11 it means the xclip and xdotool family, which is why the dependency script installs different packages per session type.

The feature I would steal first arrived in v1.44.0. The most recent prepared dictation is kept in memory so a failed paste can be retried:

```bash
hyprwhspr record copy-last
```

`paste-last` delivers it to the focused application, and `clear-last` forgets it. Only text is retained, never audio and never transcript history on disk, which keeps the convenience from quietly turning into a recording archive. The same command group gives external hotkey control with start, stop, toggle, cancel, capture and status, so a compositor binding can drive dictation without the application holding focus.

There is an escape hatch for audio files too. `hyprwhspr transcribe INPUT` takes a WAV or an MP3 and prints to stdout, with an `-o` flag to write a file. That turns a background service into a batch tool.

## Status bar, overlay, and desktop quirks

Integration is where a dictation tool shows whether it actually lives on your machine. hyprwhspr exposes a status for Waybar and one for Noctalia, both managed from the CLI with `hyprwhspr waybar` and `hyprwhspr noctalia`. Recording state is visible without switching windows.

The visualizer is a separate concern and a separate dependency file. PyGObject and PyCairo live in requirements-visualizer.txt and are installed on a best effort basis, with a comment in requirements.txt explaining why: a GUI build failure, such as a missing girepository package on Ubuntu 24.04, must never abort the install of the core runtime. That is a small engineering decision with a large effect on whether an install completes on a distro you did not test.

Compositor behaviour still varies and the README is upfront about it. On layer shell compositors, Hyprland, Sway, niri and KDE, you get an animated overlay. On Noctalia and Omarchy it matches the live shell theme. On GNOME and Mutter you may need extra work, because those compositors do not hand out the same overlay surface. Expect the visualizer to be the first thing that needs tuning on GNOME.

A few more knobs exist for people who care about them: audio ducking through pulsectl so music steps down while you record, word overrides so names and jargon survive dictation, custom prompts and hotkeys, and translation from a non English source language to English output from a single config entry.

## The command line surface as an API

Because every action in the README has a subcommand, the tool is scriptable without the GUI. Configuration reads and writes through `hyprwhspr config` with show, show all, edit and secondary shortcut. `hyprwhspr model` handles downloads and VRAM. `hyprwhspr record` drives capture. `hyprwhspr status` reports overall health and `hyprwhspr validate` checks the installation.

Diagnostics are unusually thorough for a desktop app. `hyprwhspr install status` shows managed release state, `hyprwhspr model status` reports what is loaded and how much memory it holds, `hyprwhspr mic-osd` toggles the overlay, and `hyprwhspr systemd` manages the user units. `hyprwhspr keyboard` lists and tests input devices, which is the fastest way to find out that your hotkey never registered because of an evdev permission problem rather than a binding typo.

Two commands are worth reading as design intent. `hyprwhspr test` runs microphone and transcription end to end, so you can separate a capture problem from a model problem in one shot. And `hyprwhspr uninstall` keeps your settings, credentials and models, with `--purge` to take the personal recordings too and `--keep-models` if you only want the files gone. Uninstaller behaviour that distinguishes config from content is a detail most projects skip.

## Where the project is still thin

hyprwhspr is a single project with a small public surface and 3 open issues against 1205 stars, which usually means either very responsive maintenance or very few reports. The release cadence suggests the former, but the risk profile is still that of one maintainer.

Documentation is real but concentrated. The README is a long feature list, and the detail lives in docs/, with separate guides for configuration, managed installation and the themed visualizer. The repository also carries an AGENTS.md for coding agents, a tests directory, a website directory and a contrib directory, so the project is set up for contributors rather than being a single script someone runs from a dotfiles repo.

The honest limitations are these. GNOME and Mutter need manual work for the overlay. Qwen3-ASR has no streaming or timestamps yet. Cloud backends exist, so verify your configuration if privacy is the reason you chose this tool. And the Python range is 3.11 through 3.14, which is broad but excludes anything older.

## Conclusion

hyprwhspr is worth installing on Linux if dictating beats typing and you want the audio to stay on your machine. Pick a backend during `hyprwhspr setup`, let it create the systemd user services, and log out once so the evdev permissions apply. Start with Whisper or onnx-asr on a machine with no GPU, switch to Qwen3-ASR if you dictate Chinese, Japanese or Korean, and set up the Waybar or Noctalia integration before you start depending on the hotkey.

## FAQ

### Does hyprwhspr send my voice recordings to a server?

No, not by default. Transcription runs locally through Parakeet, Whisper, onnx-asr or Qwen3-ASR, and the README states that transcription never leaves the machine unless you configure a cloud backend such as Cohere Transcribe or the ElevenLabs realtime mode.

### Which transcription backend should I choose in hyprwhspr?

Choose Whisper or onnx-asr without a GPU, Parakeet TDT V3 for speed on current hardware, and Qwen3-ASR for Chinese, Japanese or Korean. The backend-specific requirements files mean only the inference stack you actually select gets installed.

### Which Linux desktops does hyprwhspr work on?

It targets Linux with systemd and needs either a Wayland or an X11 session. Hyprland, Sway, niri and KDE Plasma get the animated overlay directly, Noctalia and Omarchy get theme matching, and GNOME and Mutter may need manual adjustment.

### Why does the hyprwhspr hotkey not work right after setup?

Setup installs systemd user services and sets evdev permissions, but group membership for input devices only applies to new sessions. Log out and back in, then confirm the keyboard is readable with the keyboard list and test command.

## Sources

- [goodroot/hyprwhspr on GitHub](https://github.com/goodroot/hyprwhspr)
- [License: MIT](https://github.com/goodroot/hyprwhspr/blob/main/LICENSE)
- [Project website](https://hyprwhspr.com)
- [README](https://github.com/goodroot/hyprwhspr/blob/main/README.md)
- [Releases](https://github.com/goodroot/hyprwhspr/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/goodroot-hyprwhspr
