Open-source project
Sharrnah/whispering-ui avatar
Sharrnah/whispering-ui

whispering-ui is a Go window over a Python backend it downloads on first launch

Native UI for the Whispering Tiger project - https://github.com/Sharrnah/whispering (live transcription / translation)

329 stars24 forksGoMIT

At a glance

What is it?
The native front end for Whispering Tiger, a Windows and Linux transcription, translation, TTS and OCR application. The Go module holds no model code: it renders a window, moves audio, and fetches the Python engine the first time you start a profile.
Who is it for?
whispering-ui is a reasonable front end if your transcription runs are local and you want results pushed into VRChat or a browser overlay rather than into a word processor, since the UI covers device selection, model and precision choices, routing and configuration save and load, and it updates itself.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly Go, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The UI framework is a personal Fyne fork, not upstream Fyne

The go.mod for module whispering-tiger-ui declares go 1.26.0 with toolchain go1.26.5, then immediately overrides the UI toolkit. The require block asks for fyne.io/fyne/v2 v2.8.1, and a replace directive above it points that path at github.com/Sharrnah/fyne/v2 at a pseudo-version, v2.0.0-20260913230347-4912b1233c4e, which encodes a commit made on 2026-09-13. A comment above the directive links to a branch path ending in feature/setting-preferred-language, so the fork exists to carry a preferred-language setting upstream has not taken. Two consequences follow. Anyone reading the require line alone will believe the project builds on released Fyne, and it does not; a build from source needs the fork to be reachable. And because the fork's version string starts at v2.0.0 while the upstream line it stands in for is at v2.8.1, no dependency scanner can tell from version numbers which toolkit is actually compiled in.

The ZIP ships the window, not the engine that runs the models

Download the matching Windows or Linux ZIP, extract it to a writable local folder with enough space for the backend and models, and do not run it from inside the archive. That instruction is the whole deployment story, and it is telling: the space requirement names two things you do not get in the ZIP. On Windows step 2 of the install is to run Whispering Tiger.exe and then accept the backend platform download when prompted. On Linux step 3 repeats it, accept the Linux backend download, before a profile can be created. The repository description and the opening line of the documentation both say the same thing from the other direction: this repository holds the native Go and Fyne interface, and the Python backend at a separate repository runs the AI models. The Go dependency list is consistent with a front end. Audio in comes from gen2brain/malgo, a binding to miniaudio, and audio out from ebitengine/oto. Neither knows what a model is. Everything that does gets fetched at runtime.

The Linux build wants x86-64, glibc 2.36 and an OpenGL X11 session

sh
chmod +x whispering-tiger-linux-amd64
./whispering-tiger-linux-amd64

The Linux download is specified as x86-64 with glibc 2.36 or newer, an OpenGL-capable X11 or XWayland desktop, and either PulseAudio or the PulseAudio compatibility service inside PipeWire. That rules out 32-bit machines, ARM hardware and distributions still carrying an older glibc, and it routes Wayland through XWayland rather than natively. Run it as your normal desktop user, not as root, which is why the two lines above are a chmod and an invocation rather than anything with sudo. Desktop audio on Linux comes from a monitor source selected through PulseAudio or PipeWire, and on Windows the equivalent is a WASAPI loopback device, so the same feature takes two different routes depending on the platform. The packaged CUDA backend carries its own CUDA runtime libraries inside the archive, but it still requires a compatible NVIDIA driver on the machine. Where NVIDIA is absent, audio.cpp offers CPU or Vulkan for speech to text and speech synthesis.

Port 5000 on 127.0.0.1 is the single-instance ceiling

Step 2 of setup is the only numeric configuration in the profile form: Websocket IP + Port, with 127.0.0.1 and 5000 as the values you can keep. The note under it says what those defaults are for, and it is not the general case. They are useful if you want to run multiple instances, or if you want the backend platform on a separate PC. The second sentence is the catch: to run multiple instances you have to change the port for each one, because the default collides with the first. So a second window is a manual edit per instance rather than a checkbox, and the IP field is what turns the backend into a network service pointed at another machine, which also means the transcription traffic leaves the loopback interface. Everything else in the profile form is device and model selection: audio input and output with a Test button that shows both level bars moving, optional push to talk, a memory estimate in the lower right corner, a compute device per task, and the model with its size and precision.

Push to talk is implemented by zeroing two speech-detection fields

Push to talk is optional and configured in one line of the profile form: click into the field and press the keys you want to use. Each key has to be pressed separately while configuring, because a chord would register as one unknown combination, and at runtime every configured key must be pressed at the same time for the shortcut to fire. The more interesting half is what happens to speech detection when you use it. The documentation does not offer a toggle; it says that to disable autodetection of speech and use push to talk only, you set Speech volume Level and Speech pause detection to 0. Two thresholds, both driven to zero. The consequence is that autodetection is not disabled by a switch that remembers your previous values, so a profile that switches between hands-free and held-key operation carries its own numeric state, and nothing in the form suggests the setting is shared across profiles. Related to the same field, the memory figure is labelled in the interface as an estimate of (V-)RAM need and is described in the documentation as rough and variable.

The memory figure in the corner is hardware introspection, not a model calculation

The lower right corner of the window shows estimated memory consumption, and the documentation is careful about it twice. It is a rough estimate that can vary, and its stated purpose is to give an idea of how much (V-)RAM the selected AI models and options need. Where the number comes from is in the dependency list: github.com/jaypipes/ghw, a hardware discovery library, is a direct dependency rather than an indirect one, so the estimate is built from what the machine reports about itself and not from the size of the model files you picked. That explains the hedge. Model selection is documented separately and honestly in the same breath: larger models usually need more memory and may be slower, and language coverage and supported tasks depend on the model. Precision is the third lever, where lower precision can cut memory use, the available choices depend on the model and runtime, and a GGUF package's precision describes how its stored weights are encoded rather than how fast anything runs.

Sentry and mDNS sit in the direct dependency list of an offline application

The opening claim is that processing runs locally after the models are downloaded, and that online services are optional plugins. Two direct dependencies describe network-facing behaviour that the feature list does not mention. github.com/getsentry/sentry-go at v0.49.0 is a crash reporting client, which means events can leave the machine independently of whether any plugin is installed. github.com/lunarhue/metallic-flock-zeroconf, pinned at a pseudo-version from 2026-06-25, is a zero-configuration discovery service, which means the application can announce itself on the local network and look for others there. A third direct dependency, github.com/hashicorp/go-cleanhttp v0.5.2, is the transport used to fetch plugin sources over HTTPS. The fourth matters for output routing: github.com/hypebeast/go-osc is pinned at a pseudo-version from 2022 for the VRChat side of things, while gorilla/websocket v1.5.3 handles browser overlays. That split is why the documentation mentions Websockets and OSC as two separate destinations rather than one.

Twenty capitalised directories sit at the root of a Go module

The top-level listing puts Audio/, BuildTools/, CustomWidget/, Fields/, LocalPlugins/, Logging/, ModelDownloader/, OscClient/, Pages/, ProfileForm/, Profiles/, RemoteAudio/, RemoteAudioView/, Resources/, RuntimeBackend/, SendMessageChannel/, Settings/, UpdateUtility/, Updater/, Utilities/ and Websocket/ alongside main.go, theme.go, go.mod and cmd/. Go convention is lowercase package directories with an import path that matches, and here every package name is capitalised, so the import list reads as a set of proper nouns. Three of them name the same subsystem from different ends, Profiles/ and ProfileForm/ for the profile list and its editor, and Settings/ and Fields/ for the same reason. The non-Go root files are just as telling: build.bat and bundle.bat are Windows batch scripts, .drone.yml is the continuous integration file, FyneApp.toml is the toolkit's application metadata, and doc/ holds the hardware, audio and realtime pages the feature list links to. The repository describes the interface layer in detail and leaves the plugin list to a separate repository.

Editorial conclusion

whispering-ui is a reasonable front end if your transcription runs are local and you want results pushed into VRChat or a browser overlay rather than into a word processor, since the UI covers device selection, model and precision choices, routing and configuration save and load, and it updates itself. It is a poor fit if you need one binary with no follow-up download, if your Linux system is older than glibc 2.36 or not x86-64, or if you need a second instance without editing a port. Before you commit, read the hardware and runtimes page for the model and runtime combinations, decide where you will extract the ZIP given that models and the backend land in that folder, and check whether the crash reporter and the zero-configuration discovery service in the dependency list are acceptable in your environment.

Frequently asked questions

Does Whispering Tiger send audio to a server?

Processing runs locally after the models are downloaded, so transcription, translation, text to speech and OCR happen on the machine. Online services exist only as optional plugins, though the manifest does include a crash reporting client and a local network discovery library as direct dependencies.

What are the system requirements for whispering-ui on Linux?

The Linux download is for x86-64 with glibc 2.36 or newer, an OpenGL-capable X11 or XWayland desktop, and either PulseAudio or the PulseAudio compatibility service inside PipeWire. It should be run as your normal desktop user rather than as root.

Can I run two instances of Whispering Tiger at once?

Yes, but not on the defaults. Websocket IP and Port start at 127.0.0.1 and 5000, and the documentation says you need to change the port for each instance. The default IP and port are otherwise kept for running one instance or for putting the backend platform on a separate PC.

Where does Whispering Tiger put the AI models on Linux?

In the folder you extracted the ZIP to, which the installation step says must be a writable local folder with enough space for the backend and models, and the application must not be run from inside the archive. The backend itself is a separate download accepted on first run.

How do I turn off hands-free speech detection in Whispering Tiger?

Set Speech volume Level and Speech pause detection to 0, which disables autodetection of speech so that only push to talk triggers it. Each push to talk key has to be pressed separately while configuring, and all of them must be held together at runtime.

Official sources

  1. License: MIT
  2. Project website
  3. README
  4. Releases
  5. Sharrnah/whispering-ui on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/sharrnah-whispering-ui.svg)](https://hysenlabs.com/projects/sharrnah-whispering-ui)