Buzz: offline Whisper transcription on your own machine
Buzz transcribes and translates audio offline on your personal computer. Powered by OpenAI's Whisper.
At a glance
- What is it?
- Buzz is a desktop app and Python package that runs OpenAI's Whisper locally to transcribe and translate audio, with GPU backends for CUDA, Apple Silicon and Vulkan. It is for people who want captions without uploading their recordings, and it is heavier than a hosted API.
- Who is it for?
- Buzz fits anyone who needs transcripts of private or large media files and can spare local disk and compute: install it from SourceForge, Flatpak, Snap or pip, then check whether your GPU path is actually active before transcribing an hour of audio. Skip it if you need a hosted API, a server-side batch pipeline, or an Intel Mac, since the README states the last version supporting Intel Macs is 1.4.5.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 7 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Buzz does that a cloud transcription API does not
The pitch is in the first line of the README: transcribe and translate audio offline on your personal computer. Every model runs locally, so the audio never leaves the machine. That matters for interview recordings, medical or legal material, and anything under an NDA where sending a file to a third-party endpoint is a non-starter.
The second difference is cost shape. A hosted API bills per minute of audio; Buzz bills in disk space and electricity. The pyproject.toml pulls in torch, torchaudio, faster-whisper, openai-whisper, transformers, accelerate, peft and bitsandbytes, plus ctranslate2 and nvidia-cudnn-cu12 on non-Darwin platforms. That is a large dependency tree, and the model weights are downloaded separately on first use.
The audience is therefore narrower than "anyone who wants captions". It is people who already have a GPU or an Apple Silicon Mac, who work with media files rather than a live API stream, and who would rather wait for a local job than pay per minute. The README also lists a command-line interface for scripting, a watch folder for automatic transcription of new files, live microphone transcription with a presentation window, speech separation before transcription, and speaker identification.
The backend layer: Whisper, faster-whisper and whisper.cpp side by side
Buzz is not a single inference path. The feature list names multiple whisper backend support: CUDA acceleration for Nvidia GPUs, Apple Silicon support on Macs, and Vulkan acceleration for Whisper.cpp on most GPUs including integrated ones. Those are three different runtimes with different model formats, and the repository layout confirms it: whisper.cpp, whisper_diarization, ctc_forced_aligner, deepmultilingualpunctuation and demucs_repo all sit at the top level as submodules or vendored directories.
The data flow is roughly: media file or microphone stream in, ffmpeg or sounddevice for decoding and capture, an optional speech-separation and diarization pass, then a Whisper-family model for the transcript, then post-processing for punctuation, timing and export to TXT, SRT or VTT. The pyproject.toml shows stable-ts and srt-equalizer in the dependency list, which is consistent with subtitle timing adjustment rather than raw token output.
Model choice is not limited to the original Whisper checkpoints. The README mentions multiple Transformer model family support via Huggingface whisper type, and the presence of peft, bitsandbytes and accelerate suggests fine-tuned or quantised checkpoints can be loaded. That flexibility is real, but it also means the quality of your output depends far more on which model you pick than on Buzz itself. Buzz is the shell around inference, not the source of accuracy.
Installing Buzz on Windows, macOS and Linux
The README points desktop users at SourceForge for both the macOS .dmg and the Windows installer. On Windows the app is not signed, so the README says you will get a warning during install and tells you to select More info, then Run anyway. On macOS, the README states Buzz now requires Apple silicon and that the last version supporting Intel Macs is 1.4.5.
Linux has three packaged routes. Flatpak:
flatpak install flathub io.github.chidiwilliams.BuzzSnap, which needs three system libraries first:
sudo apt-get install libportaudio2 libcanberra-gtk-module libcanberra-gtk3-module
sudo snap install buzzFor the PyPI route, the README says to install ffmpeg and to use a Python 3.12 environment, then:
pip install buzz-captions
python -m buzzNote the discrepancy: the README says Python 3.12, while pyproject.toml declares requires-python = ">=3.13,<3.14". If you install from PyPI and the interpreter does not match what the package metadata demands, pip will refuse or resolve to an older release. Check which of the two applies to the version you are pulling before you build a workflow around it.
For Nvidia GPU acceleration on a PyPI install, the README gives explicit torch and CUDA library pins:
pip3 install -U torch==2.8.0+cu129 torchaudio==2.8.0+cu129 --index-url https://download.pytorch.org/whl/cu129
pip3 install nvidia-cublas-cu12==12.9.1.4 nvidia-cuda-cupti-cu12==12.9.79 nvidia-cuda-runtime-cu12==12.9.79 --extra-index-url https://pypi.nvidia.comIf you are on Buzz 1.4.6 or later with an Nvidia GPU, the README describes a simpler path: go to Help, then About Buzz, and install CUDA Acceleration from there.
First run: pick a model, then a file, then an export format
After launching, the workflow the README implies is: import a file, a video, or a YouTube link, choose a model in preferences, run the transcription, then export. The screenshots in the repository show an import screen, a main screen, a preferences pane, a model preferences pane, a transcript view, a live recording view and a resize view, so the settings you need are split across at least two dialogs.
The model preference matters most. Larger Whisper checkpoints give better accuracy on accented or noisy speech and run slower; the README's mention of CUDA, Apple Silicon and Vulkan acceleration exists precisely because CPU-only inference on a large model is slow enough to be impractical for long recordings. If you have an Nvidia card, confirm the CUDA Acceleration install from Help, About Buzz actually completed before you judge the speed.
Export is the least surprising part: TXT for a plain transcript, SRT or VTT for subtitles. The README also lists a watch folder for automatic transcription of new files, which is the feature to reach for if you have a recurring pipeline rather than one-off files, and a CLI for scripting and automation if you would rather not touch the GUI at all.
Where Buzz gets in the way
The dependency list is the first real cost. torch, torchaudio, transformers, peft, bitsandbytes, faster-whisper and openai-whisper all land in the same environment, and the PyPI instructions pin specific CUDA builds. If you already have a Python environment with a different torch version, expect a conflict. The README's own advice, a dedicated environment, is not optional in practice.
The second limitation is hardware. Intel Macs are out as of the current release line. Integrated GPUs are only covered through the Vulkan path via Whisper.cpp, which the README describes as supporting most GPUs including integrated ones, but does not promise parity with CUDA on speed. If you are on a machine without a discrete GPU or Apple Silicon, Buzz will run, but the large models will be slow.
The third is that live transcription and file transcription are different workloads sharing one app. Live microphone capture goes through sounddevice and is bounded by how fast the model decodes in real time; a model that is fine for a five-minute file may fall behind on a continuous stream. The README does not document rollback or recovery behaviour for an interrupted live session, and it does not state memory requirements for any model size. Plan your model choice around your hardware rather than around the accuracy you would like.
Buzz versus running Whisper directly or through a hosted API
The obvious alternative is OpenAI's own Whisper repository, which Buzz is built on. Running whisper directly gives you a Python script and a model, and nothing else: no GUI, no watch folder, no export to SRT through a viewer, no live presentation window, no diarization, no plugin system. Buzz adds all of that on top, at the cost of a much larger dependency set and a Qt desktop application you have to install.
A hosted transcription API is the other direction. It removes the GPU requirement and the model download entirely, and it scales to hundreds of files without your laptop being involved. The trade-off is the one Buzz exists to avoid: your audio leaves your machine, and you pay per minute. For a one-off transcript of a public video, the API is simpler. For a folder of internal recordings, Buzz is the one that keeps the files local.
There is also a middle path the README gestures at without naming: the CLI. If you want Whisper's output in a script but not the Qt app, the command-line interface is the part worth evaluating, and the plugin system with AI summary generation and automated transcript resizing is the part that has no equivalent in plain Whisper.
Maintenance, licensing and what an upgrade actually costs
Buzz is MIT licensed, stated in the README badge and in pyproject.toml as license = { text = "MIT" }. That permits commercial and closed-source use, but it does not cover the models you download. Whisper checkpoints and any Hugging Face model you load carry their own licences, and the README does not enumerate them. If you are shipping transcripts commercially, check the model licence separately from the app licence.
The release cadence visible here is roughly every few months: v1.4.3 in January 2026, v1.4.4 in March, v1.4.5 in August. The last push to the repository was on 2026-09-19. Upgrades are not trivial, because the pinned dependencies move together: torch, torchaudio, ctranslate2 and the CUDA libraries are all version-locked in pyproject.toml, and the README's GPU instructions name exact versions. A minor Buzz release can require a matching torch rebuild. Budget for re-testing your transcription pipeline after each upgrade rather than assuming a drop-in replacement.
Editorial conclusion
Buzz fits anyone who needs transcripts of private or large media files and can spare local disk and compute: install it from SourceForge, Flatpak, Snap or pip, then check whether your GPU path is actually active before transcribing an hour of audio. Skip it if you need a hosted API, a server-side batch pipeline, or an Intel Mac, since the README states the last version supporting Intel Macs is 1.4.5. Verify two things first: that your Python environment is 3.13 for the PyPI route, and that ffmpeg is on PATH, because the README names it as a prerequisite.
Frequently asked questions
Is Buzz transcription free?
The application is MIT licensed and the README does not describe any paid tier or subscription. You still pay in hardware: a GPU or Apple Silicon Mac makes larger models practical, and model weights are downloaded separately.
How do I install Buzz?
On Windows and macOS the README points to SourceForge for the installer or .dmg. On Linux you can use Flatpak, Snap or AppImage, and there is a PyPI route with pip install buzz-captions followed by python -m buzz.
Is the Buzz app free?
Yes, the repository is MIT licensed and the README lists no paid edition. Note that the Windows installer is unsigned, so the README tells you to select More info, then Run anyway during installation.
What is the best free transcription app for Windows?
The README does not compare Buzz against other transcription apps, so no ranking can be given. What the README does state is that Buzz runs Whisper offline on your own machine and offers CUDA acceleration for Nvidia GPUs on Windows.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/chidiwilliams-buzz)