Model or dataset
CheshireCC/faster-whisper-GUI avatar
CheshireCC/faster-whisper-GUI

faster-whisper-GUI: a PySide6 desktop front end for faster-whisper and WhisperX

faster_whisper GUI with PySide6

3,003 stars173 forksPythonAGPL-3.0

At a glance

What is it?
CheshireCC/faster-whisper-GUI wraps faster-whisper, WhisperX, Silero VAD and Demucs in a Qt desktop application that writes srt, txt, smi, vtt and lrc files. It is aimed at people who want the full parameter set exposed without writing Python.
Who is it for?
Adopt it if you want faster-whisper's parameters, Silero VAD and WhisperX alignment exposed in a desktop window and you accept AGPL-3.0 obligations. Skip it if you need a headless server or a packaged installer, because the repository ships a requirements.txt and a FasterWhisperGUI.py entry point rather than a release binary.
Can I use it commercially?
Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
Is it still maintained?
Yes. The repository last received commits 37 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What faster-whisper-GUI adds on top of the faster-whisper library

faster-whisper is a Python library. Using it means writing a script, choosing a model size, passing a VAD filter, and then converting the returned segments into a subtitle file yourself. faster-whisper-GUI exists to remove that step. The README describes it as "a GUI software of faster-whisper" and lists the concrete outputs: transcription of audio or video files to srt, txt, smi, vtt or lrc. It also exposes, in the README's wording, "all paraments of VAD-model and whisper-model", which is the part most thin wrappers skip.

The audience is narrow and identifiable. It is for someone who has a folder of recordings, wants word-level timestamps for karaoke-style lyrics, and does not want to write a loop over files. The README shows a batch process screen and a file list screen, so multi-file runs are a first-class feature rather than an afterthought. If you are building a service that transcribes uploads on demand, this is the wrong shape of tool: it is a desktop application, not a daemon.

How the pipeline is assembled: models, VAD, Demucs and WhisperX

The repository layout tells you most of the architecture. FasterWhisperGUI.py is the entry point. faster_whisper_GUI/ holds the application package, config/ holds configuration, whisperx/ holds the WhisperX integration, and resource/ holds UI assets. A separate translate_zh_to_en.pro and an en.ts file indicate Qt translation files, which is consistent with the README's UI language screenshot.

At runtime the data flow is a chain of independent stages. A file is loaded, optionally passed through Demucs for vocal separation (the README labels this AVE, audio/vocal extraction, and links to UVR and Demucs-Gui as alternatives for that job). Silero VAD then trims silence before the whisper model runs. The faster-whisper model produces segments and, where enabled, word-level timestamps. WhisperX is offered as a separate path, which matters because WhisperX adds alignment and diarization-style workflows that plain faster-whisper does not perform. Finally the result is written out in one of the five subtitle formats, and the README notes that word-level timestamps are what make the karaoke lyric output work in VTT, LRC and SMI.

The Demucs stage is the one worth thinking about. It is a second neural model in the pipeline, so a transcription job can end up loading two models and using the GPU twice. The README itself points to UVR and Demucs-Gui as "more and better AVE" options, which reads as an acknowledgement that separation is not this project's strongest area.

Installing faster-whisper-GUI from requirements.txt on Windows or Linux

There is no packaged installer or release binary in the repository listing. Installation means a Python environment plus the dependencies in requirements.txt. Note that the file pins torch==1.13.1+cu117 and torchaudio==0.13.1+cu117, which are CUDA 11.7 builds; on a machine with a different CUDA runtime or no NVIDIA GPU, those lines will need attention before anything runs. The same file pins faster-whisper==0.10.0, pyside6 > 6.5.0 and pyside6-fluent-widgets>=1.3.2.

Create an environment and install the dependencies:

bash
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

On Windows the activation line is .venv\Scripts\activate instead. ffmpeg is required by ffmpeg-python and pyAV, and the README does not bundle it, so it has to be on PATH separately.

Start the application from the repository root:

bash
python FasterWhisperGUI.py

What you should see is the main window with the file list and the model controls. Before transcribing, a model has to exist locally. The README gives two routes: download from the Hugging Face search page for faster-whisper models, or use the software's own download and convert screens. It also links a converted large-v3 float32 build at CheshireCC/faster-whisper-large-v3-float32 on Hugging Face and a Baidu netdisk mirror. A minimal first run is: load one short audio file, pick a small model, leave VAD on, and export to srt to confirm the toolchain works before pointing it at a long recording.

Where faster-whisper-GUI gets in the way

The dependency pins are the first real obstacle. torch==1.13.1+cu117 is an old CUDA 11.7 build, and requirements.txt also carries nuitka<=1.8.6 and joblib==1.2.0 as exact or capped versions. Installing into an environment that already has a newer PyTorch will conflict. This is not a packaging detail you can ignore; it decides whether the install succeeds at all.

The second limitation is scope. The README documents a desktop workflow with screenshots for loading, downloading and converting models, batch processing, file filtering, editing timestamps and settings. It does not document a command line interface, a REST API, a Docker image, or a headless mode. If your goal is to transcribe files on a server or inside a container pipeline, this project gives you nothing to call. Related searches for a Docker setup exist, but the repository listing contains no Dockerfile and the README does not describe one.

The third is maintenance cadence. The last push to main was on 2026-08-25, but the most recent release is 0.8.5 from 2024-12-08. Source activity and tagged releases are moving at different speeds, so anyone depending on a stable version number should expect to track main rather than wait for tags. The README also carries a user agreement section listing prohibited uses; it is a usage condition attached by the author, not a technical control, and it does not change what the AGPL-3.0 licence permits.

faster-whisper-GUI compared with Faster-Whisper-XXL and faster-whisper-server

The closest alternatives people search for are Faster-Whisper-XXL and Fedirz/faster-whisper-server, and the difference is architectural rather than cosmetic.

Faster-Whisper-XXL is a standalone executable distribution. The practical difference is that you download one binary and run it, with no Python environment and no torch pin to reconcile. faster-whisper-GUI takes the opposite route: it is a Python application installed from requirements.txt, which means you inherit the CUDA 11.7 torch pins and the ffmpeg dependency but also get a source tree you can modify. If your machine already has a working CUDA 11.7 PyTorch, the install cost is low. If it does not, the executable route avoids that problem entirely. The README does not compare itself to XXL and does not mention it.

Fedirz/faster-whisper-server is a server, not a desktop application. It exposes transcription over an API so other programs can call it. faster-whisper-GUI does the reverse: a human drives it through windows, and the output is a file on disk. Choosing between them is choosing between an integration point and an operator interface. The one place they overlap is batch work, and there the GUI's file list and filter screens are the more direct tool.

Licence and the cost of keeping it running

The repository is licensed AGPL-3.0. For an individual running transcriptions locally, that is unremarkable. For a company that wants to embed this code in a product or expose a modified version over a network, the AGPL's source-availability condition applies to the modified work. This article cannot tell you what your obligations are; a lawyer can. The practical point is that AGPL-3.0 is a different proposition from the MIT-licensed faster-whisper library underneath it, and mixing the two in one product needs a deliberate decision.

Upgrade cost is dominated by the torch pins. Moving to a newer faster-whisper or a newer PyTorch means editing requirements.txt and retesting model loading, VAD and the Demucs stage, because all of them sit on top of the same runtime. Because tags lag source, there is no clean upgrade boundary to aim at; you either follow main or freeze a commit. The README offers no migration notes or rollback instructions, so a freeze-and-pin strategy is the only documented-safe option.

Editorial conclusion

Adopt it if you want faster-whisper's parameters, Silero VAD and WhisperX alignment exposed in a desktop window and you accept AGPL-3.0 obligations. Skip it if you need a headless server or a packaged installer, because the repository ships a requirements.txt and a FasterWhisperGUI.py entry point rather than a release binary. Before committing, check that the torch==1.13.1+cu117 and torchaudio==0.13.1+cu117 pins in requirements.txt match your CUDA setup, and confirm the ffmpeg binary is on PATH.

Frequently asked questions

Which Whisper model is the fastest?

The README does not rank model sizes by speed. It links the Hugging Face faster-whisper model search and notes that models can be downloaded and converted inside the application, and that the large-v3 model is supported. Model choice is exposed as a parameter rather than prescribed.

What are the differences between Faster-Whisper and Whisper?

The README does not draw that comparison. It describes this project as a GUI software of faster-whisper, and separately lists WhisperX support, so both faster-whisper and WhisperX are available as backends inside the interface.

Does Faster-Whisper use GPU?

The README does not state GPU usage directly, but requirements.txt pins torch==1.13.1+cu117 and torchaudio==0.13.1+cu117, which are CUDA 11.7 builds, and CTranslate2>=3.21.0 is listed. Those pins imply a CUDA-capable setup is the expected configuration.

Official sources

  1. CheshireCC/faster-whisper-GUI on GitHub
  2. Issues
  3. License: AGPL-3.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/cheshirecc-faster-whisper-gui.svg)](https://hysenlabs.com/projects/cheshirecc-faster-whisper-gui)