whisper-diarization: Speaker-Labeled Transcription with Whisper and NeMo
Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper
At a glance
- What is it?
- whisper-diarization combines OpenAI's Whisper ASR with NVIDIA NeMo's speaker diarization pipeline to produce transcriptions that identify which speaker said each sentence. It is aimed at developers and researchers who need speaker-attributed subtitles or meeting transcripts from audio files, without building the multi-stage alignment and embedding pipeline themselves.
- Who is it for?
- whisper-diarization is the right tool for a developer who needs speaker-attributed transcripts and is comfortable running a Python pipeline with GPU dependencies. It is not production-ready software in the sense that the parallel mode is experimental, overlapping speakers are not handled, and NeMo is a heavy dependency.
- Can I use it commercially?
- Yes. BSD-2-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 46 days ago.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The Problem: ASR Without Speaker Identity
Automatic speech recognition tools like Whisper produce a transcript of what was said but not who said it. For a conversation, interview, meeting recording, or podcast with multiple speakers, a flat transcript is difficult to follow because the reader cannot tell where one speaker ends and another begins. Manually labeling speakers in a long recording is time-consuming, and building the combination of voice activity detection, speaker embedding, and timestamp alignment to do it automatically requires integrating several research-grade components.
whisper-diarization wraps those components into a single pipeline. The README describes the full sequence: vocals are first extracted from the audio using a source separation model to improve embedding accuracy, Whisper generates the transcription, ctc-forced-aligner corrects and aligns the timestamps, MarbleNet performs voice activity detection and segmentation to remove silences, TitaNet extracts speaker embeddings to identify each segment's speaker, and a punctuation model performs a final realignment pass to compensate for minor time shifts. The result is a transcription where each sentence or segment is labeled with a speaker identifier.
Installation Requirements
whisper-diarization requires Python 3.10 or newer. Python 3.9 can be used but requires installing the requirements one by one rather than using the provided requirements file. The README lists two system-level prerequisites that must be installed before the Python packages.
Cython must be installed first:
pip install cythonAlternatively, on Debian or Ubuntu:
sudo apt update && sudo apt install cython3FFMPEG is also required and can be installed through the system package manager on most platforms:
# on Ubuntu or Debian
sudo apt update && sudo apt install ffmpeg
# on Arch Linux
sudo pacman -S ffmpeg
# on MacOS using Homebrew (https://brew.sh/)
brew install ffmpeg
# on Windows using Chocolatey (https://chocolatey.org/)
choco install ffmpeg
# on Windows using WinGet (https://github.com/microsoft/winget-cli)
winget install ffmpegWith those prerequisites in place, install the Python requirements:
pip install -c constraints.txt -r requirements.txtThe requirements.txt includes nemo_toolkit[asr] (version 2.5.0 or newer but below 3), faster-whisper (version 1.1.0 or newer), and three packages from the author's own GitHub forks: a modified Demucs for vocal separation, a deepmultilingualpunctuation fork, and ctc-forced-aligner.
Running the Pipeline
The main entry point is diarize.py, which accepts an audio file path and optional configuration flags:
python diarize.py -a AUDIO_FILE_NAMEThe command line options let the caller override the Whisper model, disable source separation, specify a language manually, and control batch size. The default Whisper model is medium.en, which is English-only. For non-English audio, pass a model that supports the target language and specify the language with --language.
The --device flag defaults to "cuda" when a CUDA GPU is available and falls back to CPU otherwise. CPU inference on NeMo is significantly slower than GPU inference, so practical use on long recordings requires a compatible GPU.
For systems with 10 gigabytes or more of VRAM, diarize_parallel.py runs the NeMo and Whisper stages concurrently rather than sequentially. The README describes this as experimental and notes that the result is the same as the sequential version because the two models do not depend on each other's output. The --batch-size option controls NeMo's batched inference; reducing it lowers memory use at the cost of throughput, and setting it to 0 disables batching entirely.
How the Pipeline Aligns Speakers to Words
The alignment step is the distinguishing part of this pipeline compared to a naive combination of Whisper and a speaker diarization model. Whisper's timestamp precision at the word level is limited; the ctc-forced-aligner step corrects this by running a CTC alignment pass on the Whisper output, producing tighter word-level timestamps. This matters because speaker attribution is assigned per word based on timestamp overlap with the speaker segments from TitaNet.
The source separation step with Demucs extracts vocals from the audio before passing it to NeMo. The README explains that this improves speaker embedding accuracy, which means better speaker identification when the original recording contains background music or noise. This step adds processing time and requires the modified Demucs fork listed in requirements.txt.
The final punctuation realignment uses a multilingual punctuation model to compensate for small time shifts that remain after the forced alignment step. The README lists this model as coming from deepmultilingualpunctuation, also installed from the author's fork. Together these steps produce a pipeline that the README describes as designed to minimize diarization error due to time shift.
Known Limitations and Incomplete Features
The README explicitly documents two current limitations. Overlapping speakers are not handled: when two speakers talk at the same time, the pipeline cannot separate them and will attribute the segment to one speaker or produce an error. The README suggests a possible approach (separating the overlapping audio and processing each speaker's channel individually) but notes this would require substantially more computation and has not been implemented.
A planned future improvement listed in the README is a maximum sentence length limit for SRT output, but this has not been added yet.
The Whisper and NeMo parameters are hardcoded in diarize.py and helpers.py. The README notes that CLI arguments to change them will be added later. Users who need to adjust NeMo's internal hyperparameters must currently edit the source files directly.
The requirements include packages installed from GitHub repository URLs rather than from PyPI. This means the installation process depends on network access to GitHub at install time, and reproducible environments require pinning commits in the requirements file rather than using version numbers.
Comparison with WhisperX
WhisperX is the most commonly cited alternative for speaker-attributed Whisper transcription. Both projects combine Whisper with forced alignment and speaker diarization, but they differ in their dependency choices and default behavior. WhisperX uses pyannote.audio for speaker diarization, which requires accepting pyannote's terms of use and obtaining an access token from Hugging Face. whisper-diarization uses NVIDIA NeMo, which has no gated model requirement but is a significantly larger dependency with its own CUDA version constraints.
For Windows users, NeMo's installation is more complex than pyannote's. For Linux environments with NVIDIA GPUs already set up for deep learning work, NeMo is a familiar dependency. The choice between the two pipelines often comes down to which dependency stack the team is already managing. Neither handles overlapping speakers well; both produce best results on clean recordings with non-overlapping speech.
Editorial conclusion
whisper-diarization is the right tool for a developer who needs speaker-attributed transcripts and is comfortable running a Python pipeline with GPU dependencies. It is not production-ready software in the sense that the parallel mode is experimental, overlapping speakers are not handled, and NeMo is a heavy dependency. Run the single-file diarize.py first on a short test file before committing to the pipeline for a large batch, because the error surface across FFMPEG, Cython, NeMo, and faster-whisper is wide enough that environment issues surface early.
Frequently asked questions
How does whisper-diarization differ from WhisperX?
whisper-diarization uses NVIDIA NeMo's TitaNet for speaker embeddings, while WhisperX uses pyannote.audio, which requires a gated Hugging Face model with terms-of-use acceptance. Both combine forced alignment with Whisper transcription. The choice depends mainly on which dependency stack is already in the project environment.
What GPU is required to run whisper-diarization at practical speed?
The README defaults the --device flag to CUDA when a GPU is available. The parallel mode (diarize_parallel.py) requires 10 gigabytes or more of VRAM. CPU inference is possible but the README does not document expected runtime on CPU for typical audio lengths.
Can whisper-diarization handle overlapping speakers?
No. The README lists overlapping speakers as a known limitation. The pipeline assigns each audio segment to one speaker and cannot separate two voices speaking simultaneously. The README describes a possible future approach using audio separation but notes it has not been implemented.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/mahmoudashraf97-whisper-diarization)