# pyannote.audio: what the speaker diarization toolkit actually does, and how to install it

> pyannote.audio is a PyTorch-based Python toolkit for speaker diarization. It ships pretrained pipelines for speech activity detection, speaker change detection, overlapped speech detection and speaker embedding, and it also fronts a paid hosted service. The open pipeline and the premium one are not the same product.

**pyannote/pyannote-audio** — Neural building blocks for speaker diarization: speech activity detection, speaker change detection, overlapped speech detection, speaker embedding 

- Repository: https://github.com/pyannote/pyannote-audio
- Website: https://www.pyannote.ai
- Stars: 10,597 · Forks: 1,109
- Language: Jupyter Notebook
- License: MIT
- Published: 2026-09-21 · Updated: 2026-09-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/pyannote-pyannote-audio

## The problem pyannote.audio solves, and who ends up using it

Given a recording with several people talking, speaker diarization answers a narrow question: who spoke when. The output is a list of turns, each with a start time, an end time and a speaker label. No transcript, no words, just segmentation and identity. pyannote.audio is a Python toolkit built on PyTorch that produces exactly that, and the repository describes it as neural building blocks for speaker diarization: speech activity detection, speaker change detection, overlapped speech detection and speaker embedding. Those four tasks are the components the pipelines assemble.

The audience is narrower than the topic suggests. This is a library for people writing Python who are comfortable with PyTorch and with downloading model weights. The README's own example imports torch, loads a pipeline by name from Hugging Face, optionally moves it to a CUDA device, and iterates over the result. If your workflow lives in a shell script or a no-code tool, the library is not aimed at you. The pyannoteAI hosted service is the path for that case, and it is a separate product with a separate key.

One detail worth noticing early: the repository classifiers list Development Status as Beta, and the package version series is 4.0.x. The pretrained pipelines are the stable part in practice; the API around them has moved between major versions, which is why the README keeps a legacy 3.1 pipeline in its benchmark table.

## How the pipeline is assembled: components, checkpoints and where inference runs

A pyannote.audio pipeline is not a single neural network. It is a composition of models, each loaded from the Hugging Face hub and wired together by the pyannote-pipeline package, which is a declared dependency. The four building blocks named in the repository description map onto stages: speech activity detection trims silence, speaker change detection proposes boundaries, speaker embedding turns each segment into a vector, and a clustering step assigns those vectors to speaker labels. Overlapped speech detection handles the case where two people talk at once, which is where segmentation-only systems tend to merge speakers.

The data flow is local by default. Audio is decoded through torchcodec, which is why ffmpeg must be installed on the machine, then the pipeline runs on CPU or on a CUDA device you select. The README's own comment on the community pipeline call says it runs locally. The premium pipeline uses the same Pipeline.from_pretrained entry point but the comment says it runs on pyannoteAI servers, and it takes a pyannoteAI API key rather than a Hugging Face token. That is the single most important design fact in the project: the same Python call site can mean local inference or a remote service, and the only thing distinguishing them is the model identifier string.

The benchmark table is worth reading as a statement of scope rather than as a scoreboard. It reports diarization error rate, lower is better, across corpora including AMI, CALLHOME, DIHARD 3, VoxConverse and REPERE. community-1 improves on the legacy 3.1 pipeline on most of those sets, though not all: REPERE phase2 goes from 7.9 to 8.9, a regression the table does not hide. The premium precision-2 column is consistently lower, and the separate speed table reports 31s per hour of audio for community-1 against 14s per hour for precision-2 on AMI IHM, measured self-hosted on an NVIDIA H100 80GB HBM3.

## Installing pyannote.audio and running a first diarization

The README gives a four-step setup. First, ffmpeg must be installed on the machine, because the torchcodec audio decoding library needs it. Second, install the package, with uv add pyannote.audio given as the recommended route or pip install pyannote.audio as the alternative. Third, accept the user conditions on the pyannote/speaker-diarization-community-1 model page. Fourth, create a Hugging Face access token.

```bash
pip install pyannote.audio
```

That installs the package and its dependencies, which include torch, torchaudio, torchcodec, lightning and the pyannote-core, pyannote-database, pyannote-metrics and pyannote-pipeline packages. Python 3.10 or later is required. Nothing downloads at this point; the model weights arrive on first use.

The README's usage example loads the community pipeline by name, passes the token, optionally moves the pipeline to a CUDA device, and iterates over the diarization output. Copy the identifier exactly, since a typo produces a hub lookup failure rather than a helpful error.

```python
import torch
from pyannote.audio import Pipeline
from pyannote.audio.pipelines.utils.hook import ProgressHook

pipeline = Pipeline.from_pretrained(
    "pyannote/speaker-diarization-community-1",
    token="HUGGINGFACE_ACCESS_TOKEN")

pipeline.to(torch.device("cuda"))

with ProgressHook() as hook:
    output = pipeline("audio.wav", hook=hook)
```

The ProgressHook context manager is optional but useful on a first run, because the initial call also fetches weights and the hook gives visible progress. The result exposes speaker_diarization, and iterating over it yields turn and speaker pairs. The README's expected output looks like start=0.2s stop=1.5s speaker_0, then speaker_1, then speaker_0 again. Speaker labels are generated per file and are not stable identities across recordings.

The premium path is shorter but different in kind. You create a pyannoteAI API key at dashboard.pyannote.ai, then load pyannote/speaker-diarization-precision-2 with that key as the token. The README notes free credits are available, and the output labels appear as SPEAKER_00 style strings rather than speaker_0.

## Where pyannote.audio is the wrong tool

The gating is the first real constraint. The community pipeline is not downloadable anonymously: you must accept user conditions on the model page and supply a Hugging Face token. That makes it awkward for fully offline or air-gapped deployment, and it means your build pipeline needs a secret. The README does not document an offline weight-mirroring procedure, so treat that as an open question rather than a solved one.

ffmpeg is a hard system dependency, not a Python one. pip will not install it for you, and a container built without it will fail at decode time rather than at import time. If you are packaging this into a minimal image, that is a real line in the Dockerfile.

Speaker labels are per-file. Diarization tells you that two segments in the same recording came from the same voice; it does not tell you that speaker_0 in Monday's call is the same person as speaker_1 in Tuesday's. Cross-recording identity is a separate problem, and the README points to voiceprinting as a pyannoteAI feature rather than a local one.

Finally, the two-pipeline arrangement is easy to misread. Anyone who copies the premium example without a pyannoteAI account, or who assumes the community example is the only one, will end up with the wrong mental model of where their audio goes. If your audio cannot leave your infrastructure, the model identifier in your code is the thing to audit, not a configuration flag.

## pyannote.audio against WhisperX and the hosted alternatives

The most common comparison is with WhisperX, and the difference is one of scope rather than accuracy. WhisperX is a transcription pipeline that bolts diarization onto word-level timestamps; the diarization segment is not the product, the transcript is. pyannote.audio is the reverse: it produces speaker turns and nothing else, and the README never mentions transcription. If you need text with speaker attribution, you need both, and you would typically use pyannote.audio's turns to label WhisperX's words. If you need only turns, pulling in a speech recognition stack is wasted work.

The other alternative is the hosted route, and pyannote.audio itself now spans both sides of it. precision-2 is the same vendor's paid service, reached through the same Pipeline.from_pretrained call. The trade is straightforward: community-1 runs on your hardware with no per-minute cost but needs a GPU to be fast, while precision-2 runs on pyannoteAI servers, is reported as 2.2x to 2.6x faster on the two corpora in the speed table, and scores lower error rates across the board. It also means your audio leaves your machine, and the README does not describe retention or data handling for that path.

A third option is the legacy speaker-diarization-3.1 pipeline, still listed and still loadable. It is worse on most benchmarks and better on REPERE phase2. Keeping it around is reasonable for reproducing older results, not for new work.

## Licence, maintenance and the cost of upgrading

The repository is MIT licensed, and the package metadata classifies it as such. That covers the code. It does not automatically cover the pretrained pipelines hosted on Hugging Face, which sit behind their own user conditions that you accept separately when you request access. The README does not state the licence of the model weights, so anyone shipping a product on top of community-1 should read the model card rather than assume MIT carries over. This is a description of what the repository says, not legal advice.

The last push to the default branch was on 2026-09-18, and the most recent release listed is 4.0.7 on 2026-06-30, following 4.0.6 and 4.0.5 in the same month. That is a recent and active release cadence for the 4.0 line.

Upgrade cost is where the project has historically been expensive. The 4.0 line requires Python 3.10 or later and pins torch>=2.8.0, torchaudio>=2.8.0 and torchcodec>=0.7.0. torchcodec is the notable one: it is the reason ffmpeg became a prerequisite, and it is a newer dependency than most audio stacks carry. Moving from 3.x to 4.x also means moving from the legacy pipeline to community-1, which changes both error rates and speaker labelling behaviour. The benchmark table records community-1 as worse than 3.1 on REPERE phase2, so an upgrade is not uniformly an improvement on every corpus.

## Conclusion

Adopt pyannote.audio if you need speaker diarization inside a Python pipeline and can accept the Hugging Face gating and the ffmpeg system dependency. Avoid it if you need a managed HTTP service with no local GPU, or if you cannot accept the user conditions on the community model: the hosted precision-2 path exists for that case, but it is a paid service with its own API key and its own terms. Before committing, verify three things: that ffmpeg is present on the target machine, that your Hugging Face token has been granted access to pyannote/speaker-diarization-community-1, and which pipeline name your code actually loads, because the string passed to Pipeline.from_pretrained decides whether inference runs on your hardware or on pyannoteAI servers.

## FAQ

### Is pyannote.audio free?

The toolkit itself is MIT licensed and free to install. The community-1 pipeline is free to use but requires accepting user conditions on Hugging Face and supplying an access token, while the precision-2 pipeline is a paid pyannoteAI service that uses a pyannoteAI API key. The README mentions free credits for pyannoteAI.

### How accurate is pyannote.audio?

The README publishes a diarization error rate table across corpora such as AMI, CALLHOME, DIHARD 3 and VoxConverse, with lower values being better. community-1 scores better than the legacy 3.1 pipeline on most listed sets, and precision-2 scores lower still on every set in that table.

### How do I install pyannote.audio?

Install ffmpeg first, since torchcodec needs it for audio decoding, then install the package with uv add pyannote.audio or pip install pyannote.audio. After that you accept the user conditions on the pyannote/speaker-diarization-community-1 model page and create a Hugging Face access token.

### How do I use pyannote.audio to diarize a file?

Load a pipeline with Pipeline.from_pretrained, passing the model identifier and your token, optionally move it to a CUDA device, then call it on an audio path. The returned object exposes speaker_diarization, which you iterate to get turn start, turn end and speaker label.

### What is pyannote.audio?

It is an open-source Python toolkit for speaker diarization built on PyTorch, described in the repository as neural building blocks for speech activity detection, speaker change detection, overlapped speech detection and speaker embedding. It ships pretrained pipelines on Hugging Face and also supports pyannoteAI premium diarization.

### How does pyannote.audio differ from WhisperX?

pyannote.audio produces speaker turns only, with start and end times and speaker labels, and the README does not describe transcription. WhisperX is a transcription pipeline that adds diarization to word-level timestamps, so the two solve different halves of the same problem.

## Sources

- [License: MIT](https://github.com/pyannote/pyannote-audio/blob/main/LICENSE)
- [Project website](https://www.pyannote.ai)
- [pyannote/pyannote-audio on GitHub](https://github.com/pyannote/pyannote-audio)
- [README](https://github.com/pyannote/pyannote-audio/blob/main/README.md)
- [Releases](https://github.com/pyannote/pyannote-audio/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/pyannote-pyannote-audio
