SpeechBrain: A PyTorch Toolkit for Speech Recognition, Speaker ID, and More
A PyTorch-based Speech Toolkit
At a glance
- What is it?
- SpeechBrain is an open-source PyTorch toolkit for building, training, and deploying speech and audio processing systems. It provides over 200 training recipes across more than 40 datasets and over 100 pretrained models, covering tasks from ASR to speaker diarization and emotion recognition.
- Who is it for?
- SpeechBrain is well suited for researchers who need a unified framework to experiment across speech tasks without switching codebases, and for engineers who need pretrained models for ASR, speaker verification, or emotion recognition with a three-line inference API. The last push to the repository was on 2026-08-27, with v1.1.1 released on the same date.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 33 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What SpeechBrain Covers and Who It Is For
SpeechBrain is a PyTorch speech toolkit described in its README as targeting conversational AI development, specifically the technology behind speech assistants, chatbots, and language models. The project covers a broad range of tasks: speech recognition (ASR), speaker recognition and verification, speech enhancement, speech separation, language modeling, dialogue, and speaker diarization. Support for EEG data processing has also been added, referenced as the MOABB benchmarks in the README.
The toolkit is used for two main purposes. Researchers use it to run training experiments on established datasets and compare new models against baselines, because SpeechBrain provides over 200 competitive training recipes across more than 40 datasets. Engineers use the pretrained models, hosted on HuggingFace, for inference in production or prototype pipelines. The README mentions use by academic institutions including Mila, Concordia University, and Avignon University for student training.
Installing SpeechBrain and Running ASR Inference
Install SpeechBrain via PyPI:
pip install speechbrainFor development or training experiments, install from the repository:
git clone https://github.com/speechbrain/speechbrain.git
cd speechbrain
pip install -r requirements.txt
pip install --editable .The `--editable` flag means changes to the speechbrain package are reflected immediately without reinstalling. Verify the installation:
pytest tests
pytest --doctest-modules speechbrainRunning inference on a pretrained ASR model requires three lines of Python:
from speechbrain.inference import EncoderDecoderASR
asr_model = EncoderDecoderASR.from_hparams(source="speechbrain/asr-conformer-transformerlm-librispeech", savedir="pretrained_models/asr-transformer-transformerlm-librispeech")
asr_model.transcribe_file("speechbrain/asr-conformer-transformerlm-librispeech/example.wav")The model downloads automatically from HuggingFace on first run and caches to the savedir. The same from_hparams pattern works for any of the 100+ models on the SpeechBrain HuggingFace organisation.
Training Recipes: The YAML and Python Architecture
SpeechBrain organises training experiments as pairs of a Python script and a YAML hyperparameter file. The README explains: hyperparameters are in the YAML file, and the training process is orchestrated through a Python script. Running any recipe follows the same pattern:
python train.py hparams/train.yamlResults are saved to the output_folder specified in the YAML file. Changing training settings means editing the YAML file rather than the Python code, which separates configuration from implementation.
SpeechBrain supports fine-tuning pretrained models from outside the project. The README lists Whisper, Wav2Vec2, WavLM, HuBERT, GPT2, and Llama2 as models from HuggingFace that can be plugged in and fine-tuned within the SpeechBrain recipe system. Training logs and checkpoints are hosted on Dropbox for replicability, so published results can be reproduced from the stored state.
The recipes/ directory in the repository is organised by dataset and task. Each dataset directory contains subdirectories for each supported task, and each task subdirectory contains the train.py and hparams/ files needed to run that experiment.
Supported Tasks and What the Pretrained Models Cover
SpeechBrain covers over 20 speech and text processing tasks across more than 40 datasets. The README mentions these categories: ASR (automatic speech recognition), speaker recognition, speaker verification, speaker diarization, speech enhancement, speech separation, language modeling, spoken language understanding, dialogue, and EEG-based brain-computer interface support via MOABB benchmarks.
For inference, the HuggingFace organisation at huggingface.co/speechbrain hosts over 100 pretrained models. Each model comes with a user-friendly Python interface. The EncoderDecoderASR class shown in the installation section is one example; similar classes exist for speaker verification, speech enhancement, and other tasks. The from_hparams method handles downloading, caching, and loading.
The RELATED SEARCHES data shows real user interest in specific models: ECAPA-TDNN for speaker recognition, SepFormer for speech separation, VAD (voice activity detection), emotion recognition, and speaker diarization. All of these are tasks supported by SpeechBrain and covered by pretrained models on HuggingFace.
Limitations and Dependency Weight
SpeechBrain's requirements.txt lists a substantial dependency set: numpy, scipy, pandas, torch, torchaudio, transformers, huggingface_hub, joblib, sentencepiece, soundfile, tqdm, hyperpyyaml, and pre-commit for development. Installing SpeechBrain with all dependencies pulls in a significant amount of code. Teams integrating it into a production container should plan for image size.
SpeechBrain is a research toolkit first. The code structure is consistent across tasks, but the recipes are primarily maintained for their research results rather than production deployment ergonomics. Engineers who need production-grade ASR with latency guarantees, streaming input, or specific hardware acceleration may find the toolkit's abstraction level is not tuned for those requirements.
The project is built on the develop branch rather than main. The pyproject.toml shows the Python requirement is 3.8.1 or newer, and the PyPI package name is speechbrain.
SpeechBrain vs Whisper for Transcription Tasks
Whisper is the most widely known open-source ASR model. SpeechBrain supports fine-tuning Whisper within its training recipe system, meaning the two are not strictly competing tools. SpeechBrain is a toolkit for building and training speech systems; Whisper is a pretrained ASR model.
For off-the-shelf transcription, Whisper (via the openai-whisper library) is simpler to use. It requires no training, no recipe configuration, and transcribes directly from audio. For tasks where Whisper does not work out of the box, including speaker diarization, speaker verification, speech enhancement, and custom domain ASR fine-tuning, SpeechBrain provides infrastructure that Whisper alone does not.
The last push to the repository was on 2026-08-27, coinciding with the v1.1.1 release. v1.1.0 shipped on 2026-03-30. The project is licensed under Apache-2.0. Citation information is documented in the CITATION.cff file in the repository root.
Editorial conclusion
SpeechBrain is well suited for researchers who need a unified framework to experiment across speech tasks without switching codebases, and for engineers who need pretrained models for ASR, speaker verification, or emotion recognition with a three-line inference API. The last push to the repository was on 2026-08-27, with v1.1.1 released on the same date. Teams doing production deployment should evaluate the full dependency list, which includes scipy, sentencepiece, soundfile, and torchaudio in addition to PyTorch itself. The Apache-2.0 license allows commercial use without restriction.
Frequently asked questions
How do I install SpeechBrain?
Run `pip install speechbrain` to install from PyPI. For training experiments or development, clone the repository and run `pip install -r requirements.txt` followed by `pip install --editable .` from the project root.
What is SpeechBrain?
SpeechBrain is an open-source PyTorch toolkit for speech and audio processing. It provides training recipes for tasks including ASR, speaker recognition, speech separation, and emotion recognition, plus over 100 pretrained models on HuggingFace for direct inference.
Is SpeechBrain free to use?
Yes. SpeechBrain is released under the Apache-2.0 license, which allows use, modification, and distribution for both open-source and commercial purposes without restriction.
How does SpeechBrain compare to Whisper?
Whisper is a pretrained ASR model optimised for off-the-shelf transcription. SpeechBrain is a training and inference framework that covers ASR, speaker recognition, diarization, speech enhancement, and more. SpeechBrain can fine-tune Whisper within its recipe system, so the two can be used together.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/speechbrain-speechbrain)