Library / SDK
pytorch/audio avatar
pytorch/audio

torchaudio: the audio I/O and transform layer for PyTorch

Data manipulation and transformation for audio signal processing, powered by PyTorch

2,950 stars800 forksPythonBSD-2-Clause

At a glance

What is it?
torchaudio is PyTorch's audio library, now in a maintenance phase with a narrower scope: loading audio, decoding it into tensors, and applying transforms that stay inside the autograd graph. This article covers what survived the 2.9 removals, how to install it, and where it stops being the right tool.
Who is it for?
Adopt torchaudio if your pipeline already produces or consumes torch tensors and you need decoding, resampling, spectrograms or Kaldi-compatible features inside that graph. Do not adopt it as a general signal-processing toolkit: the project states it is primarily a machine learning library, and the 2.9 transition removed user-facing features rather than adding them.
Can I use it commercially?
Yes. BSD-2-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What torchaudio is for, and what it deliberately is not

The README is direct about the boundary: torchaudio exists to apply PyTorch to the audio domain, and it is "primarily a machine learning library and not a general signal processing library." That sentence decides most adoption questions. If you want to filter a recording, measure loudness, or cut a file, torchaudio is the wrong layer. If you want a waveform as a tensor that carries gradients, feeds a DataLoader, and lives on the same device as your model, it is the intended layer.

The audience follows from that. You are training or fine-tuning a speech, audio-tagging, source-separation or self-supervised model and you need the boring parts handled: reading files, resampling, turning waveforms into spectrograms or MFCCs, and preparing common research datasets. The repository's examples directory reflects this, with subdirectories for asr, avsr, hubert, self_supervised_learning, source_separation, and pipeline examples for Tacotron 2, wav2letter and WaveRNN. Those are training pipelines, not utilities.

The project also ships compliance interfaces so that code can align with other libraries, specifically Kaldi spectrogram, fbank and mfcc computation. That matters when you are porting a recipe and need feature values to match a reference implementation rather than merely be reasonable.

The maintenance phase and the 2.9 removals you inherit

A note near the top of the README states that TorchAudio has transitioned into a maintenance phase, and that the migration removed some user-facing features which were deprecated from 2.8 and removed in 2.9. The stated goals were reducing redundancies with the rest of the PyTorch ecosystem, easing maintenance, and scoping the library more tightly to processing audio data for ML.

This is the single most important fact for anyone evaluating the project, and it cuts both ways. On one side, a smaller surface is easier to depend on: fewer APIs means fewer things that can change under you. On the other, code written against older tutorials may simply not run. The README points to a community message for details rather than enumerating the removals, so the practical step is to check the API reference for whatever you call before you build on it. The library is not archived and the last push was on 2026-09-10, so work continues, but the direction is subtraction rather than expansion.

Read the release cadence with the same lens. v2.9.1 arrived on 2025-11-12, v2.10.0 on 2026-01-21, and v2.11.0 on 2026-03-23. Releases keep coming; the scope is what changed.

How versioning works now: torchaudio does not pin torch

The README makes a claim that is unusual in the Python audio ecosystem and worth understanding precisely. TorchAudio 2.11 works with torch 2.11 and, in the project's words, with every future torch release such as 2.12 and 2.13. Installing TorchAudio does not pin torch to a specific version, and there is nothing to upgrade in TorchAudio when you upgrade torch.

That removes the dependency-resolution dance that usually accompanies a torch upgrade: you do not wait for a matching audio release before moving the framework forward. The constraint runs the other way instead. TorchAudio 2.11 requires torch 2.11 or newer, so the floor moves with the audio release while the ceiling does not exist. If you are on an older torch, this is not a library you can adopt incrementally; you upgrade torch first.

The README also notes that the compatibility matrix for older releases lives on the installation page, which is where you should look if you need to stay pinned to an older pairing rather than track the newest.

Installing torchaudio and running a first transform

The README gives a single installation command. The runtime dependency list in requirements.txt names torch as the minimum and torchcodec as optional, so a plain install pulls the framework and leaves the optional decoder out.

bash
pip install torchaudio

After that, the fastest way to confirm the install and see the core data flow is to run a documented transform. The transforms listed in the README include Spectrogram, AmplitudeToDB, MelScale, MelSpectrogram, MFCC, MuLawEncoding, MuLawDecoding and Resample. The README's API reference is where each of these is documented.

python
import torchaudio
import torchaudio.transforms as T

The import above is the entry point the README's transform list describes; from there you construct the transform you need and apply it to a waveform tensor. Because every step is a torch operation, the result stays on the same device as the input and remains part of the autograd graph, which is the property that makes this library different from calling out to a separate DSP package.

For real files rather than synthetic tensors, the repository's datasets module provides loaders for common audio datasets, and the README's disclaimer is explicit that the library downloads and prepares public datasets but does not host or distribute them, and that determining whether you have permission to use a given dataset is your responsibility.

Where torchaudio stops being the right tool

The maintenance phase is a real limitation, not a footnote. If your roadmap assumes the audio library will absorb new model families, new codecs or new dataset formats at the pace of the wider ecosystem, the README's own framing tells you otherwise: the goal was to reduce redundancies with the rest of PyTorch, and features were removed to get there. Plan for a stable, narrow dependency rather than a growing one.

The second limitation is scope. Because the project positions itself as a machine learning library, the general signal-processing work you might expect from an audio package is out of frame. Reaching for torchaudio to trim a file, normalize loudness for delivery, or inspect a codec is a mismatch, and the transform list in the README is a fair map of what is actually on offer.

The third is licensing, which is easy to miss because the library itself is permissive. The README warns that pre-trained models provided in the library may carry their own licenses derived from their training data, and gives SquimSubjective as an example, released under CC-BY-NC 4.0. A BSD-2-Clause library can therefore hand you a non-commercial model. If your product ships audio models, that distinction decides whether a given model is usable, and the README says other pre-trained models with different licenses are noted in the documentation.

How it compares to a general DSP stack

The honest alternative for the non-ML half of audio work is a general-purpose signal-processing library, and the difference is architectural rather than a matter of features. A DSP library treats audio as arrays and returns arrays; its functions are designed to be called once, outside any training loop, and its value is breadth of analysis and filtering operations.

TorchAudio inverts that. Its computations run through PyTorch operations, which the README presents as the reason the library feels like a natural extension of the framework. The consequence is that a transform is a module you can place in a model, batch across a DataLoader, move to a GPU and differentiate through. A DSP library gives you none of that without you writing the glue.

The trade-off is real in both directions. Choose torchaudio and you accept a narrower operation set and a project that has explicitly stopped expanding. Choose the DSP route and you keep the breadth, but you own the conversion into tensors, the device placement, and the batching. For a training pipeline the second option is usually more work than it looks; for a one-off analysis script the first option is the wrong dependency to install.

Upgrade cost, licence, and what to check before you depend on it

Upgrade cost is unusually low on the framework axis and potentially high on the API axis. Because TorchAudio does not pin torch and needs no upgrade when torch moves, the mechanical part of tracking new releases is close to free. The API part is where the 2.8 deprecations and 2.9 removals landed, so the risk is concentrated in code that calls features the project decided to drop. The mitigation is mundane: check each torchaudio symbol you use against the current API reference before you upgrade, rather than after.

On licensing, the library is BSD-2-Clause, which is permissive, but two separate notices sit alongside it. The dataset disclaimer states that the library does not host or distribute the datasets it prepares, does not vouch for their quality or fairness, and does not claim you have licence to use them. The pre-trained model notice says the same in substance for models. Neither is legal advice, and both put the determination on you.

If you maintain an older deployment, the README points to the installation page for the compatibility matrix of older releases. That page, not the README, is where a pinned pairing should be confirmed.

Editorial conclusion

Adopt torchaudio if your pipeline already produces or consumes torch tensors and you need decoding, resampling, spectrograms or Kaldi-compatible features inside that graph. Do not adopt it as a general signal-processing toolkit: the project states it is primarily a machine learning library, and the 2.9 transition removed user-facing features rather than adding them. Before committing, check the installation page for the compatibility matrix of older releases, confirm that whatever API you depend on still exists in 2.11, and read the licence notes on any pre-trained model you plan to ship, since SquimSubjective is CC-BY-NC 4.0 while the library itself is BSD-2-Clause.

Frequently asked questions

What is torchaudio used for?

It applies PyTorch to audio, providing dataset loaders, audio and speech processing functions such as forced_align, and common transforms including Spectrogram, MelScale, MelSpectrogram, MFCC and Resample. The README describes it as primarily a machine learning library rather than a general signal processing library.

How do I install torchaudio?

The README gives a single command, pip install torchaudio. The full installation and build instructions, along with the compatibility matrix for older releases, are on the project's installation page.

Does torchaudio pin a specific version of torch?

No. The README states that TorchAudio 2.11 requires torch 2.11 or newer, including future releases, and that installing TorchAudio does not pin torch to a specific version. There is nothing to upgrade in TorchAudio when you upgrade torch.

Is torchaudio still being developed?

The README states that TorchAudio has transitioned into a maintenance phase, with features deprecated from 2.8 and removed in 2.9. The repository is not archived and the last push was on 2026-09-10, but the stated goal was to reduce redundancies with the rest of PyTorch and scope the library more tightly.

Official sources

  1. License: BSD-2-Clause
  2. Project website
  3. pytorch/audio on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/pytorch-audio.svg)](https://hysenlabs.com/projects/pytorch-audio)