# pydub: high level audio editing in Python, with ffmpeg doing the heavy lifting

> pydub wraps ffmpeg or libav behind an immutable AudioSegment object so slicing, concatenation and format conversion become ordinary Python arithmetic. It is a good fit for scripts that cut and join audio; it is not an analysis or DSP toolkit.

**jiaaro/pydub** — Manipulate audio with a simple and easy high level interface

- Repository: https://github.com/jiaaro/pydub
- Website: http://pydub.com
- Stars: 9,802 · Forks: 1,133
- Language: Python
- License: MIT
- Published: 2026-09-21 · Updated: 2026-09-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/jiaaro-pydub

## What pydub is for, and who ends up using it

The problem pydub addresses is that editing audio from Python normally means either shelling out to ffmpeg with hand-built argument lists or dropping into a sample level library like numpy and scipy. pydub sits in between. The README describes it as letting you "do stuff to audio in a way that isn't stupid", and the API is built around one class, AudioSegment, that behaves like a Python sequence of audio.

The intended audience is developers writing scripts, not audio engineers working in a DAW. Typical work is programmatic: cut a podcast into chapters, trim silence from a batch of recordings, concatenate segments, apply a fade, convert a directory of files from one format to another, and attach metadata tags on export. The setup.py classifiers list the project under Sound/Audio Conversion, Editors and Mixers, which matches that scope.

What it is not is an analysis library. Every operation in the documented API is expressed in milliseconds and decibels. There is no documented interface for reading raw sample arrays, computing spectrograms or detecting pitch. If that is the job, pydub is the wrong layer and you want the samples themselves.

## The AudioSegment model: immutable, chainable, millisecond based

An AudioSegment is created from a file and then never modified. The README states this explicitly and shows it in the examples: calling reverse() on a segment returns a new object and leaves the original alone. That has a practical consequence for memory. A long file that passes through several operations can exist several times over in RAM at once, because each step produces a distinct segment. The README does not document any streaming or chunked read path, so a two hour recording is loaded whole.

Time is measured in milliseconds throughout. The quickstart defines ten_seconds = 10 * 1000 and then slices with song[:ten_seconds] and song[-5000:]. Volume is arithmetic too: adding 6 to a segment boosts it by 6 dB, subtracting 3 reduces it by 3 dB. Concatenation is the plus operator, and multiplication repeats a clip. Because every operation returns an AudioSegment, calls chain, which the README demonstrates with fade_in(2000).fade_out(3000).

Underneath, pydub is a front end. Opening or saving WAV files works in pure Python, but the README is clear that non-WAV formats such as mp3 require ffmpeg or libav. For those, pydub writes a temporary file, invokes the external binary as a subprocess, and reads the result back. The README notes that these temporary files are cleaned up automatically. The debugging section shows the actual call pydub makes, with the ffmpeg arguments and the temporary WAV path in the subprocess.call list, which is the clearest view of the data flow the project offers.

## Installing pydub and running a first edit

Installation from PyPI is a single command. The README gives it as pip install pydub. There is also a development install from git, where you can pin a release tag instead of tracking master.

```bash
pip install pydub
```

The second half of the setup is the external binary, and this is where most first attempts fail. For anything other than WAV you need ffmpeg or libav on the system. The README lists the platform commands. On Linux with aptitude it is apt-get install ffmpeg libavcodec-extra. On macOS with Homebrew it is brew install ffmpeg. On Windows the README points at prebuilt libav binaries, and instructs you to add the libav bin folder to PATH before running pip install pydub.

```bash
apt-get install ffmpeg libavcodec-extra
```

With both pieces in place, a first real edit is short. This opens an mp3, takes the first ten seconds, appends the last five seconds of the same file, and writes the result back out as an mp3 with a 192k bitrate.

```python
from pydub import AudioSegment

song = AudioSegment.from_mp3("never_gonna_give_you_up.mp3")
first_10_seconds = song[:10000]
last_5_seconds = song[-5000:]
without_the_middle = first_10_seconds + last_5_seconds
without_the_middle.export("mashup.mp3", format="mp3", bitrate="192k")
```

According to the README, without_the_middle.duration_seconds should equal 15.0, which is a cheap way to confirm the slicing behaved as expected before you export. If the export step raises an error about ffmpeg or avconv, the binary is missing or not on PATH, and the debugging logger below is the fastest way to see what pydub actually tried to run.

## Debugging conversion failures with the pydub.converter logger

The README states that most issues people run into are related to converting between formats with ffmpeg or libav, and it provides a logger for exactly that. Enabling DEBUG on the pydub.converter logger prints the subprocess call pydub constructs, including the temporary input path, the -vn flag and the output format.

```python
import logging

l = logging.getLogger("pydub.converter")
l.setLevel(logging.DEBUG)
l.addHandler(logging.StreamHandler())

AudioSegment.from_file("./test/data/test1.mp3")
```

What you should see is a line beginning with subprocess.call and the full ffmpeg argument list. That output is what you take to the ffmpeg documentation when a codec is missing, because the error pydub raises on its own is usually less specific than the binary's own stderr. This is the single most useful diagnostic the project documents, and it is worth turning on before filing an issue rather than after.

## Export parameters are passed through without validation

The export method accepts a bitrate argument in any syntax ffmpeg supports, such as bitrate="192k". Beyond that, arbitrary ffmpeg switches go into a parameters list, with the switch first and its argument second. The README shows parameters=["-q:a", "0"] for lame V0 quality and parameters=["-ac", "2", "-vol", "150"] to mix down to two channels and set hard output volume.

The README is direct about the trade-off: no validation takes place on these parameters, and you may be limited by what your particular build of ffmpeg or avlib supports. That means a typo in a switch, or a codec your distro's ffmpeg was compiled without, surfaces as a subprocess failure rather than a Python error. It also means the same pydub code can behave differently across machines with different ffmpeg builds. If your pipeline depends on a specific encoder, you are depending on the system binary, not on pydub, and that should be pinned and tested as part of deployment.

The Ogg case is a related design decision the README calls out. The Ogg specification does not fix a codec, so pydub defaults to vorbis when you export to ogg without specifying one. That is a convenience, but it is a choice the format itself leaves open, so an ogg file produced by pydub is not guaranteed to be what a different tool would have produced.

## Where pydub stops: analysis, streaming and heavy DSP

The clearest limitation is scope. pydub's documented vocabulary is slicing, concatenation, gain, fades, crossfades, repetition, playback and format conversion. There is no documented FFT, no spectral analysis, no pitch or tempo detection, and no frame level access to the underlying samples. Questions about audio analysis in Python are a different problem, and pydub does not claim to answer them.

A second limitation is the external dependency. Pure Python covers WAV only. Every mp3, ogg, flv, mp4, wma or aac operation goes through a subprocess call to ffmpeg or libav. That adds process startup cost per operation, requires the binary to exist in every environment including containers and CI runners, and makes the library's behaviour partly a function of which ffmpeg build is installed. For a serverless function with a small deployment package, that is a real constraint.

A third is the immutability model combined with whole-file loading. It is elegant for short clips and awkward for very large files, because intermediate segments are not freed until you drop the references. The README does not describe a streaming API, so there is nothing to reach for when the file does not fit comfortably in memory.

Finally, playback needs a backend. The README lists simpleaudio, pyaudio, ffplay and avplay, and says simpleaudio is strongly recommended even if you are installing ffmpeg or libav. So play(sound) is not a guarantee; it is a guarantee conditional on one of four libraries being present.

## pydub against raw ffmpeg and against sample level libraries

The obvious alternative is calling ffmpeg directly, through subprocess or a wrapper. The difference in approach is that ffmpeg's model is a filter graph over files, and pydub's model is an in-memory object you manipulate in Python. For a one-off conversion, a direct ffmpeg command is shorter and has no Python dependency at all. For a script that decides at runtime which parts of which files to keep, based on data from somewhere else, building the equivalent ffmpeg filter graph by hand is considerably more work than slicing an AudioSegment and adding the pieces together.

On the other side, libraries that expose raw sample arrays occupy a different layer entirely. They give you the frames, so you can compute anything, at the cost of writing the audio handling yourself. pydub deliberately does not go there. Choosing between them is really choosing whether your problem is editing or measuring. pydub is an editing tool that happens to be scriptable; it is not a measurement tool that happens to be able to write files.

There is a middle path worth noting: pydub and a sample level library are not mutually exclusive, since the export step hands you a file you can load with whatever reads samples. The cost is an extra encode and decode round trip, which for lossy formats means generation loss.

## Maintenance, versioning and the MIT licence

The repository is not archived, and the last push was on 2026-03-19. The most recent tagged release in the list is v0.25.1 from 2021-03-10, with v0.25.0 five days earlier and v0.24.1 before that in 2020. So the codebase sees commits, but tagged releases are years apart, and the version pinned in setup.py is 0.25.1. If you install from PyPI you get that release; if you install from git master you get whatever has landed since, which the README acknowledges by offering both installation routes.

For upgrade planning, the practical implication is that the dependency you should watch is not pydub but ffmpeg. A pydub upgrade is a small surface change; an ffmpeg upgrade can change encoder defaults, deprecate switches and alter what your parameters list does. The README's warning that no validation takes place on passed-through parameters is the reason.

The licence is MIT, stated in setup.py and present as a LICENSE file at the repository root. MIT is permissive and imposes no copyleft obligation on your own code. That covers pydub itself. It does not cover ffmpeg or libav, which are separate programs with their own licensing, and it does not cover the codecs those binaries may be built against. If you are shipping a product that bundles an ffmpeg build, the terms you need to check are that build's, not pydub's. This is a description of what the files say, not legal advice.

## Conclusion

Use pydub when the job is cutting, joining, fading or converting audio inside a Python script and ffmpeg or libav is already available or acceptable as a system dependency. Do not pick it for spectrum analysis, pitch detection or anything that needs sample level DSP, because the API works in milliseconds and decibels, not in frames and FFTs. Before committing, verify that ffmpeg or libav is on the PATH of every machine that will run the code, including CI containers, and check the export parameters against the ffmpeg build actually installed, since pydub passes them through without validation.

## FAQ

### How do I install pydub?

Install the Python package with pip install pydub, or install from git with pip install git+https://github.com/jiaaro/pydub.git@master if you want the development version. You also need ffmpeg or libav for any format other than WAV. The README gives platform commands such as apt-get install ffmpeg libavcodec-extra on Linux and brew install ffmpeg on macOS.

### What is pydub used for?

pydub manipulates audio through a high level interface built around the AudioSegment class. The README shows slicing by milliseconds, concatenating segments, applying gain in decibels, crossfading, repeating clips, fading in and out, playing audio, and exporting to any format ffmpeg supports with optional metadata tags and bitrate.

### How can I play an MP3 file using Python with pydub?

Load the file with AudioSegment.from_mp3 and pass it to pydub.playback.play. Playback requires one of simpleaudio, pyaudio, ffplay or avplay to be installed, and the README strongly recommends simpleaudio. Note that opening an mp3 at all requires ffmpeg or libav, since pure Python only handles WAV.

### How do I install pydub on Windows?

The README's Windows instructions are to download and extract libav from the Windows binaries it links, add the libav bin folder to the PATH environment variable, and then run pip install pydub. The order matters, since the binary needs to be reachable before pydub tries to convert non-WAV formats.

### How do I install pydub in Python using pip?

The README gives the install command as pip install pydub. A development install from git is also documented, as pip install git+https://github.com/jiaaro/pydub.git@master, where the tag can be replaced with a release version. Neither route installs ffmpeg or libav, which you need separately for non-WAV formats.

### What is pydub in Python?

It is a Python library for manipulating audio through a high level interface, with AudioSegment as the central object and operations expressed in milliseconds and decibels. It is published on PyPI, licensed MIT, and documented through the README and a separate API.markdown file in the repository.

## Sources

- [jiaaro/pydub on GitHub](https://github.com/jiaaro/pydub)
- [License: MIT](https://github.com/jiaaro/pydub/blob/master/LICENSE)
- [Project website](http://pydub.com)
- [README](https://github.com/jiaaro/pydub/blob/master/README.md)
- [Releases](https://github.com/jiaaro/pydub/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/jiaaro-pydub
