# Essentia: a C++ audio analysis and MIR library with Python bindings

> Essentia is an AGPLv3 C++ library for audio analysis and music information retrieval, wrapped in Python and shipped with command-line extractors. It fits research pipelines and batch descriptor computation; it is the wrong tool if you cannot accept the licence or need a stable tagged release.

**MTG/essentia** — C++ library for audio and music analysis, description and synthesis, including Python bindings

- Repository: https://github.com/MTG/essentia
- Website: http://essentia.upf.edu
- Stars: 3,754 · Forks: 634
- Language: C++
- License: AGPL-3.0
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/mtg-essentia

## What Essentia solves, and who writes it into a pipeline

Essentia is a C++ library for audio analysis and audio-based music information retrieval, released under the Affero GPLv3. The README describes an extensive collection of reusable algorithms covering audio input and output, standard digital signal processing blocks, statistical characterisation of data, and a large set of spectral, temporal, tonal and high-level music descriptors. That list is the product. If your job is to turn a folder of audio files into numbers that describe them, the library already contains most of the blocks you would otherwise write yourself.

The intended users are visible in how the project is packaged. It ships a Vamp plugin for Sonic Visualiser, so musicologists and researchers who want to see descriptors plotted against a waveform are a first-class audience. It ships Python bindings and predefined executable extractors, which the README says facilitate fast prototyping and allow research experiments to be set up very rapidly. It also ships Docker images and prebuilt static binaries for command-line extractors, which points at batch jobs and larger deployments. The README frames the range as covering both research experiments and large-scale industrial applications.

The design emphasis stated in the README is on robustness of the provided music descriptors and on computational cost of the algorithms. That is a claim about the descriptor implementations rather than about the framework around them, and it is the reason a team might pick this over assembling DSP code from smaller pieces.

## Algorithms, extractors and bindings: how the pieces fit

The architecture is a library of algorithms plus a set of wrappers around it. At the bottom is libessentia, the C++ core. The Python module is a binding over that core, and the repository layout shows the packaging machinery: pyproject.toml declares the build system, setup.py drives the extension build, and a waf script sits at the top level alongside wscript. The setup.py file shows the extension build shelling out to waf and to a third-party build script under packaging/, which tells you the Python wheel is not a thin ctypes shim but a compiled artefact.

Above the core sit executable extractors. The README points to doc/sphinxdoc/extractors_out_of_box.rst for command-line tools that compute common music descriptors, and says you can download prebuilt static binaries for a number of them instead of installing the complete library. That is a meaningful split: a researcher who only needs standard descriptors can run a binary, while someone building a custom pipeline works against the library or the Python bindings.

A separate distribution, essentia-tensorflow, is published on PyPI next to the plain essentia package. The README lists both install commands without explaining the difference in detail, so the practical reading is that one variant links TensorFlow for models that need it and the other does not. The pyproject.toml dependencies for the base package are numpy, pyyaml and six, which is a small runtime footprint for a library of this scope.

For visualisation, the Vamp plugin plugs the same algorithms into Sonic Visualiser. That gives a third entry point into the same descriptor code, useful when you want to look at a signal before writing a batch job over thousands of files.

## Installing Essentia and computing a first descriptor set

The README gives the shortest path for Python on Linux x86_64 and i686: install the wheel from PyPI. This pulls the compiled extension, so no C++ toolchain is needed.

```bash
pip install essentia
```

If your work needs the TensorFlow-linked variant, the README lists a second package name rather than a flag.

```bash
pip install essentia-tensorflow
```

For the full library, the README points at the online installation instructions and at doc/sphinxdoc/installing.rst in the repository, and states that the library is cross-platform, currently supporting Linux, macOS, Windows, iOS and Android. It also says to install from master for the latest updates, which is worth reading twice: the newest tagged release listed in the repository is v2.1_beta5 from 2019-09-05, so the master branch is where current work lives.

There is also a Docker route. The README links Docker images at hub.docker.com/r/mtgupf/essentia, which avoids building the C++ core and the third-party dependencies yourself. The setup.py file shows why that matters: unless the environment variables ESSENTIA_WHEEL_SKIP_3RDPARTY or ESSENTIA_WHEEL_ONLY_PYTHON are set to 1, the build runs packaging/build_3rdparty_static_debian.sh before configuring with waf.

Once installed, the README points Python users at the online tutorial and at src/examples/python/essentia_python_tutorial.ipynb inside the repository. That notebook is the place to start for the actual call sequence, because the README itself does not print a Python snippet. For descriptor computation without writing code, the extractors documented in doc/sphinxdoc/extractors_out_of_box.rst are the entry point.

## Version drift is the real cost of building on Essentia

The README carries an explicit warning: if you use the example extractors in src/examples, or your own code employing Essentia algorithms to compute descriptors, you should be aware of possible incompatibilities when using different versions of Essentia. That is not a hedge, it is an operational constraint. Descriptor values and algorithm interfaces can differ across versions, which means a result you published or a model you trained against one build may not reproduce against another.

The release history makes this sharper. The three most recent releases in the repository are v2.1_beta5 from 2019-09-05, v2.1_beta4 from 2018-05-23 and v2.1_beta3 from 2016-09-29. All are betas, and the newest is years behind the last push to master on 2026-09-21. The README's advice to install from master is consistent with that gap, but it also means a production system is tracking a branch rather than a version. Pin the commit you build, and treat a rebuild as a change that needs revalidation of your descriptor outputs.

The second limitation is scope. Essentia is a descriptor and DSP library, not a training framework or a model zoo. If your problem is source separation, transcription or embedding-based similarity, the algorithms here are building blocks and feature extractors rather than an end-to-end solution. Teams expecting a pretrained pipeline will spend their time on the model side regardless.

The third is the licence, covered below, and it is the one that most often disqualifies the library for a commercial product rather than delaying it.

## Essentia versus librosa, and what each one assumes

The natural comparison for a Python user is librosa, and the difference is in what each library is built around. Essentia is a compiled C++ core with Python bindings, and its centre of gravity is a fixed catalogue of music descriptors plus command-line extractors and a Vamp plugin for Sonic Visualiser. librosa is a Python library built on NumPy and SciPy, and its centre of gravity is spectral and time-frequency representations that you compose into your own analysis.

The practical consequences follow from that split. With Essentia you get descriptor definitions that the project maintains and that are intended to be consistent across users, which is what makes results comparable between research groups. With librosa you get primitives and the freedom to assemble them, plus a pure-Python stack that is easier to read and modify. If you want a specific published descriptor computed the way the MIR community computes it, Essentia's extractors are the shorter route. If you want to prototype a new representation and iterate on the maths, the Python-native stack is less friction.

There is also a deployment difference. Essentia can be embedded in a C++ application, run from prebuilt static binaries, or pulled as a Docker image, and it supports Linux, macOS, Windows, iOS and Android according to the README. A NumPy-based library is tied to a Python runtime. For an offline batch job over a large archive, the compiled path is the one that avoids shipping an interpreter.

## Licence and maintenance: what to check before you ship

Essentia is released under the Affero GPLv3, and the repository also contains a file named Essentia Licensing.txt at the top level, which suggests separate terms exist for some uses. The pyproject.toml declares the package licence as AGPL-3.0-only. The Affero variant is the part that matters for product decisions: it extends the copyleft obligation to users who interact with the software over a network. If you run Essentia behind a web service, the licence question is not the same as embedding a permissively licensed DSP library. Read Essentia Licensing.txt and take your own legal advice; nothing here is legal advice.

On maintenance, the repository is not archived and the last push was on 2026-09-21, so the codebase is being touched. The release tags tell a different story: the newest is v2.1_beta5 from 2019-09-05, and the README explicitly directs users to master for the latest updates. That combination means upgrade cost is real but manageable if you pin commits and re-run your descriptor validation whenever you move the pin. It is a poor fit for a team that needs a vendor-style support contract or a semantic versioning guarantee.

The project asks for contributions through pull requests against the GitHub repository, following a contribution policy, and requires that submitted code complies with the Developer's Certificate of Origin. There is also an issue thread referenced in the README for suggesting improvements, including proposals for new algorithms.

## Conclusion

Adopt Essentia when you need a wide set of spectral, temporal and tonal descriptors inside a C++ or Python pipeline and the AGPLv3 terms work for your distribution model. Do not adopt it if you need a stable numbered release with a documented upgrade path, since the newest tagged release is v2.1_beta5 from 2019-09-05 and the README tells you to install from master. Before committing, verify the licence obligations for your product, confirm that the algorithm you need exists in the current master tree, and pin the commit you build against.

## FAQ

### What is Essentia?

Essentia is an open-source C++ library for audio analysis and audio-based music information retrieval, released under the Affero GPLv3. It includes DSP blocks, statistical characterisation and a large set of spectral, temporal, tonal and high-level music descriptors, with Python bindings and command-line extractors on top.

### How do I install Essentia for Python?

On Linux x86_64 and i686 the README gives pip install essentia, or pip install essentia-tensorflow for the TensorFlow-linked variant. Other platforms are covered by the installation instructions linked from the README and by the Docker images at hub.docker.com/r/mtgupf/essentia.

### How does Essentia compare with librosa?

Essentia is a compiled C++ core with Python bindings and a maintained catalogue of music descriptors plus command-line extractors and a Vamp plugin. librosa is a Python library built on NumPy and SciPy, oriented toward composing your own spectral analysis rather than using predefined descriptors.

### Which Essentia version should I use?

The README says to install from master for the latest updates. The newest tagged release in the repository is v2.1_beta5 from 2019-09-05, and the README warns that using different versions of Essentia can produce incompatibilities in descriptor computation.

### What licence does Essentia use?

The project is released under the Affero GPLv3, and pyproject.toml declares the package licence as AGPL-3.0-only. The repository also contains a top-level file named Essentia Licensing.txt, which is worth reading if you plan to use the library in a product.

### Is there an Essentia alternative for computing music descriptors?

librosa is the closest alternative for Python users, but it takes a different approach: it provides spectral and time-frequency primitives rather than a fixed catalogue of maintained descriptors with extractor binaries and a Vamp plugin.

## Sources

- [License: AGPL-3.0](https://github.com/MTG/essentia/blob/master/LICENSE)
- [MTG/essentia on GitHub](https://github.com/MTG/essentia)
- [Project website](http://essentia.upf.edu)
- [README](https://github.com/MTG/essentia/blob/master/README.md)
- [Releases](https://github.com/MTG/essentia/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/mtg-essentia
