Library / SDK
libAudioFlux/audioFlux avatar
libAudioFlux/audioFlux

audioFlux: a C and Python library for time-frequency analysis and MIR feature extraction

A library for audio and music analysis, feature extraction.

3,371 stars149 forksCMIT

At a glance

What is it?
audioFlux pairs a C core with a Python API and covers transforms, spectral features and MIR tasks in one package. It suits engineers who need many time-frequency representations behind a single interface, and it expects you to know which one you want.
Who is it for?
Adopt audioFlux when you want one Python package that spans BFT, NSGT, CWT, CQT and the feature and MIR layers on top, and when you are comfortable choosing a frequency scale yourself. Do not adopt it as a drop-in replacement for librosa or aubio if your existing pipeline depends on their exact function signatures and defaults; the module split here is different.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly C, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What audioFlux is for, and who ends up using it

audioFlux is a library for audio and music analysis and feature extraction. The README describes it as a deep learning tool library, which is a little misleading: nothing in the described architecture trains a model. What it does is produce the inputs a model would consume. The README says the extracted features "can be provided to deep learning networks for training" for tasks such as classification, separation, music information retrieval and ASR.

The intended user is someone building a feature pipeline rather than a model. If you are preparing mel spectrograms, MFCCs or chroma vectors for a classifier, or you need pitch, onset or harmonic-percussive separation as intermediate steps, this library covers that ground in one dependency. The README frames the design goal as systematic and multi-dimensional feature extraction, so the value proposition is breadth of representations behind a consistent interface rather than a single best-in-class algorithm.

It is less obviously aimed at real-time audio applications. The README's commented-out feature list mentions mobile real-time calculation, but the shipped documentation does not describe a streaming API, and the installation section points at build instructions for iOS and Android rather than at a real-time usage guide. Treat the mobile story as build support, not as a documented streaming contract.

The three-module split: transform, feature, mir

The README lays out three modules, and the split is the clearest thing about the design. The transform module produces time-frequency representations. The feature module derives coefficients from those representations. The mir module handles pitch, onset and harmonic-percussive separation. The README describes the library as based on data stream design, with each algorithm module decoupled in structure, which is what allows features to be combined across dimensions.

Transform is where the breadth sits. Four algorithms, BFT, NSGT, CWT and PWT, support seven frequency scale types: Linear, Linspace, Mel, Bark, Erb, Octave and Log. A second group does not accept a scale parameter and stands alone: CQT, VQT, ST, FST, DWT, WPT and SWT. That distinction matters in practice. If you want a mel-scaled wavelet transform you will not find it here; scale selection is a property of the first group only. Three sharpening routines, reassign for STFT, synsq for CWT data and wsst for CWT, are listed separately.

The feature module has four entries. spectral and xxcc (cepstral coefficients) accept all spectrum types. deconv works on spectra. chroma is the constrained one: the README says it supports only CQT spectrum, plus Linear and Octave spectra based on BFT. If your pipeline is built on mel spectrograms and you need chroma, you have to add a second transform path.

The mir module lists pitch with YIN and STFT among others, onset with spectral flux and novelty, and hpss with median filtering and NMF. The release notes add PitchShift and TimeStretch in v0.1.8 and a TuneTrack algorithm in v0.1.10, described as an instrument tuner for guitar, ukulele, bass, banjo, mandolin and violin.

Installing audioFlux and running a first transform

The README gives two package-manager routes. Python 3.6 or newer is required. The PyPI route is a single command.

bash
pip install audioflux

The Anaconda route uses two channels, tanky25 and conda-forge.

bash
conda install -c tanky25 -c conda-forge audioflux

Building from source is a different proposition. The setup.py in the repository raises a ValueError on Windows during a build, with the message that the platform is not supported and that you should use macOS or pip instead. The compile path shells out to scripts/build_macOS.sh on darwin and scripts/build_linux.sh on linux, producing libaudioflux.dylib or libaudioflux.so. So source builds are a macOS and Linux affair; Windows users go through the wheel.

Runtime dependencies are listed in requirements.txt: numpy, scipy>=1.2.0, soundfile>=0.12.1 and matplotlib. That matplotlib entry is worth noticing. It is declared as a hard requirement, not an extra, so a headless feature-extraction container will still pull in a plotting stack unless you install with --no-deps and manage the list yourself.

The README's quickstart is a set of links rather than inline code: Mel & MFCC, CWT & Synchrosqueezing, CQT & Chroma, Different Wavelet Type, Spectral Features, Pitch Estimate, Onset Detection and Harmonic Percussive Source Separation, each pointing into docs/examples.md. Those example scripts are the place to start, because the README itself does not show a single call. If you want the exact constructor arguments for BFT or the correct scale constant for Erb, the documentation site at audioflux.top is where the README sends you.

Where audioFlux gets in your way

The documentation gap is the first obstacle. The README is an index, not a guide. It names roughly twenty algorithms and then defers every usage question to the docs site or to docs/examples.md. For a library whose main selling point is the number of available transforms, the absence of even one worked example in the README means you cannot evaluate the API shape without leaving the repository.

The chroma constraint is the second. Chroma is the feature most people reach for when doing key or chord work, and the README restricts it to CQT spectrum or to Linear and Octave spectra based on BFT. A mel-based pipeline cannot feed it. You either run a second transform or you compute chroma elsewhere.

The third is the gap between the README and the releases. The new-features section lists v0.1.10 with TuneTrack, but the most recent release in the repository's release list is v0.1.9 from 2024-05-24. The README also describes v0.1.8 features that do correspond to a shipped release. So the tuner may not be in the package you install from PyPI. Check what pip actually resolves before designing around TuneTrack.

Finally, the last push to the repository was on 2026-03-06. That is recent enough that the project is not abandoned, but the release cadence tells its own story: three releases between December 2023 and May 2024, then nothing in the release list for well over a year. Development activity and release activity are not the same thing here, and if you need a pinned, versioned dependency you are pinning to v0.1.9.

audioFlux compared with librosa and aubio

librosa is the default choice for Python audio analysis, and the practical difference is language and packaging. librosa is pure Python on top of numpy and scipy; audioFlux puts its core in C and ships a compiled extension. That is the whole trade: audioFlux gives you FFT acceleration through native code and a single import surface for transforms, features and MIR, while librosa gives you a larger ecosystem, far more tutorials and no build step. If your team cannot maintain a compiled dependency across platforms, that alone decides it.

aubio is the closer comparison on the MIR side. aubio is also a C library with Python bindings, and it targets onset detection, pitch detection and beat tracking. audioFlux overlaps on onset and pitch, and the README's pitch list is long: YIN, CEP, PEF, NCF, HPS, LHS, STFT and FFP across releases. Where they diverge is the transform layer. audioFlux treats the time-frequency representation as the primary object and builds features on top of it, with seven selectable frequency scales and wavelet transforms including CWT, DWT, WPT and SWT. aubio is oriented toward extracting events from a stream. If you are doing offline spectral research, audioFlux's transform catalogue is the reason to pick it; if you are tracking onsets in a live signal, that is aubio's territory and the README does not claim it here.

Licence, maintenance and what an upgrade costs

audioFlux is MIT licensed, per the repository's LICENSE.md and the licence badge in the README. MIT is permissive: it allows use, modification and redistribution with the copyright notice and licence text preserved, and it includes no copyleft obligation on your own code. The repository also carries a Zenodo DOI, which is useful if you need to cite a specific version in academic work. This is a description of the licence text, not legal advice; if you are redistributing a modified binary, read LICENSE.md yourself.

The dependency footprint is the practical cost of adoption. numpy, scipy, soundfile and matplotlib are all pulled in by default, and the scipy floor is 1.2.0. Those are common packages, but scipy and matplotlib together are a meaningful image-size addition for a container that only needs to compute MFCCs.

Upgrade cost is currently low because there is little to upgrade to. The release list ends at v0.1.9 from 2024-05-24, so a pinned dependency is stable by default. The risk runs the other way: if you build against a README feature that has not shipped, you are depending on the master branch, and the source build path for that only works on macOS and Linux. Windows users who need unreleased functionality have no documented route through setup.py.

Editorial conclusion

Adopt audioFlux when you want one Python package that spans BFT, NSGT, CWT, CQT and the feature and MIR layers on top, and when you are comfortable choosing a frequency scale yourself. Do not adopt it as a drop-in replacement for librosa or aubio if your existing pipeline depends on their exact function signatures and defaults; the module split here is different. Before committing, check the docs page for the specific transform you need, confirm that your platform is covered by the install instructions, and verify whether the feature you want already exists in the released package rather than only in the README's new-features list.

Frequently asked questions

What is audioFlux?

audioFlux is a C and Python library for audio and music analysis and feature extraction. It groups its algorithms into transform, feature and mir modules, and the README says the resulting features can be fed to deep learning networks for tasks such as classification, separation, MIR and ASR.

How do I install the audioFlux Python library?

The README gives two routes: pip install audioflux, or conda install -c tanky25 -c conda-forge audioflux. Python 3.6 or newer is required, and the runtime dependencies listed in requirements.txt are numpy, scipy>=1.2.0, soundfile>=0.12.1 and matplotlib.

Can I build audioFlux from source on Windows?

The setup.py in the repository raises a ValueError on Windows during a build, telling you the platform is not supported and to use macOS or pip install audioflux instead. The compile path only handles darwin and linux, producing libaudioflux.dylib or libaudioflux.so.

Which frequency scales does audioFlux support?

The README states that BFT, NSGT, CWT and PWT support Linear, Linspace, Mel, Bark, Erb, Octave and Log scales. CQT, VQT, ST, FST, DWT, WPT and SWT do not take a scale parameter and are used as independent transforms.

Does audioFlux support chroma features from a mel spectrogram?

No. The README says the chroma feature supports only CQT spectrum, plus Linear and Octave spectra based on BFT. A mel-based pipeline has to run a second transform to produce chroma.

Official sources

  1. libAudioFlux/audioFlux on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/libaudioflux-audioflux.svg)](https://hysenlabs.com/projects/libaudioflux-audioflux)