# MTG-Jamendo Dataset: 55,000 Tagged Music Tracks for Auto-Tagging Research

> MTG-Jamendo is an open dataset of over 55,000 full-audio music tracks from Jamendo, tagged across 195 genre, instrument, and mood/theme categories. It ships with pre-built data splits, Python download scripts, and precomputed mel-spectrograms, making it a ready-made benchmark for music auto-tagging and classification research.

**MTG/mtg-jamendo-dataset** — Metadata, scripts and baselines for the MTG-Jamendo dataset

- Repository: https://github.com/MTG/mtg-jamendo-dataset
- Stars: 408 · Forks: 54
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/mtg-mtg-jamendo-dataset

## What MTG-Jamendo Is and Who It Serves

MTG-Jamendo Dataset was created by the Music Technology Group at Universitat Pompeu Fabra as an open resource for music auto-tagging research. It contains over 55,000 full-audio tracks, all distributed under Creative Commons licenses as they appear on Jamendo. Tags come from content uploaders rather than editorial annotation, which reflects how music is labeled in a real consumer catalog rather than under controlled lab conditions.

The dataset covers three broad tag categories: 87 genre tags, 40 instrument tags, and 56 mood/theme tags in the main splits. A separate top-50 subset restricts the tag vocabulary to the 50 most frequent tags, which is more practical for researchers who want a well-populated label space without handling rare classes. Subsets are also available for genre, instrument, and mood/theme independently, each with matching audio downloads.

The primary audience is music information retrieval researchers who need a large, multi-label auto-tagging benchmark with actual audio content. The dataset was used in the Emotion and Theme Recognition in Music Task within MediaEval from 2019 through 2021, and it has a Zenodo DOI (10.5281/zenodo.3826813) for formal citation in research publications.

The repository also documents related datasets from the same research group. Song Describer provides natural-language descriptions for a subset of the tracks. MuChoMusic contains multiple-choice listening comprehension questions. ManyMusic offers additional annotations. These companion datasets share track IDs with MTG-Jamendo, enabling cross-dataset experiments that combine audio tags with textual descriptions or evaluation questions.

## Repository Structure and Metadata Files

The repository does not contain audio files. What it provides is the metadata, scripts, and pre-built splits that define how the dataset is used. The data/ directory holds the core metadata files.

Raw metadata goes through several postprocessing stages, each producing a named file. The base file raw.tsv has 56,639 entries. After filtering for tracks longer than 30 seconds, cleaning tag synonyms using tag_map.json, and requiring at least 50 unique artists per tag, the file autotagging.tsv emerges with 55,609 entries and 195 tags. This is the standard file for auto-tagging experiments.

Subsets built from autotagging.tsv include autotagging_top50tags.tsv (54,380 tracks, 50 tags), autotagging_genre.tsv (55,215 tracks, 95 tags), autotagging_instrument.tsv (25,135 tracks, 41 tags), and autotagging_moodtheme.tsv (18,486 tracks, 59 tags). The splits/ folder contains training, validation, and testing splits for each of these subsets, with consistent tag lists enforced across all splits.

The stats/ directory holds statistics on track, album, and artist counts per tag, sorted by the number of contributing artists. Additional metadata including artist names, album names, track titles, release dates, and Jamendo URLs appears in raw.meta.tsv.

The README documents a split-level constraint: a few tags are discarded in the splits to guarantee the same tag list across all splits. For the main autotagging.tsv splits, this leaves 55,525 tracks annotated by 87 genre tags, 40 instrument tags, and 56 mood/theme tags. This differs from the total counts in the full autotagging.tsv file, so researchers should use the split files rather than the parent file when training models that require consistent label dimensions. The scripts/ directory includes tools for reproducing the postprocessing pipeline, regenerating statistics, and recreating subsets from scratch.

## Setting Up the Environment and Downloading Data

The repository requires Python 3.7 or later. Clone it and set up a virtual environment:

```bash
git clone https://github.com/MTG/mtg-jamendo-dataset.git
cd mtg-jamendo-dataset
```

```bash
python3 -m venv venv
source venv/bin/activate
pip install -r scripts/requirements.txt
```

Downloading the audio uses the download.py script in scripts/download/. Pass the -h flag to see all options:

```bash
python3 scripts/download/download.py -h
```

To download audio for the mood/theme subset, unpack TAR archives, and remove the archives after unpacking:

```bash
python3 scripts/download/download.py --dataset autotagging_moodtheme --type audio /path/to/download --unpack --remove
```

The script supports two mirror servers for downloads (--from mtg and --from mtg-fast). Audio is available at 320 kbps MP3 or as a lower-bitrate mono version produced with LAME VBR 2. Precomputed mel-spectrograms are distributed as NumPy arrays in NPY format. Essentia-based statistical features are also available as JSON files.

The README notes that the unpacking process runs after each TAR archive is downloaded, so the peak storage requirement during download is higher than the final dataset size.

## Storage Requirements and Coverage Gaps

The full audio set for raw_30s requires 508 GB. The low-quality mono version reduces this to 156 GB. Mel-spectrograms for the same subset take 229 GB. The mood/theme subset alone, which contains 18,486 tracks, requires 152 GB for full-quality audio and 68 GB for mel-spectrograms. These figures make MTG-Jamendo unsuitable for experiments that require quick iteration or that run on machines with limited storage.

The tag distribution is not uniform. Tags that pass the 50-unique-artists threshold are included in autotagging.tsv, but many tags at the tail of the distribution have very few examples. The README notes that a few tags are discarded in the splits to keep the tag list consistent across training, validation, and testing sets. Researchers working on rare tags should examine the statistics in stats/ before designing experiments that assume all 195 tags are represented with equal density.

The audio on Jamendo is user-uploaded, which means quality is inconsistent. The dataset does not filter for production quality, recording level, or recording environment. This is realistic for a consumer tagging benchmark but is a known limitation for tasks that depend on clean audio signals.

## GTZAN as an Alternative for Smaller-Scale Experiments

The GTZAN dataset is the most commonly cited alternative for music genre classification. It contains 1,000 tracks across 10 genres, with 30-second clips at 22 kHz WAV format. Its total size is well under 1 GB, making it practical for experiments where storage is limited or where researchers want a quick benchmark. The GTZAN dataset has been widely criticized for duplicate tracks and labeling errors in the original release, but corrected versions are available and it remains the standard small-scale genre benchmark.

MTG-Jamendo differs from GTZAN in three important ways: it includes full-length tracks rather than 30-second excerpts, it has multi-label annotations rather than single-class genre labels, and it covers instrument and mood tags in addition to genre. Researchers who need multi-label auto-tagging at scale should use MTG-Jamendo. Researchers who need a single-label genre classification benchmark with a familiar setup and manageable size are better served by GTZAN.

## License and Maintenance

The code in this repository is released under the Apache-2.0 license. The audio tracks themselves are distributed by Jamendo under their respective Creative Commons licenses. The repository includes an audio_licenses.txt file and a CITATION.bib file for academic citation.

The last push to this repository was on 2026-03-18. The repository has no GitHub releases. The project does not document a roadmap or planned updates. Researchers adopting the dataset for a new project should verify that the download script's Jamendo endpoint still works before committing to it, since the repository is not receiving active maintenance and the external download infrastructure could change without notice.

## Conclusion

MTG-Jamendo is the right dataset for researchers who need full-audio tracks with multi-label tagging across genre, instrument, and mood categories at scale. It is not the right choice for a project that needs a small, easily downloadable benchmark: the full audio set requires 508 GB of storage. The last push to this repository was on 2026-03-18, so verify whether the scripts still work against the current Jamendo download endpoint before starting a new project that depends on them.

## FAQ

### Is it legal to download music from Jamendo for research?

The audio in MTG-Jamendo is drawn from Jamendo tracks released under Creative Commons licenses. The repository includes an audio_licenses.txt file that records the specific license for each track. The README notes the dataset is built using music available at Jamendo under Creative Commons licenses, but does not give legal advice; users should verify the specific license terms for each track before redistribution.

### What file should researchers use as the base for MTG-Jamendo auto-tagging experiments?

The README identifies autotagging.tsv as the base file for auto-tagging. It is the result of all postprocessing steps: filtering for tracks over 30 seconds, cleaning tag synonyms, and requiring at least 50 unique artists per tag. The splits/ folder contains training, validation, and testing sets built from this file.

### What audio formats does MTG-Jamendo provide?

The README describes three audio formats: 320 kbps MP3 (full quality), a lower-bitrate mono version converted using LAME VBR 2, and precomputed mel-spectrograms as NumPy arrays in NPY format. Essentia-based statistical features in JSON format are also available.

## Sources

- [Issues](https://github.com/MTG/mtg-jamendo-dataset/issues)
- [License: Apache-2.0](https://github.com/MTG/mtg-jamendo-dataset/blob/master/LICENSE)
- [MTG/mtg-jamendo-dataset on GitHub](https://github.com/MTG/mtg-jamendo-dataset)
- [README](https://github.com/MTG/mtg-jamendo-dataset/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/mtg-mtg-jamendo-dataset
