# AudioCraft: MusicGen, AudioGen and EnCodec in one PyTorch library

> AudioCraft is Meta's research library for audio generation and compression, bundling MusicGen, AudioGen, EnCodec, MAGNeT, AudioSeal and JASCO. It is a training and inference codebase first, and a usable pip package second.

**facebookresearch/audiocraft** — Audiocraft is a library for audio processing and generation with deep learning. It features the state-of-the-art EnCodec audio compressor / tokenizer, along with MusicGen, a simple and controllable music generation LM with textual and melodic conditioning.

- Repository: https://github.com/facebookresearch/audiocraft
- Stars: 23,651 · Forks: 2,697
- Language: Jupyter Notebook
- License: MIT
- Published: 2026-09-21 · Updated: 2026-09-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/facebookresearch-audiocraft

## What AudioCraft actually is, and who it is built for

AudioCraft is a PyTorch library for deep learning research on audio generation. That word, research, sets the boundary. The repository ships inference and training code for a set of models rather than a finished product: MusicGen for text-to-music, AudioGen for text-to-sound, EnCodec as the neural audio codec underneath, Multi Band Diffusion as an EnCodec-compatible diffusion decoder, MAGNeT as a non-autoregressive model for text-to-music and text-to-sound, AudioSeal for watermarking, MusicGen Style for text-and-style conditioning, and JASCO for text-to-music conditioned on chords, melodies and drum tracks.

The intended user is someone who wants to run or retrain these models from Python, not someone who wants a browser tab that produces a song. If your task is generating a short piece of music from a text prompt inside a script, or fine-tuning a compression model on your own audio, the layout makes sense. If your task is shipping a consumer audio product, the licence on the weights, described below, will probably decide the question before the code does.

The primary language of the repository is Jupyter Notebook, which is a fair signal of how the project is meant to be read: notebooks and documentation pages that walk through each model, with the library underneath providing the reusable pieces. The README points to per-model documents under docs/ for MUSICGEN, AUDIOGEN, ENCODEC, MBD, MAGNET, WATERMARKING, MUSICGEN_STYLE and JASCO, so the entry point for any real work is one of those files rather than the top-level README.

## How the pieces fit: EnCodec as the substrate, language models on top

The architecture visible in the repository is layered. EnCodec is a neural audio codec and tokenizer: it compresses audio into discrete tokens and reconstructs it. Everything generative sits above that representation. MusicGen is described in the README as a controllable music generation language model with textual and melodic conditioning, and AudioGen is the equivalent for sound effects. Because the generative models operate on EnCodec tokens rather than raw waveforms, the codec is a shared dependency across the stack, and the Makefile makes this concrete: the integration targets for MusicGen and AudioGen both pass compression_model_checkpoint=//sig/5091833e, reusing one trained compression model.

The training side is organised around Hydra configs and Dora, Meta's experiment runner. The repository has config/, egs/, dataset/, scripts/ and a Makefile whose integration targets invoke python3 -m dora run with overrides for solver, dataset and optimiser settings. That is the data flow a researcher works with: a solver config selects the model, a dataset config selects the audio, and a checkpoint from a previous stage can be injected as a dependency. If you only want inference, none of this matters; if you want to reproduce or extend training, this structure is the whole interface, and the README points to docs/TRAINING.md for the design principles.

One detail worth noting for anyone planning to extend the stack: the Makefile sets AUDIOCRAFT_DORA_DIR to a temporary directory per user, which suggests training runs are expected to be disposable and reproducible rather than long-lived stateful jobs.

## Installing AudioCraft and generating your first clip

The README pins the environment tightly: Python 3.9 and PyTorch 2.1.0. It advises installing torch first, in particular before xformers, and only if PyTorch is not already present. The stable release is on PyPI.

```bash
python -m pip install 'torch==2.1.0'
python -m pip install setuptools wheel
python -m pip install -U audiocraft
```

If you cloned the repository and intend to train, the README requires an editable install instead, and adds an extra for watermarking work:

```bash
python -m pip install -e .
python -m pip install -e '.[wm]'
```

ffmpeg is recommended as well, either from the system package manager or from conda:

```bash
sudo apt-get install ffmpeg
conda install "ffmpeg<5" -c conda-forge
```

Model weights are downloaded from Hugging Face. The cache location can be overridden with the AUDIOCRAFT_CACHE_DIR environment variable, and the README notes that models relying on Demucs, such as musicgen-melody, follow Torch Hub's download location instead. The README does not give a runnable inference snippet in the top-level document; the per-model pages under docs/ are where the usage examples live, so a first real run means opening docs/MUSICGEN.md or docs/AUDIOGEN.md and following the example there. Expect the first invocation to spend its time downloading weights.

## The licence split is the constraint that matters most

The repository separates two licences explicitly. The code is MIT, as stated in the LICENSE file and in setup.py. The model weights are released under CC-BY-NC 4.0, as stated in LICENSE_weights. That is a non-commercial licence, and it applies to the artefacts you actually want to use for generation. A permissive code licence does not grant you commercial rights to the weights, and the README presents the split without further commentary.

For an academic group or a hobbyist, this is unremarkable. For a company evaluating AudioCraft as the basis of a product, it is the first thing to resolve, and it is why the answer to whether AudioCraft is free depends on what you mean: the code is free under MIT, the weights carry a non-commercial condition. There is no separate commercial licensing path described in the README, and no statement about what happens to audio you generate. This is not legal advice; read LICENSE and LICENSE_weights yourself before deciding.

## Where AudioCraft is the wrong tool

AudioCraft is a research library, and several ordinary engineering expectations are simply not addressed. The README documents no rollback or version-pinning story beyond installing a specific release from PyPI, and no releases are listed alongside the project, so tracking what changed between versions means reading CHANGELOG.md rather than relying on semantic versioning guarantees.

The dependency pins are aggressive: torch==2.1.0, torchvision==0.16.0, torchtext==0.16.0, torchaudio>=2.0.0,<2.1.2, numpy<2.0.0, spacy==3.7.6, xformers<0.0.23, av==11.0.0. If your project already depends on a newer PyTorch or on NumPy 2.x, installing AudioCraft means either a separate environment or a downgrade. That is a normal cost for research code, but it is a real integration cost and the README does not present it as optional.

The README also says nothing about Windows. Installation instructions use apt-get and conda, and the environment assumptions read as Linux or macOS. People searching for how to install AudioCraft on Windows will not find an answer in the documentation provided. If you need a hosted endpoint, a web UI, or a service with an uptime commitment, this is not that project; the demos/ directory and the Gradio dependency suggest local demo apps, not a deployment target.

## AudioCraft compared with Suno and with MusicGen alone

The comparison people reach for is AudioCraft versus Suno, and the difference is categorical rather than incremental. Suno is a hosted product: you give it a prompt and it returns finished audio, with no code, no checkpoint and no training loop in your hands. AudioCraft gives you the library and the weights. You run the model on your own hardware, you control the sampling parameters, and you can fine-tune or train from the provided pipelines. The trade is that you also own the environment problems: the pinned torch version, ffmpeg, the weight download, and the licence question. If you want a song in a browser, the hosted product wins. If you want MusicGen inside a batch job that processes ten thousand prompts, AudioCraft is the only one of the two that gives you that.

The other comparison is AudioCraft versus MusicGen, which is a category error worth clearing up: MusicGen is one of the models inside AudioCraft, not a separate project. The README lists MusicGen alongside AudioGen, EnCodec, Multi Band Diffusion, MAGNeT, AudioSeal, MusicGen Style and JASCO, all as components of the same library. Choosing AudioCraft and choosing MusicGen are the same decision unless you intend to work with one of the other models.

## Maintenance, training cost and the upgrade path

The repository is not archived, and the last push was on 2026-03-03. That is the concrete maintenance fact available; the README makes no commitment about release cadence or support.

The upgrade path is a reinstall. The README offers a stable release through pip and a bleeding-edge install from the git URL, and the editable install for local development. Because the requirements file pins exact versions of torch, torchvision, torchtext, spacy, av and xformers, an upgrade is not a matter of bumping one number; it is a coordinated change, and the Makefile's integration targets exist to check that the pieces still work together on CPU with tiny sample counts. Running make tests_integ is the closest thing the repository offers to a compatibility check, and it covers compression, diffusion, MusicGen, AudioGen and watermarking in sequence.

Training cost is not quantified anywhere in the README, and the Makefile's integration targets deliberately shrink the work to ten training samples and one epoch. Anyone planning a real training run should treat the per-model docs as the source for configuration and expect to supply their own hardware budget.

## Conclusion

Adopt AudioCraft if you need training code for EnCodec, MusicGen, Multi Band Diffusion or JASCO, or if you want MusicGen and AudioGen inference inside a Python pipeline you control. Do not adopt it if you need a permissive licence for the model weights, a hosted service, or a supported Windows install path. Before you commit, check your Python and torch versions against the pinned requirements, confirm ffmpeg is present on the machine, and read LICENSE_weights to see whether CC-BY-NC 4.0 fits how you intend to use the generated audio.

## FAQ

### What is AudioCraft AI?

It is a PyTorch library for deep learning research on audio generation, containing inference and training code for AudioGen and MusicGen along with EnCodec, Multi Band Diffusion, MAGNeT, AudioSeal, MusicGen Style and JASCO.

### How do I install AudioCraft?

The README requires Python 3.9 and PyTorch 2.1.0, advises installing torch first and before xformers, then installing the stable release with python -m pip install -U audiocraft, or an editable install with python -m pip install -e . if you cloned the repository and want to train.

### How do I use AudioCraft?

The top-level README does not include a runnable inference example; it points to per-model documentation such as docs/MUSICGEN.md and docs/AUDIOGEN.md, which hold the usage instructions and configuration for each model.

### Is AudioCraft free?

The code in the repository is released under the MIT license, while the model weights are released under CC-BY-NC 4.0, a non-commercial license, so the answer depends on whether you mean the code or the weights.

### What is the difference between AudioCraft and MusicGen?

MusicGen is one of the models contained in AudioCraft, listed alongside AudioGen, EnCodec, Multi Band Diffusion, MAGNeT, AudioSeal, MusicGen Style and JASCO, so it is a component rather than a separate library.

## Sources

- [facebookresearch/audiocraft on GitHub](https://github.com/facebookresearch/audiocraft)
- [Issues](https://github.com/facebookresearch/audiocraft/issues)
- [License: MIT](https://github.com/facebookresearch/audiocraft/blob/main/LICENSE)
- [README](https://github.com/facebookresearch/audiocraft/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/facebookresearch-audiocraft
