stable-audio-tools: training and inference code for Stability AI audio models
Generative models for conditional audio generation
At a glance
- What is it?
- The repository is a Python toolkit for training and fine-tuning generative audio models and running them locally through a Gradio interface. It is a research-grade codebase: JSON configs, PyTorch Lightning, Weights & Biases logging, and an explicit unwrap step before any checkpoint can be used for inference.
- Who is it for?
- Adopt stable-audio-tools if you already work in PyTorch and need to fine-tune or train a conditional audio model on your own dataset, and you accept that a Weights & Biases account, a model config, a dataset config and an unwrap step sit between you and a usable checkpoint. Do not adopt it if you want a one-command audio generator: the repository ships no released checkpoints, no versioned releases, and no CLI for batch generation.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 14 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What stable-audio-tools is for, and who it is not for
stable-audio-tools is the code behind Stability AI's audio generation models, packaged as a repository you clone rather than a service you call. The README describes it as "Training and inference code for audio generation models", and the pyproject.toml description is narrower still: "Training and inference tools for generative audio models from Stability AI".
The audience is someone who has a dataset of audio and a reason to condition a generative model on it, or someone who wants to run a published Stable Audio checkpoint locally instead of through a hosted endpoint. Both paths exist in the repository: train.py for training and fine-tuning, run_gradio.py for a local browser interface over a pretrained model.
It is not a library you import to add audio generation to an application. There is no documented Python API surface for inference in the README, no versioned release, and the pyproject.toml version is 0.0.20. If you want a drop-in generator, this repository is the wrong layer; it is the layer underneath one.
The config-driven model architecture
Everything in stable-audio-tools is defined by JSON configuration files. The README states that "Training and inference code for stable-audio-tools is based around JSON configuration files that define model hyperparameters, training settings, and information about your training dataset." Two files matter: a model config and a dataset config. train.py takes both as flags.
The model config's top level carries a model_type key, and the README limits it to one of "autoencoder", "diffusion_uncond", "diffusion_cond", "diffusion_cond_inpaint", "diffusion_autoencoder" or "lm". That list is the actual shape of the project: it is a family of architectures, not one model. A latent diffusion setup is expressed as a diffusion model plus a separately trained autoencoder used as a pretransform, which is why the training scripts accept --pretransform-ckpt-path.
The second config key the README documents is sample_size, the length of audio fed to the model during training, measured in samples. Because that value lives in a config rather than in code, changing the training segment length is a config edit, and a checkpoint trained at one sample_size is tied to that choice.
Training itself runs through PyTorch Lightning. During training the model is wrapped in what the README calls a "training wrapper", a pl.LightningModule holding discriminators, EMA copies and optimizer state. That wrapper is what makes the training loop multi-GPU and multi-node capable, and it is also what makes checkpoints large, which brings us to the unwrap step.
Installing stable-audio-tools with uv and running the Gradio UI
The repository uses uv for dependency management. The README says to install it with pip install uv, then clone and sync. Three sync targets exist: inference only, training, and everything including the Gradio UI. Development is done in Python 3.10, and pyproject.toml pins requires-python to ">=3.10,<3.11", so a 3.11 or 3.12 interpreter will not resolve.
git clone https://github.com/Stability-AI/stable-audio-tools.git
cd stable-audio-tools
# Inference only
uv sync
# Training
uv sync --extra train
# Everything (training + Gradio UI)
uv sync --extra train --extra uiScripts then run through uv run. The README's example for the Gradio interface uses the stable-audio-open-1.0 checkpoint, and it notes that you must accept the model's terms on Hugging Face first.
uv run python run_gradio.py --pretrained-name stabilityai/stable-audio-open-1.0If you prefer pip, the README gives a direct install for the training extra:
pip install "stable-audio-tools[train]"Flash Attention is recommended for performance but is not part of the sync. The README says to follow the installation instructions in the Flash Attention repository after running uv sync, which means it is a manual step you own. Note also the uv source configuration in pyproject.toml: torch and torchaudio are pulled from a pytorch-cu126 index, but only for Linux on x86_64. On other platforms the default index applies and you should expect a different build.
Training a model and the unwrap step nobody skips
Training requires a Weights & Biases account; the README instructs you to create one and run wandb login before starting. The training command takes a dataset config, a model config and a project name:
python3 ./train.py --dataset-config /path/to/dataset/config --model-config /path/to/model/config --name harmonai_trainThe --name flag sets the Weights & Biases project name. Useful defaults are documented: --checkpoint-every defaults to 10000 steps, --batch-size to 8 samples per GPU, --num-gpus and --num-nodes to 1, and --precision to 16. --accum-batches exists for gradient accumulation when VRAM is tight, and --strategy deepspeed enables DeepSpeed ZeRO Stage 2; otherwise the strategy defaults to ddp when more than one GPU is used.
The design decision with the largest practical consequence is checkpoint wrapping. Checkpoints written during training contain the LightningModule wrapper, so they are larger than the model itself, and they cannot be used as-is. The README is explicit that unwrapped checkpoints are required for inference scripts, for use as a pretransform in another model, and for fine-tuning a pretrained model with a modified configuration.
python3 ./unwrap_model.py --model-config /path/to/model/config --ckpt-path /path/to/wrapped/ckpt --name model_unwrapFine-tuning then splits two ways. Continuing an existing run from a wrapped checkpoint uses --ckpt-path on train.py. Starting fresh from a pretrained unwrapped model uses --pretrained-ckpt-path. Mixing those two up is an easy way to spend an afternoon on a failed load, because the wrapper and the unwrapped file are not interchangeable.
Where stable-audio-tools will not help you
The repository ships no model weights. run_gradio.py with --pretrained-name expects a Hugging Face repository name, and the README's own example points at stabilityai/stable-audio-open-1.0, whose terms you must accept on Hugging Face before the download works. If you have no checkpoint and no dataset, the install gets you a training harness and nothing to run.
There are no retrieved releases for the repository, and the package version in pyproject.toml is 0.0.20. If your process depends on semantic versioning or a changelog, neither is available here; you pin a commit or you pin the version string and hope it moves predictably.
The dependency set is opinionated and heavy. torch and torchaudio are pinned to 2.7.1, transformers is unpinned, and the training extra pulls pytorch_lightning 2.5.5, wandb 0.15.4, webdataset 0.2.100 and several codec packages including encodec and descript-audio-codec. Several pins are exact (sentencepiece==0.1.99, PyWavelets==1.4.1, vector-quantize-pytorch==1.14.41). That is a reproducibility choice, and it means this repository will not simply ride along inside a larger environment that wants different versions of those same packages.
Finally, the README documents no rollback or resume semantics for interrupted training beyond the checkpoint flags, and it does not document a batch or headless inference entry point. The Gradio script is the documented inference path.
How it compares to a general audio generation library
The closest point of comparison in the ecosystem is a general-purpose audio generation library that exposes a Python API and wraps pretrained models for immediate use. The difference is where the abstraction sits. A library like that treats the model as a fixed artifact and the caller as the variable: you import, you pass a prompt, you get audio.
stable-audio-tools inverts that. The model is the variable, defined by a JSON config whose model_type selects one of six architecture families, and the caller is expected to supply the dataset, the training loop settings and the checkpoint lifecycle. The unwrap_model.py step has no equivalent in a library that only does inference, because a library never produces a wrapped checkpoint in the first place.
The trade-off is direct. You get the ability to train an autoencoder, use it as a pretransform for a latent diffusion model, and fine-tune the decoder independently, which is what --pretransform-ckpt-path is for. You give up the ability to treat audio generation as a function call. If your task is "generate a clip from a text prompt", the library is the shorter path. If your task is "train something on my audio", this repository is the only one of the two that offers it.
Licence and the cost of keeping a fork current
The repository is MIT licensed, and the top-level tree contains a LICENSES/ directory alongside LICENSE, which suggests bundled components carry their own terms. The README does not enumerate those, so if you redistribute the package or a trained model, read that directory rather than assuming MIT covers everything in the tree.
The licence on the code says nothing about the weights. The README's example checkpoint, stabilityai/stable-audio-open-1.0, sits behind a Hugging Face terms acceptance step, and those terms are separate from the MIT grant on this repository. Treat them as two different questions.
Maintenance cost is mostly dependency drift. The last push to main was on 2026-09-18, so the repository is current, but there are no retrieved releases to diff against and the version string is 0.0.20. An upgrade means reading commits. The exact pins in pyproject.toml cut both ways: they protect your training run from a surprise in vector-quantize-pytorch, and they make it your job to unpin deliberately when you need a newer torch. The uv.lock file in the tree is what makes the sync reproducible; regenerating it is the upgrade.
Editorial conclusion
Adopt stable-audio-tools if you already work in PyTorch and need to fine-tune or train a conditional audio model on your own dataset, and you accept that a Weights & Biases account, a model config, a dataset config and an unwrap step sit between you and a usable checkpoint. Do not adopt it if you want a one-command audio generator: the repository ships no released checkpoints, no versioned releases, and no CLI for batch generation. Before committing, verify that your Python version is 3.10, that PyTorch 2.7.1 and torchaudio 2.7.1 resolve on your platform, and that you have accepted the terms for stabilityai/stable-audio-open-1.0 on Hugging Face, because run_gradio.py cannot load it otherwise.
Frequently asked questions
How do I install stable-audio-tools?
Install uv with pip install uv, clone the repository, and run uv sync for inference only, uv sync --extra train for training, or uv sync --extra train --extra ui for everything including the Gradio interface. The project requires Python 3.10, and pip install "stable-audio-tools[train]" is documented as an alternative.
Does stable-audio-tools include pretrained models?
No. run_gradio.py takes a --pretrained-name pointing at a Hugging Face repository, and the README's example is stabilityai/stable-audio-open-1.0, whose terms you must accept on Hugging Face before it will load.
Why can't I use my training checkpoint for inference in stable-audio-tools?
Checkpoints written during training contain a PyTorch Lightning training wrapper with discriminators, EMA copies and optimizer state, which makes them larger than the model. The README states that unwrapped checkpoints are required for inference scripts, for use as a pretransform, and for fine-tuning with a modified configuration, and unwrap_model.py produces them.
What Python and PyTorch versions does stable-audio-tools need?
pyproject.toml pins requires-python to >=3.10,<3.11 and pins torch and torchaudio to 2.7.1. The README says PyTorch 2.5 or later is required for Flash Attention and Flex Attention support, and that development is done in Python 3.10.
Is stable-audio-tools free to use?
The repository is MIT licensed, with a separate LICENSES/ directory in the top-level tree for bundled components. That licence covers the code, not the model weights: the README's example checkpoint requires accepting terms on Hugging Face separately.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/stability-ai-stable-audio-tools)