Library / SDK
acids-ircam/RAVE avatar
acids-ircam/RAVE

RAVE's config matrix runs from a Raspberry Pi to 32GB of VRAM

Official implementation of the RAVE model: a Realtime Audio Variational autoEncoder

1,813 stars224 forksPythonNOASSERTION

At a glance

What is it?
RAVE is the official implementation from IRCAM's ACIDS team of a variational autoencoder for neural audio synthesis, and its real interface is a table of eight model configurations with stated GPU memory floors, from a 5GB Raspberry Pi profile to a 32GB style-transfer model, combined with switchable objectives, discriminators and augmentations. The detail worth knowing before you train anything is that the export step has a flag, and forgetting it produces audible artefacts.
Who is it for?
Adopt RAVE if you want a real-time neural audio model you can train on your own data and drive from a DAW or Max, and if a GPU with 16GB or more is available for the v2 configuration the project recommends. Do not adopt it for discrete token modelling without reading the redirect to msprior, which the README calls experimental, and do not adopt it on a modern stack without checking the pinned pytorch_lightning==1.9.0 and scipy==1.10.0 in requirements.txt.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Activity is slowing. The repository last received commits 7 months ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 9, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The configuration table is the interface

RAVE's user-facing surface is a table of named configurations, each with a stated minimum GPU memory, and reading it tells you what the project is for. There are eight architectures. v1 is the original continuous model at 8GB. v2 is the improved continuous model, described as faster and higher quality, at 16GB. v2_small is v2 with a smaller receptive field, adapted adversarial training and a noise generator, adapted for timbre transfer on stationary signals, back down to 8GB. v2_nopqmf drops the pseudo-quadrature filter mirror from the generator and is marked experimental, at 16GB, with the note that it is more efficient for bending. v3 is v2 with Snake activation, a descript discriminator and adaptive instance normalisation for real style transfer, and it asks for 32GB. The discrete model, described as similar to SoundStream or EnCodec, needs 18GB. onnx is a noiseless v1 configuration for ONNX use at 6GB, and raspberry is a lightweight configuration for real-time inference on a Raspberry Pi 4 at 5GB. That range is the honest scope of the project, from a Pi to a workstation.

Objectives, discriminators and three more switches

The architecture configs are only the first column. Several of them combine with a second set, and the README says many configuration files live in rave/configs and can be combined, which is what makes the CLI composable rather than a menu. Regularization applies to the v2 family only and there are three objectives: default, which is the variational autoencoder objective, the ELBO; wasserstein, a Wasserstein autoencoder objective using MMD; and spherical. A single discriminator option, spectral_discriminator, brings in the multi-scale discriminator from EnCodec. Three more switches cover the parts people forget. causal swaps in causal convolutions, which is how you lower the model's overall latency. noise enables the version 2 noise synthesiser. hybrid enables a mel-spectrogram input, which is the way to condition the model on something other than raw audio. Training is a single command with those flags appended, for example a v2 run with two augmentations:

bash
rave train --config v2 --augment mute --augment compress

The three augmentations were new in 2.3 and exist for low-data regimes: mute randomly mutes data batches at a default probability of 0.1 so the model learns silence, compress randomly compresses the waveform like a light non-linear amplification, and gain applies a random gain with a default range of -6 to 3. These change training only, not the exported model.

Preprocess, train, export

The README describes training as three separate steps, and keeping them separate is what makes the lazy option possible. Preprocessing turns a folder of audio into a dataset, and the command names an input path, an output path and a channel count:

bash
rave preprocess --input_path /audio/folder --output_path /dataset/path --channels X

The alternative is a lazy mode, which lets RAVE train directly on raw files such as mp3 and ogg without converting them first. The trade-off is spelled out: lazy loading increases CPU load by a large margin during training, especially on Windows, and is useful when a large corpus would not fit on a hard drive uncompressed. So the decision is between disk space and CPU time, and on a laptop with a big library of lossy files the lazy path is the only one that works. Training then points at the dataset and an output directory, with a name and a channel count, and export turns a finished run into a TorchScript file:

bash
rave export --run /path/to/your/run (--streaming)

A Colab notebook for training version 2 also exists, contributed by hexorcismos, which removes the GPU requirement from a first experiment entirely.

The streaming flag is the whole realtime story

Export has one option and the README attaches a warning to it with unusual clarity. Setting the streaming flag enables cached convolutions, making the model compatible with real-time processing, and the next sentence says that if you forget to use streaming mode and try to load the model in Max, you will hear clicking artefacts. That is the single most useful sentence in this documentation, because the failure is audible rather than an exception, and it would be easy to spend an afternoon blaming the model. Cached convolutions trade a bounded amount of extra computation for the ability to reuse previous activations, which is what turns a model that must recompute an entire receptive field per sample into one that can be called sample by sample from a live audio callback. Everything the project claims about being a real-time audio model rests on that mechanism, which is also why the raspberry configuration exists as a separate lightweight profile, and why the causal option exists: non-causal convolutions are the default because they are better for quality, and causal is the compromise you make for latency.

Discrete models are a redirect, and v1 is a branch

Two scope boundaries are stated plainly. The first is discrete modelling. If you want a token-based codec in the style of SoundStream or EnCodec, the README redirects you to the separate msprior library and adds that it is still experimental, which means the discrete configuration in this repository is not the supported path for that work. The second is history. The original implementation of the model can be restored with a single checkout of the v1 branch, and since the v2 configuration is described as the improved version of v1, the project has kept the first generation runnable rather than deleting it. That is worth something for reproducibility, since a paper's code is easier to cite when the code that produced the paper's numbers is still a command away. The v1 branch is also the reference for what the v2 changes actually were, which is the kind of information a changelog gives you badly and a branch gives you well. The rest of the usage path is outside this repository: a VST plugin for Windows, Mac and Linux is distributed as a beta from the IRCAM forum, with issues taken both here and on the forum's discussion page.

Requirements that are pinned for a reason, and one that is not

Installation is one command, and the README attaches a warning to it that most projects would bury.

bash
pip install acids-rave

You are strongly advised to install torch and torchaudio before acids-rave, so that you can choose the appropriate version of torch for your device, and the package no longer enforces a specific torch version for future compatibility with new devices and modern Python environments. ffmpeg is also required, with conda install ffmpeg given as the way to get it inside a virtual environment. The dependency list explains what you are agreeing to: pytorch_lightning pinned to 1.9.0 and scipy pinned to 1.10.0, both old, plus librosa, cached-conv, nn-tilde for the cached convolution work, GPUtil for device monitoring, gin-config because the configuration files are gin files, udls for downloading datasets, tensorboard for logging, and Flask, which is a hint that something in the project serves over HTTP. Those two exact pins are the compatibility risk to check first on a modern interpreter, since a 2023-era training framework is where a fresh environment will break first.

Licence, citation and a two-year gap

The licence needs a moment of attention because the two signals disagree. setup.py lists the package as MIT in its trove classifier, and there is a LICENSE file at the repository root, while the repository's own licence metadata is recorded as a custom licence that could not be classified. The classifier is not the licence and the file is, so if you are building on this, read the LICENSE file rather than the metadata, and reconcile it with the classifier. Separately, the README asks that if you use RAVE as part of a music performance or installation, you cite either the repository or the article, which is a request rather than a licence condition but a sensible one for a research tool that ends up on stage. The maintenance picture is the other thing to note. There is a single release, v2.3.1, dated 2023-12-18, and the last push to the repository is 2026-03-07. So the code is being touched, tutorials are being published on the IRCAM forum, and no version has been cut in between, which means a new user installs a two-year-old package and should expect to read the source before trusting it with a new environment.

Against a hosted model, and against building from the paper

The alternatives split cleanly. A hosted neural audio service gives you a voice or an instrument in a browser with no GPU and no training, and you get none of the model, which is the thing RAVE exists to give you. A general deep learning framework plus the paper gives you full control and full responsibility, and for a lab that is often the right answer, since the architecture is described in a paper rather than hidden in a repository. RAVE sits between them as a research implementation with a training pipeline, a configuration matrix, an export path and a plugin, and that middle position is exactly where its value is: it saves you the parts that are not the research contribution, namely dataset preparation, a real-time-capable export and a DAW-facing wrapper. What it does not save you is the choice of objective, the latency trade-off or the training data, and the README's own warnings about lazy loading, GPU memory floors and the streaming flag are the places where that work has been moved rather than removed.

Editorial conclusion

Adopt RAVE if you want a real-time neural audio model you can train on your own data and drive from a DAW or Max, and if a GPU with 16GB or more is available for the v2 configuration the project recommends. Do not adopt it for discrete token modelling without reading the redirect to msprior, which the README calls experimental, and do not adopt it on a modern stack without checking the pinned pytorch_lightning==1.9.0 and scipy==1.10.0 in requirements.txt. Verify four things first: that you have ffmpeg available, that you installed torch and torchaudio before installing the package as the README strongly advises, that your model configuration fits the memory you actually have, since the floors range from 5GB to 32GB, and that you export with --streaming, because the README says a model exported without it produces clicking artefacts when loaded in Max. There is one release, v2.3.1 from 2023-12-18, and the last push was 2026-03-07.

Frequently asked questions

How do I install RAVE?

With pip install acids-rave. The README strongly advises installing torch and torchaudio first so you can pick the version matching your device, and the package no longer enforces a specific torch version. ffmpeg is also required, installable with conda install ffmpeg.

How much GPU memory does RAVE need?

It depends on the configuration. The README's table gives floors of 8GB for v1 and v2_small, 16GB for v2 and v2_nopqmf, 18GB for the discrete model, 6GB for the onnx configuration and 5GB for the Raspberry Pi one, while v3 asks for 32GB.

What does the RAVE streaming export flag do?

It enables cached convolutions, which makes the model compatible with real-time processing. The README warns that a model exported without it will produce clicking artefacts when loaded in Max.

What is lazy dataset preprocessing in RAVE?

It lets RAVE train directly on raw files such as mp3 and ogg without converting them first. The README warns it increases CPU load by a large margin during training, especially on Windows, and that it is useful for large corpora that would not fit uncompressed on a hard drive.

Can RAVE train discrete audio tokens?

There is a discrete configuration described as similar to SoundStream or EnCodec, but for discrete models the README redirects users to the separate msprior library, which it describes as still experimental. For audio synthesis, the continuous v2 configuration is the recommended path.

What licence is RAVE released under?

The package metadata in setup.py lists the MIT licence in its trove classifier and a LICENSE file sits at the repository root, while the repository licence metadata is recorded as unclassified, so read the file rather than the metadata. The only release is v2.3.1 from 2023-12-18.

Official sources

  1. acids-ircam/RAVE on GitHub
  2. Issues
  3. README
  4. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/acids-ircam-rave.svg)](https://hysenlabs.com/projects/acids-ircam-rave)