DDSP: a differentiable DSP library for audio generation, not a VST plugin
DDSP: Differentiable Digital Signal Processing
At a glance
- What is it?
- Magenta's DDSP wraps synthesizers, filters and losses as TensorFlow layers so a network can output audio directly. It is a research library, and the search traffic around it is mostly about a plugin it is not.
- Who is it for?
- Adopt DDSP if you are building a model whose output layer is audio and you want interpretable DSP in the middle: the Processor API, the Harmonic synth and the loss functions are the parts worth reading first. Do not adopt it if what you actually want is the DDSP-VST plugin, which this repository does not contain, or if you need a library that tracks current TensorFlow, since setup.py pins tensorflow<=2.11, numpy<1.24 and protobuf<=3.20.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 5 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What DDSP actually is, and who the library is for
DDSP is a library of differentiable versions of common DSP functions: synthesizers, waveshapers and filters. The README's framing is that these interpretable elements can be used inside a deep learning model, especially as the output layers for audio generation. So the intended user is someone training a model that produces sound, not someone looking for an instrument to play.
The README's own example is three lines of intent: get synthesizer parameters from a neural network, initialize a Harmonic synth, generate audio from amplitudes, harmonic distribution and f0_hz. That is the whole pitch. A network predicts physically meaningful parameters, and a DSP block turns them into a waveform.
This matters because the search traffic around the name is dominated by plugin questions. People type "magenta ddsp vst", "DDSP-VST models", "magenta ddsp plugin free download". Those queries are about a separate product. This repository is the Python library, and nothing in the README, the module list or setup.py describes a plugin binary, a VST bundle or a downloadable instrument. If you arrived here looking for the plugin, you are in the wrong repository.
Who it is for, concretely: researchers and engineers who already work in TensorFlow, who want the output layer of a generative audio model to be a synthesizer they can reason about rather than a raw waveform decoder.
The Processor contract: inputs become controls, controls become signal
The central abstraction is the Processor. It inherits from tfkl.Layer, so it composes like any other Keras layer, but it does something other layers do not: it formats its inputs into controls that are physically meaningful. The README's example is a synthesizer removing frequencies above the Nyquist frequency to avoid aliasing, or making amplitudes strictly positive.
The split is three methods. get_controls() maps inputs to controls. get_signal() maps controls to a signal. __call__() is get_signal(**get_controls()), so calling the processor runs both stages. Inputs are a variable number of tensors, usually network outputs. Controls are a dictionary of tensors scaled and constrained for that processor. Signal is the output tensor, usually audio or a control signal for another processor.
That intermediate dictionary is the design decision worth noticing. Because controls are explicit and constrained, you can inspect them between the two stages, and the README's worked example shows why you would want to: with a 16 kHz sample rate and an 8 kHz Nyquist limit, only 18 harmonics of the harmonic distribution remain nonzero after the processor removes the ones above Nyquist and normalizes the rest. The constraint is not hidden inside a black box.
The cost is that you must know what controls each processor expects. The README points to the test file in each module as the place to find usage examples, which is a fair signal that the prose documentation is thinner than the code.
Installing DDSP and generating audio in a first script
Installation is short. The README requires TensorFlow version >= 2.1.0 and notes the core library runs in either eager or graph mode. The system package comes first because the audio stack needs libsndfile.
sudo apt-get install libsndfile-dev
pip install --upgrade pip
pip install --upgrade ddspAfter that, the README's getting-started example is the smallest real use: a network produces a dict of parameters, a Harmonic processor turns them into audio.
import ddsp
# Get synthesizer parameters from a neural network.
outputs = network(inputs)
# Initialize signal processors.
harmonic = ddsp.synths.Harmonic()
# Generates audio from harmonic synthesizer.
audio = harmonic(outputs['amplitudes'],
outputs['harmonic_distribution'],
outputs['f0_hz'])The three keyword arguments are the contract: amplitudes, harmonic_distribution and f0_hz. What you should see is a tensor of audio. The README does not state the sample rate of that output in this snippet; the Harmonic example elsewhere uses 16 kHz, and the tutorial notebooks are where the full pipeline, including resampling, is worked through.
If you would rather not write the training loop yourself, the repository ships a self-contained training library under ddsp/training and a set of Gin configs, and setup.py declares package_data for *.gin plus a top-level update_gin_config.py script. The README directs you to the Colab demos for the guided path: the train_autoencoder demo converts audio files into a dataset and trains an autoencoder, and the timbre_transfer demo applies pretrained models to new audio.
Pinned dependencies are the real constraint on adoption
The install line is one command. The dependency set behind it is not casual. setup.py pins tensorflow<=2.11, numpy<1.24, protobuf<=3.20, scipy<=1.10.1, librosa<=0.10, pydub<=0.25.1, hmmlearn<=0.2.7, crepe<=0.0.12, note_seq<0.0.4 and tensorflowjs<3.19, and the protobuf pin carries the comment "temporary fix for proto dependency bug".
Read that as a compatibility envelope, not a suggestion. If your environment already runs a newer TensorFlow or a newer NumPy, installing ddsp into it will either downgrade those or fail to resolve. In a shared environment that is a real cost, and it is the most likely reason an adoption attempt stalls.
The dependency list also reveals scope. apache-beam, cloudml-hypertune and google-cloud-storage are present because the training library targets Google Cloud infrastructure. If you only want the core synth and effect layers, you are still installing the training stack's dependencies.
One more practical note: the README's install steps are written for a Debian-style system, since libsndfile-dev is an apt package. The README does not document a Windows or macOS equivalent, and it does not document a rollback or uninstall procedure.
Where DDSP is the wrong tool
DDSP is not a real-time audio plugin, and it is not a synthesis engine you drop into a DAW. Nothing in the README describes a VST, an audio callback, a host integration or a latency figure. The library's unit of work is a tensor inside a TensorFlow graph or an eager context. If your requirement is a playable instrument with predictable latency, this is the wrong layer, regardless of what the search results suggest.
It is also not a pretrained model you can just run. The README links to Colab demos with pretrained models for timbre transfer, but the library itself is the differentiable DSP layer. The models live in the demos and in the training library, and the README does not document a Python API for loading a released checkpoint outside that demo path.
A third boundary is the dependency pin. If you are building on current TensorFlow, the pin to <=2.11 means you are either maintaining a separate environment or waiting for the pin to move. The last tagged release listed is v3.5.1 from 2023-04-26, while the repository's last push was on 2026-09-22, so the code is moving but the release tags are not.
Finally, this is research code with a research documentation style. The README repeatedly points at test files as the source of usage examples. That is workable for someone comfortable reading tests, and frustrating for someone who wants a reference manual.
How DDSP differs from a neural vocoder
The obvious alternative approach is a neural vocoder: a model that predicts a waveform, or a spectrogram followed by a learned waveform generator, with no explicit DSP in between. The difference is what sits between the network and the sound.
In a vocoder pipeline the network's output is a representation you then invert. In DDSP the network's output is a set of parameters for a named processor, and the processor's job is to constrain those parameters into something physically sensible before synthesis. The README makes this explicit: the Harmonic synth logarithmically scales amplitudes, removes harmonics above Nyquist and normalizes the remaining distribution. Those are DSP operations, written as differentiable tensor code.
The practical consequence is interpretability and control. You can look at the controls and see how many harmonics are active, or that amplitudes are positive, because the processor guarantees it. You also get a strong prior: a harmonic model will not produce arbitrary noise, which is a benefit for pitched instruments and a limitation for anything percussive or noisy that a harmonic-plus-noise or filtered-noise model does not cover.
The trade-off is that you are choosing the synthesis model in advance. A vocoder can learn whatever the data contains. A DDSP processor can only produce what its controls and its DSP structure allow.
Maintenance, licensing and upgrade cost
The repository is not archived, and the last push was on 2026-09-22. The most recent tagged release listed is v3.5.1 from 2023-04-26, with v3.1.0 in 2022 and v1.9.0 in 2021 before that. So there is a gap between commit activity and release tagging, and the README does not document a release cadence or a deprecation policy.
Upgrade cost is dominated by the pins. Moving to a newer ddsp means moving the whole TensorFlow, NumPy, protobuf and SciPy set together, because those bounds are declared in setup.py rather than left open. The protobuf pin is annotated as a temporary fix, which tells you the maintainers were working around an upstream resolution problem rather than choosing that version deliberately. There is no documented migration guide between v1.9.0, v3.1.0 and v3.5.1 in the README or the release list.
The licence is Apache-2.0, stated in the README badge area and in setup.py, and the source files carry the standard Apache header. Apache-2.0 is permissive and includes an explicit patent grant. That is a summary of what the repository states, not legal advice; if you are shipping a product, have your own counsel read the licence text and check the licences of the dependencies, several of which are separate projects with their own terms.
Editorial conclusion
Adopt DDSP if you are building a model whose output layer is audio and you want interpretable DSP in the middle: the Processor API, the Harmonic synth and the loss functions are the parts worth reading first. Do not adopt it if what you actually want is the DDSP-VST plugin, which this repository does not contain, or if you need a library that tracks current TensorFlow, since setup.py pins tensorflow<=2.11, numpy<1.24 and protobuf<=3.20. Before committing, run the pip install on your target machine, check that libsndfile-dev is available there, and confirm your TensorFlow version resolves against those pins. The README documents no rollback path and no upgrade procedure, so treat the pinned dependency set as the contract you are accepting.
Frequently asked questions
What is Differentiable Digital Signal Processing (DDSP)?
DDSP is a library of differentiable versions of common DSP functions such as synthesizers, waveshapers and filters, so those interpretable elements can be used as part of a deep learning model, especially as output layers for audio generation.
What is DDSP in the magenta/ddsp repository?
It is a Python library consisting of a core DSP library under ddsp/ and a self-contained training library under ddsp/training/, licensed Apache-2.0 and installed with pip.
How do I install DDSP?
The README gives three steps: install libsndfile-dev with apt, upgrade pip, then run pip install --upgrade ddsp. TensorFlow version 2.1.0 or higher is required.
Is DDSP the same as the DDSP-VST plugin?
No. The README, module list and setup.py describe a Python library of differentiable DSP functions and a training library. Nothing in the repository describes a VST plugin or a downloadable instrument binary.
What can I use DDSP processors for?
The README's example uses a Harmonic synthesizer to turn network outputs for amplitudes, harmonic distribution and f0_hz into audio. Processors expose get_controls() to map inputs to physically meaningful controls and get_signal() to map controls to a signal.
Where do I find DDSP tutorials and demos?
The README links step-by-step Colab tutorials under ddsp/colab/tutorials covering the Processor class, synths and effects, ProcessorGroup, training on a single sound and core functions, plus demos for timbre transfer, training an autoencoder and pitch detection.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/magenta-ddsp)