Vocos: a Fourier-domain vocoder that skips the time domain entirely
Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
At a glance
- What is it?
- Most GAN vocoders generate waveform samples directly. Vocos generates spectral coefficients and inverts them, which is why it reconstructs audio in a single forward pass and why its dependency list is short.
- Who is it for?
- Vocos is small in a way that matters. Two pretrained checkpoints under ten million parameters each, eight runtime dependencies, and a single forward pass from acoustic features to waveform.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 38 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 20, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The one architectural sentence that defines the project
The README's opening paragraph contains the whole design in four sentences. Vocos is a fast neural vocoder designed to synthesize audio waveforms from acoustic features. It is trained with a generative adversarial network objective and can generate waveforms in a single forward pass. Unlike other typical GAN-based vocoders, it does not model audio samples in the time domain. Instead it generates spectral coefficients and reconstructs audio through an inverse Fourier transform.
That last sentence is the project. Generating spectral coefficients rather than raw samples means the network's output lives in a compact frequency representation, and getting back to audio is a transform rather than a learned generation step. The single forward pass claim follows from that: there is no autoregressive rollout over thousands of samples.
For anyone who has read a vocoder paper, this is the interesting contrast. Time-domain GAN vocoders like the HiFi-GAN family spend their capacity on waveform detail. A Fourier-domain model spends it on magnitude and phase. Which wins depends on your content and your data, but the two approaches fail differently, and a vocoder that picks one deliberately is easier to reason about than an ensemble.
The paper is arXiv 2306.00814, authored by Hubert Siuzdak, and audio samples are hosted on the project's GitHub Pages site.
Installing for inference versus installing to train
The README splits installation into two paths, and the difference matters more than usual here.
For inference only, it is a single command:
python -m pip install vocosFor training, the entry point, configs and metrics live in the source repository, so you clone and install with the training extras:
git clone https://github.com/gemelo-ai/vocos.git
cd vocos
python -m pip install "setuptools<80"
python -m pip install -e ".[train]"That `setuptools<80` pin is not cargo cult. The README explains it: the checked-in training stack pins PyTorch Lightning 1.8.6, which still uses `pkg_resources`, and holding Setuptools below version 80 in the training environment preserves that compatibility.
So there are two support postures in one project. Inference is a normal pip install against a maintained package. Training is an archaeology exercise on a pinned 2022-era framework. Anyone planning to retrain Vocos on a custom corpus should read that before budgeting the work.
The runtime dependency set is correspondingly tight:
torch
torchaudio
numpy
scipy
einops
pyyaml
huggingface_hub
encodec==0.1.1Eight entries, with PyTorch and torchaudio doing the heavy lifting, `einops` for the tensor reshaping a frequency-domain architecture needs, `huggingface_hub` for checkpoint loading, and EnCodec pinned to an exact version because the discrete codec path depends on its exact behavior.
Two ways in: mel spectrograms or EnCodec tokens
The package exposes two front doors, and choosing between them is the first real decision you make.
The mel path reconstructs audio from a mel spectrogram. The README's example is deliberately minimal:
import torch
from vocos import Vocos
vocos = Vocos.from_pretrained("charactr/vocos-mel-24khz")
mel = torch.randn(1, 100, 256) # B, C, T
audio = vocos.decode(mel)The comment on the random tensor is the useful part: batch, channels, time. A 100 by 256 mel goes in, a waveform comes out.
The EnCodec path takes discrete tokens instead. It requires a `bandwidth_id` drawn from the list `[1.5, 3.0, 6.0, 12.0]`, corresponding to EnCodec's bandwidth embeddings:
vocos = Vocos.from_pretrained("charactr/vocos-encodec-24khz")
audio_tokens = torch.randint(low=0, high=1024, size=(8, 200)) # 8 codebooks, 200 frames
features = vocos.codes_to_features(audio_tokens)
bandwidth_id = torch.tensor([2]) # 6 kbps
audio = vocos.decode(features, bandwidth_id=bandwidth_id)The two models are not interchangeable. `vocos-mel-24khz` was trained on LibriTTS for 1M iterations with 13.5M parameters. `vocos-encodec-24khz` was trained on the DNS Challenge dataset for 2M iterations with 7.9M parameters. The EnCodec model is less than two-thirds the size, which tracks with decoding from a lower-dimensional discrete representation.
One documentation wart worth knowing about: the copy-synthesis example under the EnCodec heading calls `vocos(y, bandwidth_id=bandwidth_id)` where the mel section calls `vocos.decode(mel)`. The mel form is the one that matches the API described everywhere else, so treat the EnCodec snippet as illustrative rather than copy-paste exact.
Two checkpoints, and what the releases changed
There are three tagged releases, and they cluster in a four month window ending in October 2023.
v0.0.3, published 2023-06-21, fixed inefficient phase unwrapping. That is a small change with real consequences for a Fourier-domain model: phase unwrapping that is correct but slow undermines the entire argument for the architecture, so fixing it matters more than the release notes suggest.
v0.0.4 followed the next day, fixing type errors. v0.1.0 on 2023-10-14 is the substantive one. It adopted a new multi-resolution, multi-band discriminator taken from the Descript Audio Codec, updated the recommended AdamW hyperparameters to `lr=5e-4` and `betas=(0.8, 0.9)`, and refreshed the pretrained models on Hugging Face.
Those earlier checkpoints are tagged for reference, and the release notes show the shape of the escape hatch: loading `charactr/vocos-encodec-24khz` with a `revision` of `v0.0.4` reaches the pre-discriminator version directly, rather than requiring a git commit pin. Reproducing a paper result or a production output from before that change is therefore still possible.
The repository has 1160 stars and 134 forks, 48 open issues, is not archived, and was last pushed on 2026-08-29. A detail that dates the project: the release notes and pull request links still reference the older `charactr-platform/vocos` repository, so the code moved to the `gemelo-ai` organization after those fixes landed.
Training means fighting a pinned framework
The training section is short and honest about its own age. You prepare a filelist for training and validation sets, fill in a config such as `configs/vocos.yaml`, and start training with `train.py`.
The filelist step uses `find` with a predicate that handles both regular files and symlinks, which is the detail that matters for anyone assembling a dataset out of symlinked directories:
find "$TRAIN_DATASET_DIR" \( -type f -o -type l \) -name '*.wav' -print > filelist.train
find "$VAL_DATASET_DIR" \( -type f -o -type l \) -name '*.wav' -print > filelist.valThe type `l` in that predicate is what includes symlinked `.wav` files without descending into symlinked directories.
Then it is one command, and a caveat:
python train.py -c configs/vocos.yamlThe README states that the checked-in training code and configs target PyTorch Lightning 1.8.6, and points at that specific version's documentation for customizing the pipeline. Combined with the `setuptools<80` requirement, that pins three things at once: the framework version, the packaging tool, and therefore what your environment can be.
The repository tree also has a `metrics/` directory, so evaluation is part of the repo rather than something you write yourself, and a `notebooks/` directory containing an integration example with the Bark text-to-audio model, which is the clearest signal of intended use: Vocos is meant to be the last stage in someone else's generation pipeline.
Maintenance is stated rather than implied
The README has a dedicated project maintenance section, which is unusual and worth reading before you depend on this.
It says Vocos remains available for research and production use, and that the maintainers are continuing to review focused, backward-compatible improvements. It names current priorities: documentation, project infrastructure, issue and pull request triage, and making the contribution path clearer. It closes by saying feedback and small, well-scoped contributions are welcome.
Read carefully, that is a maintenance-mode announcement with a service guarantee attached. No new features are promised, the compatibility promise is explicit, and the work named is the unglamorous kind that keeps a dependency-light library installable. The commit cadence backs it up: the last push was 2026-08-29, close to the release-freeze period implied by the last tag in October 2023.
For a package whose pre-trained weights are the actual product, that is a reasonable posture. The model files do not rot. The code needs to keep installing against current PyTorch, and that is what the declared priorities are aimed at.
The license is MIT, `setup.py` names Hubert Siuzdak as author, and contributing asks you to read `CONTRIBUTING.md` before opening a pull request and to include enough detail for maintainers to reproduce the change. The project also asks that you cite the arXiv paper if the code contributes to your research.
Editorial conclusion
Vocos is small in a way that matters. Two pretrained checkpoints under ten million parameters each, eight runtime dependencies, and a single forward pass from acoustic features to waveform. The architectural bet is that you do not need to model audio samples in the time domain, and the README states that plainly rather than burying it. The practical caveats are all about training rather than inference. Inference is `pip install vocos` and a call to `Vocos.from_pretrained`, while training means cloning the repository, holding Setuptools below version 80, and accepting PyTorch Lightning 1.8.6 as the pipeline. If you want to retrain on your own data, budget for fighting that pin. If you want to use the models, there is nothing to negotiate, and the maintained status is stated plainly on the README as backward-compatible improvements plus documentation and triage.
Frequently asked questions
What is Vocos and how does it differ from other neural vocoders?
Vocos is a fast neural vocoder that synthesizes waveforms from acoustic features in a single forward pass, trained with a GAN objective. Unlike typical time-domain GAN vocoders, it does not model audio samples in the time domain. It generates spectral coefficients and reconstructs audio through an inverse Fourier transform instead.
How do I install Vocos and load a pretrained model?
For inference, install the package from PyPI with python -m pip install vocos, then create a model with Vocos.from_pretrained using either the charactr/vocos-mel-24khz or charactr/vocos-encodec-24khz checkpoint and call decode on a mel spectrogram or on features converted from EnCodec tokens.
What is the difference between the two pretrained Vocos models?
The mel model, charactr/vocos-mel-24khz, was trained on LibriTTS for 1M iterations with 13.5M parameters and decodes mel spectrograms. The EnCodec model, charactr/vocos-encodec-24khz, was trained on the DNS Challenge dataset for 2M iterations with 7.9M parameters and decodes discrete EnCodec tokens, which requires passing a bandwidth_id from the list [1.5, 3.0, 6.0, 12.0].
Can I still train Vocos on my own audio dataset?
Yes, but expect constraints. The training entry point, configs and metrics live in the source repository, so you clone it and install the training extra. The checked-in training code and configs target PyTorch Lightning 1.8.6, which still uses pkg_resources, so the README asks you to pin Setuptools below version 80 to keep that compatibility.
Is Vocos still maintained?
The README has a project maintenance section stating that Vocos remains available for research and production use, with maintainers reviewing focused, backward-compatible improvements. Current priorities are documentation, project infrastructure, issue and pull request triage, and clarifying the contribution path. The repository is not archived and was last pushed on 2026-08-29.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/gemelo-ai-vocos)