DDSP-SVC: what the 6.3 branch actually ships
Real-time end-to-end singing voice conversion system based on DDSP (Differentiable Digital Signal Processing)
At a glance
- What is it?
- A rectified-flow singing voice conversion project whose default branch is named 6.3 while the newest published release is 5.0. Three pretrained models must be placed by hand before anything runs, the changelog sits in the Chinese readme, and the English text ends partway into a mix example.
- Who is it for?
- DDSP-SVC fits a specific setup: training singing voice conversion on a single consumer GPU, with the vocoder, pitch extractor and feature encoder fetched by hand. Before an overnight run, settle three things.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 2 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The default branch is 6.3 and the newest tag is 5.0
The default branch is named 6.3, yet the newest published tag is 5.0 from 8 February 2024, with 4.0 in August 2023 and 3.0 in May 2023. Nothing in that list records what changed across the 6.x series, so any statement about current behaviour has to come from the branch itself rather than from release notes. The entry points still carry the rectified-flow generation in their names: `train_reflow.py`, `main_reflow.py`, `gui_reflow.py`, `gui_reflow_locale.py` and `configs/reflow.yaml`. The top-level listing holds no diffusion-named module, which leaves the quality path the introduction walks through, from a pre-trained vocoder based enhancer to a shallow diffusion model and then to a rectified-flow model, told in prose while every runnable command stays on the reflow set. The branch was last pushed on 26 September 2026 and is not archived.
Three pretrained pieces and four paths that have to match
Three pretrained pieces are placed by hand before anything runs, and each one has a location the configuration file names. A feature encoder goes under `pretrain/contentvec` for ContentVec or `pretrain/hubert` for HubertSoft. They are alternatives, not a pair, and the HubertSoft route also says the configuration file must be modified at the same time without naming the key to change, so that switch stays implicit. The vocoder goes wherever the `vocoder.ckpt` parameter points, `pretrain/nsf_hifigan/model` by default, and its `config.json` has to sit in that same directory, which means a checkpoint moved on its own loads against nothing. Pitch extraction uses an RMVPE archive unzipped straight into `pretrain/`. Two vocoder options sit on top: the openvpi release archive, or one you fine-tune through the openvpi/SingingVocoders project for better quality.
requirements.txt pins one version and leaves PyTorch out
The dependency file carries seventeen entries, and `numpy==1.26.4` is the only one pinned to an exact version while librosa, scipy, transformers, torchcrepe, torchfcpe and the rest float. PyTorch is deliberately absent from the list, and the install step opens by telling you to take it from the official PyTorch site before running:
pip install -r requirements.txtThat ordering matters, because torchaudio has to match the torch build instead of resolving on its own. One combination is written down as known good: python 3.11 on Windows with cuda 13.0, torch 2.9.1 and torchaudio 2.9.1. The interface layer arrives as `FreeSimpleGUI`, and `sounddevice` and `soundfile` sit beside it, so the GUI and the audio path are both first-class dependencies rather than extras pulled in later.
A sample rate mismatch runs to completion and just takes forever
Training clips go into `data/train/audio` and validation clips into `data/val/audio`, and `python draw.py` exists to help you choose which files land in the validation set. After that:
python preprocess.py -c configs/reflow.yaml -j <number of processes>The shipped configuration targets a 44.1kHz synthesizer on an RTX-4060, and the notes wrapped around that command hold the rough edges. A clip whose rate differs from the yaml still runs to the end, so nothing fails loudly; the resampling pass is simply very slow. About a thousand clips are advised, long clips can be cut into shorter segments to speed training up, none should run under two seconds, and a dataset too large for internal memory asks you to set `cache_all_data` to false. Validation wants roughly ten clips because more makes validation slow. Recordings of lower quality are handled by setting `f0_extractor` to `rmvpe`.
n_spk swaps flat folders for numbered speaker ids
Whether the model carries one voice or several comes down to `n_spk` in the configuration file. Above one, the audio folders have to be named with positive integers no greater than that count, and each number becomes the speaker id that inference later selects with `-id`:
# training dataset
# the 1st speaker
data/train/audio/1/aaa.wav
data/train/audio/1/bbb.wav
...
# the 2nd speaker
data/train/audio/2/ccc.wav
data/train/audio/2/ddd.wav
...
# validation dataset
# the 1st speaker
data/val/audio/1/eee.wav
data/val/audio/1/fff.wav
...
# the 2nd speaker
data/val/audio/2/ggg.wav
data/val/audio/2/hhh.wav
...Set to one, the older flat layout keeps working with clips sitting directly under the audio directory. The validation tree mirrors the numbering, so each voice needs its own held-out clips. No conversion between the two shapes is described, which makes this a decision made before preprocessing rather than something the scripts sort out afterwards.
Two intervals decide which weights are throwaway
A weight is written temporarily every `interval_val` steps and permanently every `interval_force_save` steps, and both can be moved to suit the situation. Training can be interrupted safely, since rerunning the identical command resumes it:
python train_reflow.py -c configs/reflow.yamlFine-tuning takes the same route. Stop the run, re-preprocess a new dataset or change parameters such as batchsize and lr, then run the same command line again. Progress is watched through TensorBoard pointed at the `exp` directory:
tensorboard --logdir=expTest audio samples show up only after the first validation, so a run that has not reached an `interval_val` boundary gives you curves without any audio attached.
The inference flags expose the ODE and the mix example stops mid line
Conversion runs through one script with the whole sampling setup on the command line:
python main_reflow.py -i <input.wav> -m <model_ckpt.pt> -o <output.wav> -k <keychange (semitones)> -id <speaker_id> -step <infer_step> -method <method> -ts <t_start>`-step` is the number of sampling steps taken by the rectified-flow ODE, `-method` is either `euler` or `rk4`, `-ts` is where the ODE begins, and `-k` shifts pitch by semitones. That start time must be greater than or equal to `t_start` in the configuration file, with keeping the two equal recommended and 0.0 as the default. Right after it a `-mix` option for designing your own vocal timbre is offered, and the example beneath it is where the English text ends: the fence holds a single comment reading `# Mix the t` and nothing after that. The flag is announced, its arguments are not. Three more top-level scripts, `batch_infer.py`, `export_onnx.py` and `slicer.py`, sit in the tree with no matching command anywhere in the visible text.
The changelog is one sentence pointing at the Chinese file
The update log is a sentence rather than a list. The author says they are too lazy to translate and sends readers to the Chinese readme, and `cn_README.md` is indeed present in the tree. Since the newest tag is still 5.0, that file is the only place branch-level change detail exists, which leaves English readers tracking a moving branch from a document whose final section never got finished. The disclaimer above the install step is the other line to read before training anything: models are to be trained only on legally obtained authorized data, the synthesized audio is not to be used for illegal purposes, and the author accepts no responsibility for infringement or fraud caused by the checkpoints or the audio they produce.
Editorial conclusion
DDSP-SVC fits a specific setup: training singing voice conversion on a single consumer GPU, with the vocoder, pitch extractor and feature encoder fetched by hand. Before an overnight run, settle three things. Which branch you are actually on, since 6.3 carries no matching release. What your own dataset licence permits, because the project asks for authorized data and every checkpoint inherits the voices inside it. Whether your clips already match the sample rate in `configs/reflow.yaml`, because a mismatch does not fail, it just grinds. The English text stops mid-example, so expect to read the configuration file rather than the prose for the last details.
Frequently asked questions
What does DDSP-SVC need installed before training starts?
PyTorch from the official site first, then the dependency file, which pins only `numpy==1.26.4`. One combination is given as known good: python 3.11 on Windows with cuda 13.0, torch 2.9.1 and torchaudio 2.9.1.
Where do the DDSP-SVC pretrained models have to be placed?
ContentVec or HubertSoft under `pretrain/contentvec` or `pretrain/hubert`, one of the two rather than both, the vocoder checkpoint at the location `vocoder.ckpt` names, which defaults to `pretrain/nsf_hifigan/model`, with that vocoder's `config.json` in the same directory, and the RMVPE extractor unzipped into `pretrain/`.
How is a multi-speaker model trained in DDSP-SVC?
Set `n_spk` above one in the configuration file and name every audio folder with a positive integer no greater than that count, each number acting as the speaker id passed to inference with `-id`. With `n_spk` at 1 the flat single-speaker layout still works.
Is there a published DDSP-SVC release for the current branch?
No. The default branch is named 6.3 while the newest release is 5.0 from 8 February 2024, preceded by 4.0 in August 2023 and 3.0 in May 2023, so the branch carries no tag of its own.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/yxlllc-ddsp-svc)