# Diff-SVC: singing voice conversion with a diffusion model, and why its own README calls it finished

> Diff-SVC converts a singing recording into a target timbre using a diffusion model, with basic pitch correction. The repository is public and not archived, but its README states the project is no longer actively maintained and points to successors.

**prophesier/diff-svc** — Singing Voice Conversion via diffusion model

- Repository: https://github.com/prophesier/diff-svc
- Stars: 2,719 · Forks: 810
- Language: Jupyter Notebook
- License: AGPL-3.0
- Published: 2026-09-28 · Updated: 2026-09-28 · Language: en
- Canonical page: https://hysenlabs.com/projects/prophesier-diff-svc

## What Diff-SVC converts, and who the repository is written for

Diff-SVC takes a singing recording and re-renders it in a target timbre, keeping the performance. The README describes it as converting input singing voices into a target timbre with basic pitch correction, and the repository name for the task is SVC, singing voice conversion. The input is audio, not sheet music and not text: you supply a vocal you already have, and the model changes who it sounds like.

The audience is narrow and the README is explicit about it. The notes state that the project was established for academic exchange purposes and is not intended for production environments, and that the author is not responsible for copyright issues arising from sound produced by the model. Redistribution of the code, or public publication of results (the README gives video site submissions as an example), requires attribution to the original author and this repository. Using it for any other project requires contacting the author in advance. That is a research artifact with conditions attached, not a product.

The repository also carries a status note that matters more than any feature list: it says the project is no longer actively maintained, that it was an early exploration of applying diffusion models to singing voice conversion, and that community projects such as so-vits-svc and ddps-svc have succeeded it. The last push to the default branch was on 2026-06-06, but the README's own statement about maintenance is the honest starting point.

## The mechanism: HuBERT content, F0, mel spectrograms, a diffusion decoder

The pipeline visible in the repository is a three-stage one. Preprocessing turns raw audio into training data; training fits a diffusion model over that data; inference samples from the model to produce converted audio.

Content comes from a self-supervised speech encoder. The README credits soft-vc, and the update log records moving HuBERT inference from ONNX to torch on 2022-10-28, with a warning that anyone who downloaded the ONNX HuBERT model must download and replace it with the pt model while leaving the config unchanged. That change is what the README says enables direct GPU inference and preprocessing on a 6 GB GTX 1060.

Pitch is handled separately. The dependency list includes praat-parselmouth 0.4.1 and torchcrepe 0.0.17, and an update on 2022-11-02 mentions integrating new vocoder code and updating the parselmouth algorithm. The update log also records an F0 disk cache added on 2022-11-13, which is a practical concession: extracting pitch is slow enough that caching it across runs is worth the disk.

The model itself is a diffusion model over mel spectrograms, with a vocoder converting mel back to waveform. The README mentions a mel spectrogram saving feature added on 2022-11-04, and contentvec support added on 2022-11-11. Training is orchestrated by run.py with a YAML config, and the repository carries a modules/, network/, infer_tools/ and utils/ layout typical of a research codebase. None of this is a stable API, and nothing in the README promises one.

## Installing Diff-SVC and running a first inference

There is no pip package and no install script. The README's inference section points at a notebook: "查看./inference.ipynb". You clone the repository, install the pinned dependencies, point the config at your audio and your checkpoint, and run the notebook or the inference entry point.

The dependency file is fully pinned, which is good for reproducibility and unforgiving on modern hardware. The torch and torchaudio entries carry a CUDA 11.3 suffix, and pytorch-lightning is pinned to 1.3.3:

```bash
pip install -r requirements.txt
```

A shorter list exists as requirements_short.txt if you want to see what the project considers essential before committing to the full set. Preprocessing is a separate step and needs the repository root on the Python path, which is why the README exports PYTHONPATH before invoking the binarizer:

```bash
export PYTHONPATH=.
CUDA_VISIBLE_DEVICES=0 python preprocessing/binarize.py --config training/config.yaml
```

The config path is training/config.yaml in both documented commands. Training uses run.py with an experiment name, and the README's example passes --reset, which discards the existing run for that name rather than resuming it:

```bash
CUDA_VISIBLE_DEVICES=0 python run.py --config training/config.yaml --exp_name [your project name] --reset
```

For inference specifically, the README gives no command line, only the notebook reference. The repository does contain infer.py and flask_api.py at the top level, but the README does not document either, so treat them as code to read rather than interfaces to rely on. The README also says checkpoints, demo audio and other files needed for inference and training are distributed through a QQ channel (channel number 5763z98e4m) and a Discord server, not through the repository or a release. The two releases, v0.1.0-alpha and v0.1.1-alpha, date from November 2022.

## Where Diff-SVC breaks down

The most serious limitation is stated by the project itself: it is no longer actively maintained, and the README redirects users to so-vits-svc and ddps-svc. Any bug you hit, including one in the preprocessing path, is yours to fix.

Hardware is the second constraint. The README's claim of direct GPU inference and preprocessing on a 6 GB GTX 1060 comes with the ONNX-to-pt migration, which means an older download is not just stale, it is wrong. The update note from 2022-11-23 describes a bug that could resample the original ground-truth audio used for inference to 22.05 kHz, and asks users to check their test audio and use the updated code. A repository whose own changelog says "请务必检查自己的测试音频" is telling you that silent quality regressions happened and were only found after release.

The environment is the third constraint. The pinned stack (torch 1.12.1+cu113, torchaudio 0.12.1+cu113, pytorch-lightning 1.3.3, numpy 1.23.4, pywin32 304) reflects a 2022 Windows and CUDA setup. Reproducing it in 2026 means either matching an old CUDA runtime or unpinning and debugging whatever breaks. The requirements.txt also contains a bare "Werkzeug=" line in the truncated listing, which is the kind of detail that suggests the file was edited by hand.

Finally, the README's own copyright position means this is the wrong tool if you need a clear commercial path. The notes disclaim responsibility for copyright issues in generated sound and require prior contact for other uses.

## Diff-SVC against so-vits-svc and RVC

The README names the alternatives directly: so-vits-svc and ddps-svc, both described as community-driven successors. The difference is architectural, not just chronological. Diff-SVC generates mel spectrograms through a diffusion process conditioned on HuBERT content features and separately extracted F0, then vocodes them. so-vits-svc, as its name indicates, is built on the VITS end-to-end architecture, which couples the acoustic model and the vocoder in a single generative framework rather than sampling a mel spectrogram and handing it to a separate vocoder. That changes both the training loop and the inference path.

RVC is the other name that comes up in searches for this project. The README does not describe RVC's internals and does not mention it, so the honest comparison is limited: RVC is a separate project with its own architecture, and Diff-SVC's README points to so-vits-svc and ddps-svc rather than to RVC.

What Diff-SVC offers that its successors do not is legibility as a study object. It is a comparatively small, diffusion-first implementation with a documented preprocessing and training split, and the README credits diffsinger, the openvpi-maintained diffsinger fork, and soft-vc as its foundations. If you want to read how diffusion was wired to singing voice conversion in 2022, this is a compact example. If you want results, the README tells you to go elsewhere.

## Licence, checkpoints and the cost of keeping it running

The repository is AGPL-3.0. That licence carries network-use obligations: if you modify the code and expose it to users over a network, the AGPL's terms apply to that service, not just to distributed binaries. Separately, the README's own notes impose conditions that sit outside the licence text: attribution to the author and this repository when you redistribute the code or publish results, and prior contact with the author for other uses. Read both, and take legal advice if the distinction matters to you; this article does not give it.

Upgrade cost is effectively unbounded in one direction and zero in the other. Because the project is unmaintained, there is no upgrade path: no security patches, no dependency bumps, no compatibility work for newer PyTorch. You are pinned to the 2022 stack unless you do the porting yourself. The one thing you can still get from the community is artifacts rather than code: checkpoints and supporting files are distributed through the QQ channel and Discord listed in the README, and those come with no versioning guarantee relative to the code you cloned.

The repository also has a note that should stop a specific kind of confusion: it states that this project has no connection with the paper of the same name, DiffSVC (arXiv 2105.13871), and asks that the two not be mixed up. If you arrived from a citation, check which one you actually want.

## Conclusion

Diff-SVC is for readers who want to study how a diffusion model was applied to singing voice conversion, or who already hold a checkpoint and want to run inference from inference.ipynb. It is not for anyone who needs a maintained tool, a packaged GUI, or a service to put in front of users: the README itself says the project is no longer actively maintained and names so-vits-svc and ddps-svc as successors, and the notes say it was built for academic exchange rather than production. Before you commit time, verify three things in this order: that the pinned requirements (torch 1.12.1+cu113, torchaudio 0.12.1+cu113, pytorch-lightning 1.3.3) can be resolved on your machine, that a checkpoint and its matching config are available to you, and that your intended use is permitted under AGPL-3.0 and the author's reuse conditions.

## FAQ

### What is singing voice conversion (SVC)?

It is the task Diff-SVC performs: taking an input singing voice and re-rendering it in a target timbre while keeping the performance, with basic pitch correction. The README frames the project as converting input singing voices into a target timbre.

### Is Diff-SVC still maintained?

No. The README's project status note states that the project is no longer actively maintained, that it was an early exploration of diffusion models for singing voice conversion, and that community projects such as so-vits-svc and ddps-svc have succeeded it.

### How do I install and train Diff-SVC?

Clone the repository, install the pinned dependencies with pip install -r requirements.txt, then run preprocessing with export PYTHONPATH=. followed by python preprocessing/binarize.py --config training/config.yaml, and train with python run.py --config training/config.yaml --exp_name [your project name] --reset. The README points to ./inference.ipynb for inference.

### Where do I download Diff-SVC checkpoints?

The README says trained checkpoints, demo audio and other files needed for inference and training are distributed through a QQ channel (channel number 5763z98e4m) and a Discord server, not through the repository or its releases. The two releases listed, v0.1.0-alpha and v0.1.1-alpha, date from November 2022.

### Is Diff-SVC the same as the DiffSVC paper?

No. The README explicitly states that this project has no connection with the paper of the same name, DiffSVC (arXiv 2105.13871), and asks that the two not be confused.

## Sources

- [Issues](https://github.com/prophesier/diff-svc/issues)
- [License: AGPL-3.0](https://github.com/prophesier/diff-svc/blob/main/LICENSE)
- [prophesier/diff-svc on GitHub](https://github.com/prophesier/diff-svc)
- [README](https://github.com/prophesier/diff-svc/blob/main/README.md)
- [Releases](https://github.com/prophesier/diff-svc/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/prophesier-diff-svc
