# RVC WebUI: three requirement files, two CUDA versions, one misspelling

> Retrieval-based-Voice-Conversion-WebUI trains and runs a voice conversion model behind a local Gradio page, and the interesting parts are all in the setup: three hardware-specific requirement files, a two-stage CUDA install pinned to torch 2.7.1, a model tree that must be downloaded in a fixed layout, and a pip index pointing at two Chinese university mirrors by default.

**RVC-Project/Retrieval-based-Voice-Conversion-WebUI** — Easily train a good VC model with voice data <= 10 mins!

- Repository: https://github.com/RVC-Project/Retrieval-based-Voice-Conversion-WebUI
- Stars: 38,552 · Forks: 5,289
- Language: Python
- License: MIT
- Published: 2026-08-17 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/rvc-project-retrieval-based-voice-conversion-webui

## Three requirement files, and the name is requirments not requirements

The install path forks on your GPU before it forks on anything else, and the three branches are three text files in the repository root: requirments_cpu_py312.txt, requirments_cu118_py312.txt and requirments_cu128_py312.txt. The spelling is part of the interface. The name is missing an e, so a correctly spelled requirements file name returns nothing.

On a CPU, an AMD card or an Intel card you use the CPU file, and the README notes that Windows can additionally use DirectML while Linux stays on CPU:

```bash
python -m pip install -r requirments_cpu_py312.txt
```

Nvidia cards split by generation. An RTX 50 series card installs CUDA 12.8 Torch first and then requirments_cu128_py312.txt. Anything older installs CUDA 11.8 Torch first and then requirments_cu118_py312.txt. Both are two-stage, and the order is explicit:

```bash
python -m pip install torch==2.7.1+cu128 torchaudio==2.7.1+cu128 \
  --index-url https://download.pytorch.org/whl/cu128 \
  --extra-index-url https://pypi.org/simple
python -m pip install -r requirments_cu128_py312.txt
```

Torch is pinned to 2.7.1 on both CUDA variants, and the README has you confirm the result rather than assume it:

```bash
python -c "import torch; print('torch:', torch.__version__); print('cuda:', torch.version.cuda); print('cuda available:', torch.cuda.is_available())"
```

Install the wrong file and you are reinstalling torch, which on a fresh machine is the longest step in the whole process.

## The default package index is a mirror at two Chinese universities

The download sources are baked into the top of all three requirement files, and they are not PyPI. The defaults resolve to https://mirrors.pku.edu.cn/pypi/simple for packages and to https://mirrors.nju.edu.cn/pytorch/whl/cpu, /cu118 and /cu128 for the Torch wheels. The README's guidance is written for users in mainland China, who are told they can keep the default mirror as is.

Everyone else has to edit the file, because the documented way to change sources is not an environment variable or a flag. The instruction is to replace only the --index-url and --extra-index-url values, and to keep the package versions, the CUDA suffix and the two-stage order. The official equivalents are given in a table: https://pypi.org/simple for packages, and https://download.pytorch.org/whl/cpu, /whl/cu118 and /cu128 for the wheels. Note that the two-stage torch command above already shows the official form, which is why the second install in that pair does not go through the mirror even before you have edited anything.

The consequence of getting this wrong is not a wrong version, it is a slow or unreachable install. A requirements file pointing at a university mirror from outside that region either crawls or times out during the resolution step, and the error message will point at the network rather than at the file you needed to edit.

## The assets tree is fixed, and hubert_base arrives through an include pattern

The WebUI creates its run directories itself, but the model files do not come with it. They are pulled from a Hugging Face repository, lj1995/VoiceConversionWebUI, and the paths are not negotiable, because the code reads them by name. The tree the README spells out includes assets/hubert_base with config.json, preprocessor_config.json and pytorch_model.bin, assets/rmvpe/rmvpe.pt, assets/pretrained/ and assets/pretrained_v2/ for the base models, assets/pymss_weights/ for the source separation models, assets/weights/ for your own RVC .pth files, assets/indices/ for your .index files, and logs/mute/ for the training silence samples.

The first download step upgrades the client and fetches the feature extractor:

```bash
python -m pip install --upgrade huggingface_hub

# Required for inference and feature extraction
hf download lj1995/VoiceConversionWebUI --revision main \
  --include "hubert_base/*" --local-dir assets
hf download lj1995/VoiceConversionWebUI rmvpe.pt --revision main \
  --local-dir assets/rmvpe
```

The pretrained and pretrained_v2 directories are needed only for v1 and v2 training, and a mute.zip is required for training as well, since the silence samples land in logs/mute/. One extra download exists for a single configuration: if you are on Windows with AMD or Intel DirectML, you also need rmvpe.onnx placed in assets/rmvpe. The CUDA paths do not need it, which is the one place where the DirectML route costs you an extra file.

## 170ms realtime, and 90ms only where ASIO and the driver cooperate

There are two front doors, and the repository root carries a batch file for each on Windows: go-webui.bat for the training and inference interface, and go-realtime_gui.bat for the realtime voice changer. Behind them sit webui.py and realtime_gui.py, and there is a third surface, the RVCRealtimeVST directory at the root, which is the plugin side of the realtime path.

The realtime interface is described as a free choice of operation, with an end-to-end latency of 170ms implemented. The same line claims 90ms end to end when ASIO input and output devices are used, and immediately qualifies it as depending heavily on hardware driver support. Read that carefully before you plan around it. 170ms is what the interface delivers as its baseline; 90ms is a number that depends on the audio stack underneath you, and the README does not name which interfaces qualify or what happens when the driver does not cooperate. There is also an online demo linked in the header, on ModelScope, which is a way to hear the result before installing anything.

The pitch extraction underneath is RMVPE, described as the InterSpeech2023 vocal pitch extraction algorithm, credited in the reference list to Dream-High/RMVPE with the pretrained model trained and tested by yxlllc and RVC-Boss. The stated reason for using it is to eliminate the muted or garbled output problem, with speed and low resource use as the secondary claims.

## top1 retrieval is the trick that keeps your source voice out of the model

The retrieval in the name refers to one specific step. Instead of training on the features of your own recording alone, the project uses top1 retrieval to replace the features of the input source with features drawn from the training set, and the stated purpose is to eliminate voice timbre leakage. That is the mechanism behind the claim that you can train a usable model from a small amount of data.

The data guidance goes with it: collect at least 10 minutes of low-noise speech, and the README frames the whole project as being able to train a good model in ten minutes or less. Training on a modest GPU is called out as a feature, and a separate item covers changing timbre through model merging, done with ckpt-merge in the ckpt tab of the interface.

Two other capabilities sit alongside. The project can call pymss and MSST models to separate vocals from accompaniment quickly, which is how you get clean training material, and audio-slicer appears in the reference list alongside ContentVec, VITS, HIFIGAN and Gradio. What the README does not spell out is the directory layout for training audio itself. It documents the download paths, the outputs you put in assets/weights/ and assets/indices/, and the silence samples under logs/mute/, and leaves the input side to the interface, which creates its run directories.

## Linux gets a documented apt line, Windows gets two files in the project root

The two platforms are set up asymmetrically and the difference is not cosmetic. This branch targets Python 3.12 x64, the README says to enter the repository root first, and Ubuntu 24.04 x86_64 is the recommended distribution. The system dependencies come in one apt line that includes ffmpeg, so the audio tooling is handled for you:

```bash
sudo apt update
sudo apt install -y python3.12 python3.12-venv python3.12-dev ffmpeg unzip libsndfile1 libportaudio2

python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip setuptools wheel
```

Windows gets the venv and nothing else from the system, because there are no system packages to install there:

```powershell
py -3.12 -m venv .venv
.venv\Scripts\activate
python -m pip install --upgrade pip setuptools wheel
```

FFmpeg is the gap that creates. Ubuntu already has it from the apt line; Windows users are told to place ffmpeg.exe and ffprobe.exe in the project root, downloaded from the Hugging Face repository. No path configuration, no PATH edit, just two executables sitting next to webui.py. libportaudio2 in the apt line is the audio I/O layer, which is also the layer the ASIO latency claim depends on.

Once the models are in place, starting the interface is one command, python webui.py, and a headless Ubuntu server uses python webui.py --noautoopen to suppress the browser launch. The service listens on port 7865 by default.

## MIT, plus a file that also covers every library the model borrows

The repository is MIT licensed, and the root carries a LICENSE entry alongside a second entry whose name covers the MIT licence together with the licences of the referenced libraries. That pairing is the part to read rather than skim, because the model is not self-contained. ContentVec, VITS, HIFIGAN, pymss, audio-slicer, Gradio and FFmpeg all appear in the reference list, and RMVPE carries its own attribution for the pretrained checkpoint.

The base model is described as trained on close to 50 hours of the open source VCTK dataset, with the note that there are no copyright concerns to worry about. A third model is promised in the header: an RVCv3 base model with larger parameters, more data, better results, roughly the same inference speed, and less training data required. That is a stated intention, not something in the tree today.

The documentation is spread across languages rather than duplicated in one place. The root README is the Simplified Chinese one, and docs/cn, docs/en, docs/jp, docs/kr, docs/fr, docs/tr and docs/pt hold the rest, with Korean in both Hangul and Hanja variants. An i18n/ directory at the root backs that. Practical upshot for an English reader: the English file is a translation, and the changelog link in the header points at the Chinese one.

## A 2024 release, a two year silence, then a 2026 rebuild on Python 3.12

The version numbers are date stamps, which makes the release history unusually easy to read. 2.1.230814 and 2.2.231006 shipped in 2023, both recorded on 2024-06-05, and 2.3.260718 shipped on 2026-07-21. Between the middle of 2024 and the middle of 2026 there is nothing on the release list. The repository itself is not archived, and the last push was on 2026-08-04, so the current line is a revival rather than a continuous one.

What changed in that revival is visible in the dependency layout. This branch is built for Python 3.12 x64 only, with a CPU file, a cu118 file and a cu128 file, where the newest file exists because RTX 50 series cards need CUDA 12.8. A project that spent two years on a 2023 stack and came back on torch 2.7.1 with a new CUDA target has rewritten its install path, and the two-stage torch command is what you would expect from that rewrite. Anything written against the older pip-only instructions has to be redone rather than adjusted.

Two things carry forward unchanged. The .gitmodules file at the root means parts of the tree arrive as submodules rather than as committed content, which is consistent with assets/ holding model files you are told to download, and the Windows batch launchers are still the documented one-click route on that platform.

## Conclusion

Adopt RVC if you want a voice conversion model trained on your own data, you have a Python 3.12 environment ready, and you can accept a CUDA version chosen by your GPU generation rather than by you. Do not adopt it expecting the documented 90ms realtime figure, since that one is conditional on ASIO and driver support, and do not plan around the default package sources, which resolve to Peking University and Nanjing University mirrors. Before you install anything, check which of the three requirments_ files matches your card, and if you are outside mainland China, rewrite the index URLs in that file first, keeping the versions, the CUDA suffix and the two-stage order intact.

## FAQ

### What is retrieval-based voice conversion?

In this project the retrieval step uses top1 retrieval to replace the features of the input source with features from the training set, with the stated purpose of eliminating voice timbre leakage. That substitution is what lets the model train well from a small amount of recorded data.

### Is RVC AI free?

The repository is MIT licensed, and the root carries both a LICENSE entry and a second entry covering the MIT licence together with the licences of the libraries it references. The interface runs locally, started with python webui.py on port 7865, and the README names no paid service or subscription.

### How do I make an RVC model?

Collect at least 10 minutes of low-noise speech, which is the recommended minimum, and the WebUI creates its own run directories. The README does not document the training audio directory layout, so that part is left to the interface, while the model downloads and output paths are spelled out exactly.

### Does RVC have a realtime voice changer?

Yes. A separate realtime interface is launched with go-realtime_gui.bat on Windows, backed by realtime_gui.py and the RVCRealtimeVST directory, and the README reports 170ms end-to-end latency. It claims 90ms with ASIO input and output devices, and says that figure depends heavily on hardware driver support.

## Sources

- [Official README](https://github.com/RVC-Project/Retrieval-based-Voice-Conversion-WebUI#readme)
- [Project repository](https://github.com/RVC-Project/Retrieval-based-Voice-Conversion-WebUI)
- [Release notes](https://github.com/RVC-Project/Retrieval-based-Voice-Conversion-WebUI/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/rvc-project-retrieval-based-voice-conversion-webui
