Open-source project
voicepaw/so-vits-svc-fork avatar
voicepaw/so-vits-svc-fork

so-vits-svc-fork: a realtime voice conversion fork that is no longer maintained

so-vits-svc fork with realtime support, improved interface and more features.

9,332 stars1,220 forksPythonNOASSERTION

At a glance

What is it?
voicepaw/so-vits-svc-fork adds realtime conversion and a GUI to the SoftVC VITS codebase, but its own README declares the project no longer maintained. Here is what it does, how to install it, and when to pick something else.
Who is it for?
Adopt so-vits-svc-fork if you already have compatible 4.0 or 4.1 models and want a pip-installable CLI and GUI for offline or realtime conversion on a machine you control; the README states the project is no longer maintained, so treat it as a frozen tool rather than a platform to build on. Do not adopt it for new production work, for non-4.0/4.1 checkpoints, or if you expect upstream fixes.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 13 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What so-vits-svc-fork actually solves

Voice conversion takes an audio clip of one singer and re-renders it in another voice. The original so-vits-svc project did that offline, with a training and inference pipeline that assumed you would prepare a dataset, train a model, and then convert files in a batch. so-vits-svc-fork keeps that pipeline but adds a realtime path and a unified interface on top.

The README lists what the fork adds over the original repository: realtime voice conversion, a partial integration of QuickVC, a fix for what the author describes as misuse of ContentVec in the original repo, pitch estimation through CREPE, a GUI plus a unified CLI, and training that the README claims is about twice as fast. It also states that pretrained models download automatically and that fairseq no longer needs to be installed separately.

The audience is narrow and technical. You need a GPU or a willingness to run on CPU, Python 3.9 or newer according to pyproject.toml, and a checkpoint in the 4.0 or 4.1 format. The README is explicit that 4.1 models are not supported and that other models are not supported either, which is a sharper boundary than most forks draw.

How the conversion pipeline is put together

The package is a Python project built with setuptools, and pyproject.toml exposes two console entry points: svc maps to so_vits_svc_fork.__main__:cli and svcg maps to so_vits_svc_fork.gui:main. That split is the whole interface story. Anything you can do in the GUI has a CLI counterpart, and the CLI is the part you can script.

Underneath, the dependency list tells you what the pipeline is made of. librosa and soundfile handle audio loading, torchcrepe and praat-parselmouth handle pitch, transformers pulls in the ContentVec and HuBERT-style encoders, lightning drives training, and sounddevice is what makes the realtime path possible because it can open a live audio stream. FastAPI is also a dependency, so there is an HTTP surface somewhere in the package even though the README does not walk through it.

The data flow is the standard SoftVC arrangement: an input waveform is encoded into content features, pitch is estimated separately, a trained VITS-style decoder conditioned on a speaker embedding reconstructs the waveform, and the result is written out or streamed back. The fork's contribution is not a new architecture. It is the plumbing around it: CREPE for pitch instead of the original estimator, automatic download of the pretrained encoder weights, and a realtime loop built on sounddevice.

Installing so-vits-svc-fork with pip and running the CLI

The README offers three install paths: a one-click Windows .bat file, pipx, and plain pip. The pipx route is marked experimental in the README, so the plain pip route inside a virtual environment is the one the project describes in most detail. The README warns that installing without a virtual environment can raise a PermissionError when Python lives under Program Files.

Create and activate a Python 3.11 environment, then install torch and torchaudio from the CUDA 12.1 wheel index before installing the package itself. The README notes that MacOS users or anyone without a GPU should drop the torch line and let the default wheels resolve.

bash
python3.11 -m venv venv
source venv/bin/activate
python -m pip install -U pip setuptools wheel
pip install -U torch torchaudio --index-url https://download.pytorch.org/whl/cu121
pip install -U so-vits-svc-fork

After that, the svc command is on your PATH. Running it with no arguments should print the Typer-generated command list, which is the quickest way to confirm the install resolved. The GUI is a separate entry point.

bash
svc --help
svcg

The Dockerfile takes a different route: it starts from pytorch/pytorch:2.0.1-cuda11.7-cudnn8-runtime, installs build-essential, then pip installs so-vits-svc-fork and sets svcg as the entrypoint. That image pins an older PyTorch than the torch>=2.8.0 floor in pyproject.toml, so the container and the pip metadata are not describing the same environment. If you build the image, expect resolution to pull a newer torch than the base image ships.

The maintenance status is the first thing to read

The README contains a section titled No Longer Maintained, and it is not a hedge. The author gives reasons: the technology moved on within a year, the goal of a more modular and easier-to-install repository was not reached for lack of skills, time and money, PySimpleGUI dropped LGPL, and Typer overtook Click in popularity. The same section notes that updates have been limited to maintenance since Spring 2023.

There is a real tension between that text and the repository's activity. The last push was on 2026-09-10, and the most recent release, v4.2.30, is dated 2026-02-02. So the repository is not archived and commits still land. What you should not conclude from that is that the project is under active development in the sense of new features or a roadmap. The README says otherwise, and the release cadence between v4.2.29 in October 2025 and v4.2.30 in February 2026 is consistent with maintenance rather than feature work.

For an adopter, this changes the calculus. You are picking up a finished tool, not joining a project. Bug reports may sit. Dependency bumps may lag. The pyproject.toml still classifies the package as Development Status 2, Pre-Alpha, which is an odd label for a 4.x release line but is what the metadata says.

Where so-vits-svc-fork breaks or is the wrong choice

The model format is the hardest limit. The README states that the fork is based on branch 4.0 (v1) or 4.1 and that models are compatible, then immediately says 4.1 models are not supported and other models are not supported. That sentence contradicts itself on its face, and the practical reading is that you should test your checkpoint before planning anything around it. If your model came from a different fork or a newer upstream line, assume it will not load.

Realtime conversion is the headline feature and also the most fragile. It depends on sounddevice, which means a working audio backend on the host, and it depends on a GPU fast enough to run the encoder and decoder inside the stream callback budget. The README does not publish latency figures, and it does not document a fallback when the machine cannot keep up. On a laptop CPU, the realtime path is not a realistic target.

There is also no rollback story. The README does not document how to revert a model or a configuration after a bad conversion, and it does not describe version pinning for the pretrained weights that download automatically. If reproducibility matters to you, capture the environment you installed into rather than trusting the package to stay stable.

Finally, the licensing metadata is ambiguous. The README badge and pyproject.toml both say MIT, but the repository's license field is reported as NOASSERTION, meaning GitHub could not classify the LICENSE file. That mismatch is worth resolving before you redistribute anything.

Alternatives the README itself points to

The maintainer does not pretend this is the only option. The alternatives section names several projects and, unusually, is candid about their status. The RVC family is listed first: IAHispano/Applio under MIT, which the README calls actively maintained, fumiama's RVC under AGPL, and the original RVC under MIT, which the README says is no longer maintained.

The difference in approach matters more than the names. RVC-family tools are retrieval-based: they build an index over the target speaker's feature vectors and retrieve from it at inference time, which tends to need less training data and converges faster than the VITS-style decoder this fork uses. so-vits-svc-fork instead trains a generative model per speaker, which is heavier to prepare but gives the decoder more freedom in how it renders pitch and timbre.

The README also lists VCClient, which offers a web-based GUI for realtime conversion and is described as not quite actively maintained, plus fish-diffusion and the DDSP-SVC and ReFlow-VAE-SVC projects from yxlllc, and coqui-ai/TTS for text-to-speech. If your actual goal is a low-latency realtime voice changer, the README's own advice is to look elsewhere, because it says this project may suit people who want to try voice conversion at all rather than people chasing latency.

Who should install this, and what to check first

Pick so-vits-svc-fork if you already hold a 4.0 or 4.1 checkpoint, you want a single pip install rather than a fairseq build, and you value having both svc and svcg available without writing glue code. The automatic pretrained model download removes a step that trips people up in the original repository, and the CREPE pitch estimator is a genuine change rather than a cosmetic one.

Skip it if you need a maintained dependency, if your models are in another format, or if realtime latency is the requirement rather than a bonus. The README's own alternatives section is the strongest argument against choosing this fork for new work.

Before you invest time, do three checks in order. Confirm your checkpoint loads through the svc CLI rather than assuming the 4.0/4.1 compatibility note covers it. Confirm torch and torchaudio resolved against the CUDA wheel index that matches your driver, since the README's example pins cu121 while the dependency floor in pyproject.toml is torch>=2.8.0. And confirm the GUI starts from the svcg entry point on your platform, because the maintainer cites the PySimpleGUI licence change as one reason for stepping back, and the project now depends on pysimplegui-4-foss.

Editorial conclusion

Adopt so-vits-svc-fork if you already have compatible 4.0 or 4.1 models and want a pip-installable CLI and GUI for offline or realtime conversion on a machine you control; the README states the project is no longer maintained, so treat it as a frozen tool rather than a platform to build on. Do not adopt it for new production work, for non-4.0/4.1 checkpoints, or if you expect upstream fixes. Before committing, verify that your model loads with the svc CLI, that torch and torchaudio install against the CUDA wheel index your GPU needs, and that the PySimpleGUI-based GUI starts on your platform.

Frequently asked questions

Can I clone my voice locally with so-vits-svc-fork?

Yes, the pipeline runs on your own machine rather than a hosted service: the README describes a pip install, automatic download of pretrained models, and a realtime path built on sounddevice. You still need a compatible 4.0 or 4.1 checkpoint, and the README states that other models are not supported.

How do I install so-vits-svc-fork with pip?

The README's manual route creates a Python 3.11 virtual environment, installs torch and torchaudio from the CUDA 12.1 wheel index, then runs pip install -U so-vits-svc-fork. It warns that installing without a virtual environment can raise a PermissionError when Python is under Program Files.

Is so-vits-svc-fork still maintained?

The README has a No Longer Maintained section and says updates have been limited to maintenance since Spring 2023. The repository is not archived and the last push was on 2026-09-10, so commits still land, but the project describes itself as finished rather than actively developed.

Which models does so-vits-svc-fork support?

The README says the fork is based on branch 4.0 (v1) or 4.1 and that models are compatible, then states that 4.1 models are not supported and other models are not supported. Test your checkpoint with the svc CLI before relying on it.

Official sources

  1. Issues
  2. README
  3. Releases
  4. voicepaw/so-vits-svc-fork on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/voicepaw-so-vits-svc-fork.svg)](https://hysenlabs.com/projects/voicepaw-so-vits-svc-fork)