GPT-SoVITS: Few-Shot Voice Cloning From a Minute of Audio
GPT-SoVITS is a few-shot voice cloning and TTS WebUI that fine-tunes a speaker's voice from one minute of audio and supports zero-shot synthesis from a five-second sample.
At a glance
- What is it?
- GPT-SoVITS is a Python WebUI for zero-shot and few-shot text-to-speech and voice conversion. The mechanism is a GPT stage plus a SoVITS stage, and the real cost is data preparation, not training.
- Who is it for?
- Adopt GPT-SoVITS if you need a self-hosted, MIT-licensed TTS stack for English, Japanese, Korean, Cantonese or Chinese and you are prepared to curate clean audio yourself. Do not adopt it if you want a managed API with consent controls, or if you expect GPU-quality training on a Mac, since the README states Mac models are trained on CPU.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 42 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
What GPT-SoVITS solves, and who it is actually for
The bottleneck in custom text-to-speech has never been the model architecture. It is the recording session. Most voice cloning pipelines ask for hours of clean, transcribed speech before they produce anything usable, which rules out archival recordings, a single podcast episode, or a voice actor you can only book once.
GPT-SoVITS targets that constraint directly. The README describes two modes: zero-shot TTS, where a 5-second vocal sample is enough for instant conversion, and few-shot TTS, where roughly 1 minute of training data is fine-tuned into a closer match. The project is a WebUI, not a library-first toolkit, and it bundles the unglamorous parts of the workflow: voice accompaniment separation to strip background music, automatic training set segmentation, and multilingual ASR through Fun-ASR-Nano, SenseVoice and FunASR, plus text labeling.
That bundling tells you the intended user. This is for someone who has a folder of audio and no pipeline. It is not aimed at teams who already have a data engineering stack for speech, and it is not aimed at people who want to type a sentence and receive a voice without touching a single file. Cross-lingual inference covers English, Japanese, Korean, Cantonese and Chinese, so the training language and the inference language do not have to match.
The GPT plus SoVITS split, and where your GPU memory goes
The name is literal. The system is two models in sequence, and the repository layout reflects that: GPT_SoVITS/ holds the model code, tools/ holds the data preparation utilities, and webui.py plus api.py and api_v2.py sit at the top level as entry points.
The first stage is autoregressive and handles the mapping from text to semantic tokens. The second, SoVITS, converts those tokens into an acoustic representation and then into waveform. Splitting the job this way is why the project can offer zero-shot and few-shot from the same checkpoint family: the GPT stage carries the linguistic generalization, while the SoVITS stage carries the timbre, which is what fine-tuning actually adjusts.
The practical consequence is that inference speed is not uniform across hardware. The README reports an RTF of 0.028 on a 4060Ti and 0.014 on a 4090 for v2 ProPlus, with 0.526 on an M4 CPU. Those numbers come from the project, not from independent measurement, and the README is pointed about people criticizing inference speed. Treat them as a ratio between devices rather than a guarantee: a 4090 is roughly twice as fast as a 4060Ti in that comparison, and Apple silicon running on CPU is an order of magnitude slower. If you plan to serve many concurrent requests, the GPU tier determines whether that is viable at all.
The requirements file pins the environment tightly. numpy is capped below 2.0, librosa is pinned to 0.10.2, gradio is held below 5, transformers must be at least 4.51 but below 5, and peft is held below 0.18.0. Those caps matter because the project depends on both funasr and modelscope. Upgrading one of those libraries independently is the most likely way to break an otherwise working install.
Installing GPT-SoVITS on Linux, Windows and macOS
The README lists tested combinations rather than a single supported environment: Python 3.10 or 3.11 with PyTorch 2.5.1 and CUDA 12.4, Python 3.11 with PyTorch 2.7.0 and CUDA 12.8, Python 3.9 with PyTorch 2.5.1 on Apple silicon, and Python 3.9 with PyTorch 2.2.2 on CPU. Start from one of those rows rather than from whatever is already on your machine.
The Linux path is a conda environment followed by the install script. The device and source flags are not optional in the documented form:
conda create -n GPTSoVits python=3.10
conda activate GPTSoVits
bash install.sh --device <CU126|CU128|ROCM|CPU> --source <HF|HF-Mirror|ModelScope>Substitute one value from each group, for example CU126 and HF. The source flag selects where model weights are fetched from, which is why HF-Mirror and ModelScope appear as alternatives to Hugging Face. Add --download-uvr5 if you want the accompaniment separation models, which the README treats as optional.
On Windows, the README offers an integrated package first: download the archive, extract it, and double-click go-webui.bat. If you prefer to install from source, the PowerShell equivalent is:
conda create -n GPTSoVits python=3.10
conda activate GPTSoVits
pwsh -F install.ps1 --Device <CU126|CU128|CPU> --Source <HF|HF-Mirror|ModelScope>macOS uses the same install.sh but accepts MPS or CPU for the device. The README attaches an explicit warning to that platform: models trained with GPUs on Macs are significantly lower quality than models trained elsewhere, so the project uses CPUs instead. Plan to train on a different machine and only run inference on the Mac.
For a first real use, launch the WebUI from the project root after installation:
python webui.pyThe Gradio interface is where the bundled tools live. The documented first pass is to feed it a vocal sample, run the separation and segmentation tools to cut the audio into training segments, let the ASR step transcribe them, and correct the labels before training. The 5-second zero-shot path skips training entirely and is the faster way to judge whether the voice quality is acceptable for your use case.
If you would rather not manage the environment, docker-compose.yaml defines four services. The full and Lite variants exist for both CUDA 12.6 and 12.8, and the Lite images omit ASR and UVR5 models. The compose file maps ports 9871, 9872, 9873, 9874 and 9880, sets is_half=true, and requests shm_size of 16g. Run a specific service with:
docker compose run --service-ports GPT-SoVITS-CU126-LiteThe README notes that Docker Compose mounts all files in the current directory, so run it from the project root after pulling the latest code, and increase shm_size on Docker Desktop for Windows if you see unexpected behaviour.
Where GPT-SoVITS fails, and the cases it was not built for
The largest limitation is not in the model. It is in the input. Zero-shot from 5 seconds works, but 5 seconds of noisy, reverberant, or music-backed audio produces a voice that carries those artifacts forward. The bundled UVR5 separation and ASR tools exist precisely because raw audio is usually unusable, and they add their own failure modes: separation can leave spectral residue, and ASR transcription of accented or domain-specific speech needs manual correction before training. The README lists text labeling as a tool, which is an admission that automatic transcription is a starting point.
Platform support is uneven in a way that affects quality, not just speed. The macOS note is the clearest example: the project chooses CPU training on Macs because GPU-trained models there are significantly worse. That is a deliberate downgrade, and it means a Mac-only user is effectively limited to inference and zero-shot work.
The dependency pins are a second constraint. numpy below 2.0, gradio below 5, peft below 0.18.0, transformers between 4.51 and 5, and torchmetrics at or below 1.5 mean this project cannot share an environment with a stack that has moved past those versions. If you are integrating TTS into an existing application, expect a separate virtual environment or container rather than a shared one.
Finally, consent and legality sit outside the code. The repository ships an MIT licence and no voice-consent mechanism. Nothing in the WebUI verifies that the speaker agreed to be cloned. That is a policy gap you have to close yourself, and it is the reason some organizations will rule this project out before evaluating quality at all.
GPT-SoVITS versus RVC, F5-TTS and CosyVoice
The most common comparison is with RVC, and the difference is the task. RVC is a voice conversion system: you supply audio of yourself speaking and it re-renders that performance in another voice. GPT-SoVITS starts from text. If your input is a recording and you want the same words in a different timbre, RVC is the shorter path. If your input is a script, GPT-SoVITS is the one that reads it.
Against F5-TTS and CosyVoice, the distinction is scope rather than architecture. GPT-SoVITS ships a complete data preparation suite in the same interface as training and inference, which is unusual. F5-TTS and CosyVoice are also few-shot TTS systems, but this project's pitch is the whole loop from raw audio to a trained voice inside one WebUI, with the ASR and separation steps included. That is a real convenience and also a real coupling: you inherit the project's choices about which ASR models to use and how to segment audio.
The API surface is worth noting in this context. api.py and api_v2.py exist at the top level, and the Docker image exposes port 9880 alongside the WebUI ports, so the project can be driven as a service rather than only through Gradio. If you are comparing TTS engines for a backend integration, that is the entry point to evaluate, not webui.py.
Editorial conclusion
Adopt GPT-SoVITS if you need a self-hosted, MIT-licensed TTS stack for English, Japanese, Korean, Cantonese or Chinese and you are prepared to curate clean audio yourself. Do not adopt it if you want a managed API with consent controls, or if you expect GPU-quality training on a Mac, since the README states Mac models are trained on CPU. Before committing, verify GPU memory against your target version, confirm your audio can be separated from background music, and read the MIT licence terms for the voice you intend to clone.
Frequently asked questions
What is GPT-SoVITS?
It is a few-shot voice conversion and text-to-speech WebUI written in Python. The README describes zero-shot TTS from a 5-second sample and few-shot fine-tuning from about 1 minute of training data, with cross-lingual inference in English, Japanese, Korean, Cantonese and Chinese.
Is GPT-SoVITS free?
The repository is licensed under MIT, so the code is free to use under those terms. Model weights are downloaded separately through the --source flag, and the README points to Hugging Face, an HF mirror and ModelScope as sources.
How do I install GPT-SoVITS?
On Linux and macOS you create a conda environment and run install.sh with a device flag such as CU126 or CPU and a source flag such as HF. On Windows the README offers an integrated package that starts with go-webui.bat, or install.ps1 for a source install.
How do I use the GPT-SoVITS WebUI?
Start it with python webui.py from the project root. The interface bundles voice accompaniment separation, training set segmentation, multilingual ASR and text labeling, so the documented workflow is to prepare and label a dataset there before training a GPT or SoVITS model.
How do I use the GPT-SoVITS API?
The repository includes api.py and api_v2.py at the top level, and the Docker Compose services publish port 9880 alongside the WebUI ports. The README does not document the request format for those endpoints.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/rvc-boss-gpt-sovits)
Community notes