Real-Time-Voice-Cloning tells you in its own README that it has gotten old
GitHub describes it as Clone a voice in 5 seconds to generate arbitrary speech in real-time. The repository metadata lists Python as its primary language. The metadata lists the NOASSERTION license. This article stays within the project description and details documented in the GitHub repository README.
At a glance
- What is it?
- A Python implementation of SV2TTS with a real-time vocoder, runnable as a GUI toolbox or a command line tool through uv. The project's own heads-up section says commercial services will give you better audio, and points you at Chatterbox instead, which is the most useful thing on the page.
- Who is it for?
- Adopt this repository if you want to read or modify SV2TTS, need the encoder and synthesizer training scripts side by side with a real-time vocoder, and are willing to hold an old Python and an old torch. Do not adopt it to produce the best voice output you can get, because the project itself says paid services will beat it and recommends Chatterbox for that.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Activity is slowing. The repository last received commits 6 months ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 26, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The heads-up section recommends a different project
The most valuable paragraph in this repository is the one warning you off it. It says that like everything else in deep learning, this repo has quickly gotten old, and that many SaaS apps, often paying ones, will give better audio quality than this repository will. It then does something unusual for a project README and names a replacement, pointing at Chatterbox as a similar project up to date with the 2025 state of the art in voice cloning, and at paperswithcode for other repositories and recent research. That framing should govern the rest of this article. The code is worth reading because it is a compact implementation of a specific paper, and it is worth running if you need exactly this architecture, but nobody should arrive expecting current output quality and be surprised later.
SV2TTS is three stages sharing one speaker representation
The architecture is a transfer learning scheme rather than a single model, and understanding the split explains the file layout. SV2TTS is described as a deep learning framework in three stages. In the first, a digital representation of a voice is created from a few seconds of audio. In the second and third, that representation is used as a reference to generate speech from arbitrary text. The point of the design is that the expensive part, learning what a voice sounds like, happens once in stage one, and stages two and three reuse it, which is what makes five seconds of audio enough. The repository description condenses this to cloning a voice in 5 seconds to generate arbitrary speech in real-time, and the real-time part comes from the vocoder rather than from the text-to-speech stages.
Two of the four papers are implemented here and two are borrowed
A table credits each component to its source, and the split matters if you plan to cite or replace anything. SV2TTS itself, arXiv 1806.04558 on transfer learning from speaker verification to multispeaker text-to-speech, and the GE2E encoder, arXiv 1710.10467, are both marked as implemented in this repository. The other two are not. The WaveRNN vocoder, arXiv 1802.08435, and the Tacotron synthesizer, arXiv 1710.10467 as cited in the table, both point at fatchord/WaveRNN as the implementation source. The repository layout follows that split directly, with encoder/, synthesizer/ and vocoder/ as sibling directories and a matching train script beside each, encoder_train.py, synthesizer_train.py and vocoder_train.py, along with separate preprocessing scripts for audio and for embeds. The work began as the author's master's thesis.
ffmpeg, then uv, then one of four run commands
Windows and Linux are both supported, and the setup is three steps. First ffmpeg, which is necessary for reading audio files, verified by running it on a command line. Second uv for package management, installed per platform or through pip. Third, one of four commands, choosing between the GUI toolbox and the command line, and between CUDA and CPU:
uv run --extra cuda demo_toolbox.pyUse the cpu extra instead of cuda if you have no NVIDIA GPU, and demo_cli.py in place of demo_toolbox.py if you do not want the GUI. The two extras conflict with each other by design in the project configuration, so you pick one. uv creates a .venv with an appropriate Python environment automatically, and the project asks you to open an issue if that fails. Pretrained models download automatically, and the manual fallback is a Hugging Face repository for the author, so the model weights are not something you assemble yourself.
torch is pinned at 1.10 and Python at 3.9
The dependency file is where the age shows, and it is worth reading before you plan an install. requires-python is >=3.9,<3.10, so this is a single minor version and a modern interpreter will refuse to resolve. Both optional dependency groups, cpu and cuda, pin torch to 1.10, and the CUDA index the project maps that torch to is the cu113 wheel index, meaning CUDA 11.3. The rest of the list is pinned the same way, with librosa at 0.9.2, matplotlib at 3.5.1, Pillow at 8.4.0, scikit-learn at 1.0.2, soundfile at 0.10.3.post1, tqdm at 4.62.3 and umap-learn at 0.5.2, while numpy and scipy take ranges. There is a uv.lock alongside it, and the configuration restricts resolution to Windows and Linux on x86_64 so the lock does not select wheels that do not exist on Windows.
The last push was 2026-03-09 and there are no releases
Two facts bound what to expect from this project. The repository has no GitHub releases at all, so there is no version to pin and no changelog to read, and the root package version in pyproject.toml is 0.0.0. The last push to the repository was on 2026-03-09, which is more than six months before now, so this is not a codebase receiving regular work, and the last-push date is the honest signal of where it stands. That is consistent with the heads-up section and with the pinned dependencies, and the three reinforce each other rather than telling different stories. The licence file is present at the repository root, though the project facts record no standard licence identifier, so anyone planning to build on the code should read that file rather than assume a grant.
Chatterbox is the alternative the project nominates
The alternative here is named by the project itself rather than by a comparison site, which makes the difference in approach unusually easy to state. This repository implements a 2018 transfer learning scheme whose speaker representation is produced by a verification encoder and consumed by a borrowed synthesizer and vocoder, and its claim to distinction is that the vocoder runs in real time inside a single local process. Chatterbox is described as a similar project that is up to date with the 2025 state of the art, meaning the architecture and the model quality are where the two diverge rather than the interface. For a reader the practical split is simple. Use this one to study or modify SV2TTS, and follow the project's own pointer if the objective is audio a listener would not know was generated.
Editorial conclusion
Adopt this repository if you want to read or modify SV2TTS, need the encoder and synthesizer training scripts side by side with a real-time vocoder, and are willing to hold an old Python and an old torch. Do not adopt it to produce the best voice output you can get, because the project itself says paid services will beat it and recommends Chatterbox for that. Verify first your Python version, since requires-python is >=3.9,<3.10 and nothing else will resolve, then check whether the CUDA 11.3 wheel index it points at still serves what you need before you plan an install around a GPU.
Frequently asked questions
What does the Real-Time-Voice-Cloning toolbox need installed?
ffmpeg, which is necessary for reading audio files, and uv for package management, which creates the .venv automatically. The run command then chooses the cuda or cpu extra, and demo_toolbox.py or demo_cli.py for the GUI or the command line.
Does Real-Time-Voice-Cloning run on Windows and on Linux?
Both are stated as supported, and the uv configuration restricts resolution to Windows and Linux on x86_64 so the lock file does not pick wheels missing on Windows. The installation instructions give separate uv commands for each platform.
Can I run Real-Time-Voice-Cloning without an NVIDIA GPU?
Yes, by using the cpu extra instead of cuda in the run command, for example uv run --extra cpu demo_toolbox.py. The two extras are declared as conflicting so that one is chosen rather than both.
Where do the pretrained models come from?
They are downloaded automatically. If that does not work, the project says to fetch them manually from a Hugging Face repository published by the author, which is where the SV2TTS weights are hosted.