# Voice-Pro assembles a full dubbing pipeline, then stops developing

> Voice-Pro chains video download, voice separation, transcription, translation and speech generation into one local web interface on Windows. Version 4 rebuilt the installer so it works without a compiler toolchain, and announced that development is paused.

**abus-aikorea/voice-pro** — Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal isolation, and multilingual translation.

- Repository: https://github.com/abus-aikorea/voice-pro
- Website: https://www.wctokyoseoul.com
- Stars: 12,966 · Forks: 1,871
- Language: Python
- License: GPL-3.0
- Published: 2026-09-18 · Updated: 2026-09-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/abus-aikorea-voice-pro

## A dubbing pipeline assembled into one local interface

Voice-Pro is a web interface that chains together the steps involved in taking a video in one language and producing it in another. Downloading the source, separating voices from background audio, transcribing, translating, then generating speech, each handled by an established open model, driven from a browser interface running on your own machine.

The value is integration rather than invention. Every component here exists separately and every one of them is a project in its own right, and wiring them into a working pipeline is a day of work that goes wrong in specific ways: mismatched sample rates, timestamps that drift, a translation step that silently drops lines. Voice-Pro's contribution is that someone did that assembly and shipped it.

The audience is a content creator or a small team producing multilingual video who has a Windows machine with an NVIDIA GPU and would rather not pay per character to a hosted service. The README positions the project explicitly as an alternative to a commercial voice platform.

## The stack, and the component that was dropped

The components are named precisely rather than described in the abstract.

Transcription runs on the Whisper family in three forms: the original implementation, a faster reimplementation, and a timestamp-focused variant. Zero-shot voice cloning covers three systems, meaning a voice can be reproduced from a short sample rather than trained. Conventional multilingual speech generation comes from two further engines. Source material arrives through a widely used video downloader, and translation across more than a hundred languages runs through a translation wrapper library. Two commercial services can be substituted for the translation and speech steps if you supply your own keys.

One component went the other way. The version 4 notes state that a timestamp alignment library was dropped because its dependency pins blocked the upgrade to the current interface framework, with existing configurations falling back to the faster transcription engine. That is a real capability loss disclosed plainly, and it is the kind of tradeoff that assembling a pipeline from many projects forces on a maintainer eventually: one component's pins hold the rest of the stack still until something gives.

## Version 4 is mostly an installer rewrite, and that is the point

The release published on 2026-07-13 is mostly housekeeping, and it is the most consequential change in the project's history, because installation is where tools like this lose most of their prospective users.

The installer moved from a conda and pip arrangement to a modern Python package manager with a committed lock file, with everything confined to a single directory inside the project. The runtime moved to Python 3.12, a recent Torch build with support for current GPUs, and a new major version of the interface framework. Most significantly, the notes state that a CUDA toolkit and Visual Studio build tools are no longer required, because the dependencies now ship prebuilt and the Torch distribution bundles its own CUDA runtime.

Compiler toolchains and a matching CUDA installation are where Windows installations of machine learning software usually fail, so dropping both removes the most common cause of a failed setup.

A related set of changes targets locked-down machines: no administrator rights needed, a portable media tool downloaded automatically when the system lacks one, model downloads that recover from interrupted transfers, and translation that retries with backoff when the free endpoint rate-limits, reporting failed lines while keeping the originals. Errors were also changed from a brief warning to a notice that stays on screen until dismissed. Every one of those addresses a way the previous version failed quietly.

## Installing it, and the narrow interpreter window

The manifest is strict about the interpreter.

```toml
requires-python = ">=3.12,<3.13"
```

That is a single minor version, not a floor. The installer supplies its own Python inside the project directory, so this is less restrictive in practice than it looks, but anyone intending to run against a system interpreter should know the window is one version wide.

The repository root carries the entry points as scripts rather than commands to compose: a launcher, a configure step, an uninstall script, each in Windows and shell variants, plus a one-click installer script. Starting the application is running the launcher.

```bash
start.bat
```

That bootstraps the environment on first run and opens the web interface afterwards. When something breaks, the documented recovery path is to delete the installer directory and run the launcher again, which the README describes as a clean reinstall taking a few minutes, with already downloaded models preserved in their own directory. Separating the environment from the model cache is a small design decision that turns reinstallation from an overnight job into a coffee break.

Optional commercial services are configured through environment variables, with an example file in the repository naming the speech and translation keys, their endpoint and their region.

## Development has stopped, and the README says so

The README carries a notice that development and updates are not possible for the time being, because the author is occupied with other work.

The project was made fully open source and free at the same time, and the code can be distributed and modified by anyone, so this is a considered handover rather than an abandonment. It still means no fixes, no dependency updates and no response to changes in the services it depends on.

The dependencies are where that bites. This pipeline depends on a video downloader that must track a video platform's changes, a free translation endpoint that rate-limits, and several model projects that continue to move. A downloader that stops working is not a bug in this repository and it will break this repository all the same, and nobody is currently positioned to update it.

The platform position narrows it further. The README states it works well on Windows with an NVIDIA GPU, and that operation on macOS and Linux has not been verified. Not unsupported, not tested. Combined with paused development, a reader on either of those platforms should treat this as a codebase to fork rather than a tool to install.

The licence is stated twice and not consistently. Repository metadata records the GPL, while an author comment in the README header names the lesser variant, and those carry meaningfully different obligations for anyone redistributing a modified build. Resolving which one governs is a prerequisite to forking the project, which is what the paused development invites people to do.

## The commercial platform is the alternative, and the split is control against upkeep

The README names a commercial voice platform as the thing this replaces.

The difference in approach is ownership of the pipeline. A hosted platform gives you current models, a supported interface, no GPU, no installation, and a per-use bill, and it keeps working when a video site changes its interface because that is their problem. It also means your audio and script pass through someone else's systems, your costs scale with output, and your voice cloning is bounded by what their terms permit.

Voice-Pro inverts each of those. The models run on your machine, the cost after hardware is electricity, nothing leaves your computer, and you can modify any stage because the code is yours under an open licence. What you take on is a GPU, an installation, and now the maintenance too, since nobody upstream is doing it.

For occasional dubbing, the hosted service almost certainly wins on effort. For steady volume, for material that cannot be uploaded, or for anyone willing to maintain a fork, this is a complete and working starting point that someone else already debugged. The version 4 installer work is what makes that starting point realistic, and it arrived in the same release that announced the pause.

## Conclusion

Voice-Pro suits a creator on Windows with an NVIDIA GPU who dubs video often enough that per-use pricing hurts, and who accepts that the project is a starting point to maintain rather than a product to receive updates from. It is the wrong choice on macOS or Linux, which the README says have not been verified, and for anyone who needs the pipeline to keep working without their attention, since development is paused and the stack depends on a video downloader and a free translation endpoint that both change underneath it. Check the interpreter window of 3.12 only before starting, and resolve the licence question first, because the repository metadata and the README header name different terms.

## FAQ

### Is Voice-Pro still being developed?

No. The README states that development and updates are not possible for the time being because the author is occupied with other work, and that the code was made fully open source and free so anyone can distribute and modify it.

### Does Voice-Pro work on Mac or Linux?

The README states it works well on Windows with an NVIDIA GPU and that operation on Mac and Linux has not been verified. The repository does ship shell script variants alongside the Windows batch files, but the platform position is untested rather than supported.

### Do I need CUDA Toolkit installed for Voice-Pro?

Not since version 4. The release notes state that the CUDA Toolkit and Visual Studio Build Tools are no longer required, because dependencies ship prebuilt wheels and the Torch distribution bundles its own CUDA runtime.

### How do I fix a broken Voice-Pro installation?

The README's troubleshooting advice is to delete the installer files directory and run the launcher again, which it describes as a clean reinstall taking a few minutes. Previously downloaded models are kept in a separate directory and are not re-downloaded.

## Sources

- [abus-aikorea/voice-pro on GitHub](https://github.com/abus-aikorea/voice-pro)
- [License: GPL-3.0](https://github.com/abus-aikorea/voice-pro/blob/main/LICENSE)
- [Project website](https://www.wctokyoseoul.com)
- [README](https://github.com/abus-aikorea/voice-pro/blob/main/README.md)
- [Releases](https://github.com/abus-aikorea/voice-pro/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/abus-aikorea-voice-pro
