# VoiceStudio (OmniVoice-Studio): local voice cloning, dubbing and dictation in one desktop app

> VoiceStudio is an AGPL-3.0 desktop studio that bundles 16 TTS engines and 11 ASR engines behind one interface. It is for people who want voice cloning and dubbing to stay on their own machine, and it asks for a real GPU and a tolerance for beta software in return.

**debpalash/OmniVoice-Studio** — Local voice clone, video dubbing, dictation and audiobook maker. The open-source ElevenLabs alternative.

- Repository: https://github.com/debpalash/OmniVoice-Studio
- Website: https://palash.dev/omnivoice
- Stars: 47,464 · Forks: 5,350
- Language: Python
- License: AGPL-3.0
- Published: 2026-08-08 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/debpalash-omnivoice-studio

## What VoiceStudio actually replaces

The project ships as VoiceStudio, previously OmniVoice-Studio, and the README frames it as the open-source answer to a hosted voice service. The workflows it bundles are voice cloning from a short reference clip, voice design from age, accent, pitch and style instructions, video dubbing, dictation, multi-voice stories and audiobooks, and batch generation. The at-a-glance table lists 16 TTS engines, 11 ASR engines and a 646-language catalogue, with the caveat that actual coverage and quality depend on the selected engine. That caveat matters more than the headline number.

The intended user is someone who already has a reason to keep audio off someone else's servers: a studio with client material under contract, a researcher working with non-public recordings, or a team doing enough volume that per-character pricing becomes the dominant cost. The README states that voices, projects, settings and outputs stay on the machine by default, and that network-backed features are explicit opt-ins. If you have no such constraint, a hosted service removes model downloads, driver problems and disk management from your week, and the comparison table in the README says as much.

## Engine registry, model catalogue and the local API

The architecture visible from the repository is a monorepo: a Python backend under backend/ and omnivoice/, a frontend/ directory, an electron/ directory for the desktop shell, and deploy/ and infra/ for packaging. The root package.json is named omnivoice-studio-monorepo and declares bun@1.4.2 as the package manager, with turbo.json driving workspace builds.

Models are not hardcoded into the app. The README describes a Model Catalogue where you install, remove, select and route TTS, ASR and LLM models, and a registry-based plugin interface for adding engines. GPU routing is automatic across CUDA, Apple Silicon MPS/MLX, ROCm and CPU, with per-engine checks. That design is why the first launch matters: the app creates a managed Python environment and downloads the default model, and later launches reuse both. Switching engines means downloading weights again, so the disk footprint grows with the number of engines you keep installed.

The app also exposes a local REST, SSE and WebSocket API, an OpenAI-compatible audio API, and an MCP server for synthesis and transcription. The dev scripts confirm the backend listens on port 3900, since the wait:api script polls http-get://localhost:3900/system/info and the predev script clears ports 3900 and 3901.

## How to install OmniVoice Studio and clone a first voice

The README directs most users to a packaged build rather than a source checkout. Download from the latest release page: a DMG for macOS 13.3 and later on Apple Silicon, an MSI for Windows 10 and 11 x64, or an AppImage for Linux x86_64 with glibc 2.39 or newer. Docker images exist for CUDA, ROCm and CPU, including worker-only GPU profiles. On macOS the first launch needs a one-time right-click then Open approval, and the README is explicit that Intel Macs cannot run the local Python backend and must point at a remote backend instead.

After launch, the documented path to a first result is three steps. Open Voice Cloning, add a clean voice sample (three seconds works, 5 to 15 seconds usually gives a better prompt), then enter text, choose a language and select Generate. The first launch is slower than later ones because it builds a managed Python environment and fetches the default model.

If you prefer to run from source, the README gives this sequence. It assumes the development prerequisites from the contributing guide are already installed.

```bash
git clone https://github.com/debpalash/VoiceStudio.git
cd VoiceStudio
bun install
bun run desktop
```

The README notes that `bun run dev` starts the browser UI instead of the desktop shell. When something fails, the README points at two things. Inside the app, Settings then About then Run self-check. From a terminal, the deep diagnostic:

```bash
uv run python backend/main.py --diagnose --deep
```

The README also says to check the install troubleshooting guide, to save a scrubbed diagnostic bundle from the app when opening an issue, and to compare measured benchmarks and performance settings if generation feels slow. Note the dependency pinning in pyproject.toml: setuptools is capped below 80 because whisperx and faster-whisper import pkg_resources at runtime, and an unpinned requirement resolves to a version without it, which breaks WhisperX transcription and makes its availability check report false. If you manage dependencies yourself, that cap is load-bearing.

## Where VoiceStudio is the wrong tool

The README labels the project an active beta and tells you to use the latest release for stable work or main for current fixes. That is a fair warning, not boilerplate: the last push was on 2026-08-28, and v0.5.1 shipped the same day, so the codebase is moving quickly and interfaces can shift between releases.

Hardware is the second constraint. The recommended path assumes a discrete GPU; CPU is listed as a supported compute target but the README's own comparison table puts performance in the depends-on-your-hardware column. If you need sub-second synthesis for a live product, a local model on consumer hardware is the wrong architecture regardless of which engine you pick.

The third constraint is licensing. The project is AGPL-3.0-only, and the pyproject metadata states that a commercial licence is available for proprietary or closed-source use without AGPL obligations. If you plan to embed this in a product you do not open-source, that is a decision to make before you build, not after. Finally, the bundled omnivoice TTS model by Han Zhu remains Apache-2.0 upstream, and optional engines keep their own model licences, so the licence picture is per-engine rather than uniform.

## VoiceStudio against a hosted API and against building your own pipeline

The honest alternative is the one the README itself names: a hosted voice service. The difference is where the audio goes and who owns the failure. With a hosted provider you create an account, call an endpoint, and pay per use; the provider handles model weights, scaling and driver compatibility. With VoiceStudio you install the app and the weights, and you manage updates, disk and compute. The README's comparison table puts offline use on the local side and fast setup on the hosted side, which is the trade in one line.

The other alternative is assembling the stack yourself: a TTS model plus WhisperX for transcription plus Demucs for vocal isolation plus Pyannote for diarization, wired together with your own scripts. VoiceStudio wraps those same categories of component behind one interface, a batch queue, and a local API. If you only ever need one engine and one language, that wrapper is overhead. If you need dubbing with speaker preservation, chapter rendering to .m4b, and an MCP server for agent clients, assembling it yourself is the larger project.

## Upgrade cost, licence obligations and what to check first

The repository carries alembic.ini, which indicates schema migrations for a local database. An upgrade that changes the schema can require a migration step, and the package.json includes a desktop-prod:upgrade script that runs the production desktop build with --keep-data, alongside --keep-models for preserving downloaded weights. Those flags are the documented way to move versions without re-downloading everything.

The AGPL-3.0-only licence is the practical constraint for commercial adopters. The pyproject metadata states plainly that a commercial licence is available for proprietary or closed-source use without AGPL obligations, with a contact address. That is a statement of the project's licensing position, not legal advice; if you distribute a modified version or expose it over a network, read the licence text in LICENSE and LICENSE-NOTICE.md and get your own counsel.

The first things worth verifying on your own machine: that the engine you intend to use reports ready in the Model Catalogue rather than silently falling back to CPU, that your GPU routing matches what the README promises for your platform, and that the languages you need are actually covered by that engine rather than merely present in the 646-language catalogue.

## Conclusion

Adopt VoiceStudio if your audio cannot leave your hardware, you already have a CUDA or Apple Silicon machine, and you can live with the README's own label of active beta. Do not adopt it if you need a hosted endpoint with an uptime guarantee, if you are on an Intel Mac (the README says the local Python backend cannot run there), or if you want to ship a closed-source product on top of it without a commercial licence. Before committing, install the model you actually intend to use, run Settings then About then Run self-check, and confirm the engine reports ready rather than falling back to CPU.

## FAQ

### Is OmniVoice Studio free to use?

The software is free and open source under AGPL-3.0, and the README states there is no account, API key, subscription or usage meter for the core workflow. You supply the hardware and the disk for model weights. A commercial licence is offered separately for proprietary or closed-source use.

### Is voice cloning illegal?

The repository does not address the legality of voice cloning; it ships an AI watermark feature using AudioSeal embedding and detection, but the README makes no legal claim about cloning a real person's voice. Consent and local law are outside what this material covers.

### How to install OmniVoice Studio?

Download a packaged build from the latest release: a DMG for macOS 13.3+ on Apple Silicon, an MSI for Windows 10/11 x64, or an AppImage for Linux x86_64 with glibc 2.39+. Docker images cover CUDA, ROCm and CPU. First launch creates a managed Python environment and downloads the default model.

### What is OmniVoice Studio?

It is a local-first desktop studio for voice cloning, voice design, video dubbing, dictation, stories and audiobooks, now named VoiceStudio. The README lists 16 TTS engines, 11 ASR engines and a 646-language catalogue, with quality depending on the engine you select.

### How to use OmniVoice Studio?

Open Voice Cloning, add a clean voice sample of roughly 3 seconds (5 to 15 seconds usually gives a better prompt), then enter text, choose a language and select Generate. Engines are switched in the Model Catalogue or with Ctrl/Cmd+E.

### What is the most realistic TTS voice in OmniVoice Studio?

The README does not rank engines by realism; it states that actual language coverage and quality depend on the selected engine. The repository points to docs/benchmarks.md for measured comparisons, which is the only place a claim like this could be checked.

## Sources

- [Official documentation](https://palash.dev/omnivoice)
- [Official README](https://github.com/debpalash/OmniVoice-Studio#readme)
- [Project repository](https://github.com/debpalash/OmniVoice-Studio)
- [Release notes](https://github.com/debpalash/OmniVoice-Studio/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/debpalash-omnivoice-studio
