# OpenVoice: a package versioned 0.0.0, no releases, and two dependency lists that disagree

> OpenVoice is a voice cloning model from MyShell and MIT, released under MIT for commercial and research use, with V2 adding native support for six languages and a different training strategy. The packaging tells a thinner story: the version string is 0.0.0, the repository has no releases, and requirements.txt installs two packages that setup.py leaves out.

**myshell-ai/OpenVoice** — Instant voice cloning by MIT and MyShell. Audio foundation model.

- Repository: https://github.com/myshell-ai/OpenVoice
- Website: https://research.myshell.ai/open-voice
- Stars: 37,749 · Forks: 4,248
- Language: Python
- License: MIT
- Published: 2026-08-17 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/myshell-ai-openvoice

## The package version is a placeholder and there are no releases

setup.py declares the distribution as MyShell-OpenVoice with version set to the literal string 0.0.0. It also points a Changes link at the GitHub releases page. That repository has no GitHub releases at all, so the link has nothing behind it, and 0.0.0 is the version every commit in the history would report.

The repository is not marked archived, and the last push is dated 2025-04-19.

The consequence is that there is no way to ask for a specific build of this project. A dependency resolver sees the same 0.0.0 whether the code is from the V1 era or the V2 era, and pip will not warn you that the weights and the inference code are from different generations. The only honest record of what you have is the commit you cloned and the date on it, which means a reproduction six months from now has to be pinned by commit hash rather than by version, and the fact that V2 arrived in April 2024 as a change of weights and configuration rather than a tagged release makes that worse.

## requirements.txt installs two packages that setup.py leaves out

There are two dependency lists and they do not match. setup.py lists fourteen pinned packages in install_requires. requirements.txt lists the same fourteen and adds two more, both without a version constraint, sitting between whisper-timestamped and pypinyin:

```
whisper-timestamped==1.14.2
openai
python-dotenv
pypinyin==0.50.0
```

So the OpenAI SDK and python-dotenv are present when you install from requirements.txt and absent when you install the package itself.

The consequence is that the two documented routes produce different environments from the same repository, and the difference is invisible until something imports openai and finds nothing. The unpinned entries are the other half of the problem: fourteen neighbours are frozen to exact versions and these two float, so a fresh environment resolves the OpenAI client to whatever is current. The README does not say which file to install from, so the choice is made by whichever file the reader opens first, and it changes what is present.

## Fourteen exact pins and a Python 3.9 floor

Every dependency in both lists is pinned with a double equals sign, and the floor declared in setup.py is python_requires >= 3.9. The pins are specific: librosa 0.9.1, faster-whisper 0.9.0, pydub 0.25.1, wavmark 0.0.3, numpy 1.22.0, eng_to_ipa 0.0.2, inflect 7.0.0, unidecode 1.3.7, whisper-timestamped 1.14.2, pypinyin 0.50.0, cn2an 0.5.22, jieba 0.42.1, gradio 3.48.0, and langid 1.1.6. There is no lock file with hashes, no constraint tier, and no separate extras block.

The consequence is an environment that resolves to those exact releases or does not resolve at all. The demo surface in particular is bound to Gradio 3.48.0, so the interface code is written against that release's API rather than a current one, and anyone who upgrades Gradio to fix a vulnerability is editing the demo rather than updating a constraint. A single unpinned transitive dependency of any of these can still change under you, because exact pins on direct dependencies do not pin what they pull in.

## The README has no install command and points at a docs file

The How to Use section is one sentence: see usage in docs/USAGE.md for detailed instructions. Common Issues points at docs/QA.md. Neither command nor code appears anywhere in the README itself, and the only code block it contains is the BibTeX citation.

The repository root shows what is on offer instead: three notebooks named demo_part1, demo_part2, and demo_part3, a docs directory, the openvoice package directory, a resources directory, requirements.txt, and setup.py.

The consequence is that a new user has three plausible entry points and no instruction for choosing. The notebooks are the fastest look at the model, the package is the route for reuse, and the usage document is the route for anything else, but the README leaves the order to the reader. For a project whose headline capability is instant cloning, the first thing a user has to go and find is how to run anything at all, and the resources directory is where the checkpoints would live without the root file explaining what belongs in it.

## V2 changes the training strategy, not just the checkpoint

The two versions are described separately. V1 claims accurate tone colour cloning, meaning the reference tone colour is cloned and speech is generated in multiple languages and accents. It claims flexible voice style control over emotion and accent plus rhythm, pauses, and intonation. And it claims zero-shot cross-lingual cloning, where neither the language of the generated speech nor the language of the reference needs to appear in the massive-speaker multi-lingual training dataset.

V2, released in April 2024, keeps all of that and changes three things: a different training strategy for better audio quality, native multi-lingual support for English, Spanish, French, Chinese, Japanese, and Korean, and free commercial use under the MIT licence for both V1 and V2.

The consequence is that swapping between versions is not a checkpoint substitution, because the training strategy moved, and the native list is the boundary of what the multilingual claim means in V2. Anything outside those six languages falls back on the V1 claim, which is about generation rather than native training data. The commercial answer also changed in April 2024, so a project that started on V1 under different terms has to re-check what its own use now falls under.

## Tone colour and voice style are separate controls

The two headline capabilities are independent, and reading them as one feature leads to the wrong debugging path. Tone colour cloning is about the reference: take a short sample and reproduce its timbre in the generated speech. Voice style control is about delivery: emotion, accent, rhythm, pauses, and intonation are granular parameters you set on the output regardless of which reference you used.

The consequence is that the two claims fail in different ways. If the speaker sounds wrong, the problem is the reference or the tone colour, and no amount of style tuning fixes it. If the line lands with the wrong intonation or the wrong accent, that is a style parameter, not a bad clone, and the reference may be perfectly good. Zero-shot cross-lingual sits alongside both: the language of the generated speech and the language of the reference do not need to match each other or appear in the training set, which is the claim that makes a mismatched reference and target language a supported case rather than an edge case.

## No commit since 2025-04-19 and one citation from 2023

The licence section is short and unambiguous: V1 and V2 are MIT licensed and free for both commercial and research use, and the project started under those terms in April 2024. The acknowledgements name three upstreams the implementation is based on, TTS from coqui-ai, VITS, and VITS2.

The citation is the part that lags. The only reference the repository offers is the BibTeX entry for the 2023 arXiv preprint 2312.01479, titled OpenVoice: Versatile Instant Voice Cloning, with four authors. Every V2 improvement described in the README postdates that preprint and has no citable artifact in this repository, and no release tags point at the moment each change landed.

The consequence is that the licence question has a clean answer and the provenance question does not. Anyone writing a methods section or an internal note about which model produced a result has the 2023 paper for V1 and only prose for V2, plus a commit date. The contributors are listed by affiliation rather than by contribution, with two from Tsinghua, one from MIT, and one from MyShell, and the deployment history offered is that the model has powered voice cloning on myshell.ai since May 2023.

## Conclusion

OpenVoice fits a researcher who wants tone colour cloning and style control locally, with a MIT licence that permits commercial work and no API keys in the cloning path. It does not fit a team that needs a pinned, reproducible build, because the package version is a placeholder and there are no releases to pin against, and it does not fit someone who needs current language coverage, because the native list is six languages and the code has not moved since 2025-04-19. Before you build on it, decide which of the two dependency files you are installing from, record the commit you used, and read the V1 paper plus the V2 notes yourself, since the only citation in the repository is the 2023 preprint.

## FAQ

### is openvoice free

Yes. The README states that OpenVoice V1 and V2 are MIT licensed and free for both commercial and research use, and that this has applied to both versions since April 2024. The repository carries a LICENSE file at its root.

### What is open voice?

It is a voice cloning model from MyShell and MIT that clones a reference tone colour and generates speech from text, with granular control over voice style including emotion, accent, rhythm, pauses, and intonation. It also does zero-shot cross-lingual cloning, where neither the generated language nor the reference language needs to be in the training dataset.

### how to use openvoice

The README contains no command and defers to docs/USAGE.md for detailed instructions, with docs/QA.md for common questions. The repository root also carries three notebooks, demo_part1.ipynb, demo_part2.ipynb, and demo_part3.ipynb, which are the quickest route to seeing the model run.

### how to install openvoice

The README shows no install command. The packaging metadata declares a distribution named MyShell-OpenVoice requiring Python 3.9 or higher, with fourteen pinned dependencies in setup.py, and requirements.txt adds openai and python-dotenv on top of the same fourteen.

### openvoice vs xtts

The README makes no comparison and names no alternative. What it credits is three upstreams the implementation is built on, coqui-ai TTS, VITS, and VITS2, which is a different relationship from a competitor comparison.

## Sources

- [Official documentation](https://research.myshell.ai/open-voice)
- [Official README](https://github.com/myshell-ai/OpenVoice#readme)
- [Project repository](https://github.com/myshell-ai/OpenVoice)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/myshell-ai-openvoice
