# Genie-TTS measures a 1.13 second first inference on one laptop CPU, pins one runtime to dodge a regression, and ships voices from three game franchises

> An ONNX inference engine and model converter for GPT-SoVITS, installable from pip with Python 3.10 or newer. The comparison table discloses its own test conditions, which is rarer than it should be and more useful than the numbers themselves, and the rest of the page is mostly about what the first run downloads and what the bundled voices are.

**High-Logic/Genie-TTS** — GPT-SoVITS ONNX Inference Engine & Model Converter

- Repository: https://github.com/High-Logic/Genie-TTS
- Stars: 1,787 · Forks: 122
- Language: Python
- License: MIT
- Published: 2026-09-16 · Updated: 2026-09-16 · Language: en
- Canonical page: https://hysenlabs.com/projects/high-logic-genie-tts

## The latency table is one cold start on one laptop processor

The page claims near-instantaneous speech synthesis and then publishes the numbers behind it, which is the right order. The row is first inference latency: 1.13 seconds for this engine, 1.35 for the official PyTorch model, and 3.57 for the official ONNX model. Two rows below it compare footprint, with a runtime size around 200MB against several gigabytes for PyTorch, and a model size around 230MB against around 750MB for the official ONNX build. Two of those cells read as similar rather than as a number, so the runtime comparison is really PyTorch against the ONNX runtime, and the model comparison is really this engine's weights against the other ONNX weights.

The test conditions sit in a note under the table: a set of 100 Japanese sentences of about 20 characters each, averaged, on an i7-13620H. One processor, one language, one sentence length, and the metric is the cold start rather than the steady state, which is what a first inference measures.

## One dependency is pinned, and the reason is written in the manifest

The dependency list is long and almost entirely unpinned, with exactly one version lock in it.

```text
onnxruntime==1.22.1
```

The comment beside it, in Chinese, says the runtime version is fixed to avoid the performance regression in 1.23. Everything else floats: the ONNX format library, the tokenizer, numpy, the audio reader and writer, a high quality resampler, YAML and schema libraries, an audio playback library, the hub client with its transfer extras, the web framework and its server, and a set of grapheme to phoneme packages, one family per language. Japanese gets one, English gets the NLP toolkit, Chinese gets a pinyin library, a model, and a fast segmenter, and Korean gets four. The whole point of that grapheme layer is that each language front end is a different library, so a language support claim on this project is really a claim about a dependency it does not write.

## The declared homepage points at a differently named repository

The package metadata lists the project URL as a repository called Genie, while the project itself lives in a repository called Genie-TTS. That is the kind of mismatch that sends a reader to an empty page from a package index, and it is worth knowing about because the model files this library depends on are not distributed through the package index at all: they are hosted in a HuggingFace space under the same shortened name, which is where the character models and the resource bundle are published. So the chain a new user follows runs from one repository name to another name for the code and a third location for the weights, and only the install line itself is unambiguous.

```bash
pip install genie-tts
```

The manifest also carries a version that matches the newest tag, and a single optional dependency for a graphical interface toolkit that the page itself never mentions.

## The data directory has to be set before the import

The first run downloads about 391MB of resource files, and the page gives two ways to handle that. The default is the library's own prompt flow. The alternative is to fetch the same files by hand and point the library at them, which requires an ordering constraint that the example spells out twice.

```python
import os

# Set the path to your manually downloaded resource files
# Note: Do this BEFORE importing genie_tts
os.environ["GENIE_DATA_DIR"] = r"C:\path\to\your\GenieData"

import genie_tts as genie
```

The constraint exists because the library resolves its resources during import rather than at first call, so an environment variable set afterwards is too late. That is also why the example uses an absolute Windows path with a raw string. Alongside this, the manifest ships model and key directories as package data for both supported model versions, so some resources travel with the wheel while the 391MB bundle does not.

## The Chinese prosody assets are opt-in and excluded from other languages

One feature has an unusually clear warning attached to it. There are optional Chinese RoBERTa assets that improve Chinese prosody, downloadable on their own or as part of the full resource flow, and the page states in bold that they are intended only for the Chinese path and should not be used for Japanese, English, or Korean inference. That is a narrower instruction than most optional features get, and it points at a real asymmetry in the pipeline: the Chinese front end has a learned text feature stage and the other three do not, so a Japanese or English user who triggers the full download pays for assets their language will ignore. The language codes accepted by the character loader are given as four short strings in the best practices example, and the Chinese one is the only path where that download matters.

## The three bundled voices are characters from commercial games

The page offers a way to try the engine without owning a model, and it does so with three ready-made characters.

```python
import genie_tts as genie
import time

# Automatically downloads required files on first run
genie.load_predefined_character('mika')

genie.tts(
    character_name='mika',
    text='どうしようかな……やっぱりやりたいかも……！',
    play=True,  # Play the generated audio directly
)

genie.wait_for_playback_done()  # Ensure audio playback completes
```

Each is named after a role from a commercial game title: a character from Blue Archive, a character from Reverse: 1999, and a character from Wuthering Waves, one per supported language. They are hosted in the project's HuggingFace space rather than in the repository, so they arrive with the first-run download and disappear if you point the library at your own folder. whoever uses them commercially is asking their own question about voice rights; the page names the titles and leaves it there.

## The documented server call binds every interface and names no authentication

A lightweight web server ships with the library, and the example that starts it looks like this.

```python
import genie_tts as genie

# Start server
genie.start_server(
    host="0.0.0.0",  # Host address
    port=8000,  # Port
    workers=1  # Number of workers
)
```

The host in the documented call is all interfaces rather than the loopback address, the worker count is one, and the port is the default for this kind of framework. Nothing on the page describes an authentication scheme, a token, or a rate limit, and the pointer for request formats goes to a tutorial file in the repository rather than to a section here. So the server is presented as something you start on your own machine, and the page leaves the exposure decision entirely to the operator who changes that host value.

## Two releases, both shipped as bundles, and ten months without one

The release record has two entries, and both are titled as integrated bundles that work after extraction rather than as source drops: version 1.0.2 in September 2025 and version 2.0.2 in December 2025. The manifest version matches the newer one. The last commit on the default branch is dated 2026-08-30, which puts roughly ten months of work on the branch with no release cut from it, so the published artifact and the repository head are far apart. Two other version surfaces exist side by side: the packaging manifest lists twenty-one dependencies and a requirements file at the root lists the same twenty-one in a different order, in both cases with Korean grapheme tools at the end. The roadmap explains the release shape as well: Windows bundles are checked off, official Docker images are not, and support for the third and later model versions is not, with the page stating that conversion supports the second version and the Pro Plus variant.

## Conclusion

Take it if you already run GPT-SoVITS models and want CPU inference without pulling in a training stack, since the runtime depends on ONNX rather than PyTorch and the conversion path is a separate step. Three checks first. The benchmark is one cold start on one laptop processor with twenty-character Japanese sentences, so measure your own text before quoting the 1.13 seconds. The first run fetches about 391MB from a HuggingFace space unless you point the library at a local folder, and the environment variable has to be set before the import or the download happens anyway. And the three ready-made voices are characters from commercial game titles, which is a licensing question for you rather than for the project.

## FAQ

### What does Genie-TTS need before it can synthesise anything?

A first run downloads about 391MB of resource files, either through the library's own prompts or by hand from the project's HuggingFace space. Setting GENIE_DATA_DIR before importing the library points it at a folder you filled yourself.

### Can Genie-TTS convert a GPT-SoVITS model, and what does conversion require?

Yes. convert_to_onnx takes a .pth model file, a .ckpt checkpoint, and an output directory, and currently supports the V2 and V2ProPlus model versions. Converting requires torch, which is not one of the package's dependencies.

### Which languages and model versions does Genie-TTS support?

Japanese, English, Chinese, and Korean on GPT-SoVITS V2 and V2ProPlus models, with Python 3.10 or newer and classifiers through 3.13. Support for the V3 and V4 model versions appears on the roadmap and is still unchecked.

## Sources

- [High-Logic/Genie-TTS on GitHub](https://github.com/High-Logic/Genie-TTS)
- [Issues](https://github.com/High-Logic/Genie-TTS/issues)
- [License: MIT](https://github.com/High-Logic/Genie-TTS/blob/master/LICENSE)
- [README](https://github.com/High-Logic/Genie-TTS/blob/master/README.md)
- [Releases](https://github.com/High-Logic/Genie-TTS/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/high-logic-genie-tts
