kokoro-onnx hands you a pip install and two model files you fetch yourself
TTS with kokoro and onnx runtime
At a glance
- What is it?
- kokoro-onnx wraps the Kokoro TTS model in onnxruntime so speech synthesis runs without a deep learning framework. The install is one command and the package is small, but the weights are not on PyPI: they come from a GitHub release you download and place next to your script.
- Who is it for?
- kokoro-onnx is a good fit if you want local text to speech without installing a framework stack, and if you are willing to fetch the weights once and keep them beside your code. Before you start, settle four things.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 34 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Pip installs the code, the two model files come from a release page
The whole installation is one console line:
pip install -U kokoro-onnxThat gets you the Python package and none of the model. The detailed walkthrough, folded away behind an Instructions summary, adds a manual download step: fetch kokoro-v1.0.onnx and voices-v1.0.bin from the model-files-v1.1 release and place them in the same directory as your script. There is no model path setting, no cache directory variable and no first-run downloader in the description, so the location of those two files is a property of your working directory rather than a configured value.
The rest of that walkthrough uses uv rather than pip. It recommends installing uv for isolated Python, creating a new project folder, then:
uv init -p 3.12
uv add kokoro-onnx soundfileThen you paste the contents of examples/save.py into a file called hello.py, keep the two downloaded files beside it, and run:
uv run hello.pyThe text to speak is edited inside hello.py, and the stated result is that audio.wav appears. So the first successful run depends on three things lining up: a script, two model files, and a working directory.
Three version numbers, and the releases carry only models
There are three separate version lines in this repository and they do not line up.
The package version is 0.6.1, written in pyproject.toml and published to the package index the project links as its homepage. The GitHub releases are a different thing entirely: the three recorded releases are named model-files, model-files-v1.0 and model-files-v1.1, dated 2025-01-03, 2025-01-28 and 2025-03-01. Those are asset releases, which is why the instructions point at the model-files-v1.1 release for the two downloads.
So the third line is the model itself. The tag called model-files-v1.1 hosts files named kokoro-v1.0.onnx and voices-v1.0.bin, and the page's headline announcement is that version 1.0 models are out. A reader comparing a tag name to a filename gets two different version numbers for the same weights.
What none of this covers is the code. No tag corresponds to 0.6.1, the most recent release of any kind is from March 2025, and the last commit on the branch is dated 2026-09-01. Anyone who needs to know which package build they are running has to read the installed package version, because the release history will not tell them.
The documentation follows the same split. The homepage link points at the package index rather than at a documentation site, the model announcement is a bare attachment link with no caption, and the usage information is the walkthrough plus the examples directory. BUILDING.md and CONTRIBUTE.md cover the two workflows a contributor needs and a user does not.
The GPU extra installs nothing on the machine the speed claim names
The feature list claims fast performance near real time on macOS M1, and then the packaging quietly rules out GPU support on that machine. The optional gpu group contains one dependency, onnxruntime-gpu at 1.20.1 or newer, behind the marker platform_machine == 'x86_64' and sys_platform != 'darwin', with a comment in the file explaining that onnxruntime-gpu is not available on Linux ARM or macOS.
Read together, the extra resolves to nothing on an Apple Silicon Mac and nothing on a Linux ARM machine. It exists for x86_64 Windows and Linux. The base dependency in every install is plain onnxruntime at the same 1.20.1 floor, so an M1 user gets the CPU build, and the near real time claim is about that build rather than about acceleration.
The examples directory does cover both paths. with_cuda.py is the GPU variant, and with_quant.py sits beside it, which lines up with the other size figure in the feature list: roughly 300 MB normally and about 80 MB quantized. Those two scripts are where the acceleration and the size claims can be looked at, rather than in anything the packaging says.
The feature list itself has four entries, and they are uneven in how far they can be checked. Multiple languages and multiple voices point at files maintained elsewhere, since the voice list is a document in the model repository. The near real time claim names one machine and no measurement, and it names the machine where the GPU extra does not apply. The size figures are the only two numbers on the page that have an artefact behind them, and one artefact is 300 MB of weights you download yourself.
The quick start uses uv, and the plain pip path lacks the audio libraries
Two installation routes are on offer and they deliver different things. The uv route adds kokoro-onnx and soundfile together, which is why the soundfile argument is spelled out explicitly in the walkthrough. The plain pip route is a single upgrade of kokoro-onnx and nothing else.
That difference matters because of where the audio libraries sit. In pyproject.toml the dependency groups list dev as ruff, sounddevice and soundfile, so the library that writes a wav file is not a runtime dependency of the package. A user who runs pip install -U kokoro-onnx and then uv run hello.py, or who copies save.py into their own project, will not have soundfile unless they add it. The playback example, play.py, is in the same position with sounddevice.
The Python range is pinned on both sides, from 3.10 up to but not including 3.14, while the walkthrough creates the project with Python 3.12. Reproducibility is set up for the uv path rather than the pip one: uv.lock and a .python-version are committed to the repository, so a uv workflow resolves to pinned versions and a pip workflow does not.
Two scripts for one phonemizer dependency
The examples directory holds twenty-five scripts, and twelve of them are variations named with a with_ prefix. The clearest pair is with_espeak_data.py and with_espeak_lib.py, which differ in how the espeak data reaches the phonemizer. The runtime dependencies behind that are espeakng-loader at 0.2.4 or newer and phonemizer at 3.4.0 or newer, plus numpy at 2.0.2 or newer. Whether you need the data variant or the library variant is decided by your platform, and the two script names are the only place that choice is spelled out.
with_provider.py belongs to the same group of decisions. The voices section recommends using the misaki g2p package from v1.0 and points back at the examples, so the grapheme to phoneme provider is selectable rather than fixed, and the phoneme path itself is exposed again in with_phonemes.py.
Nine of the twenty-five scripts are per-language: chinese, english, french, hebrew, hindi, italian, japanese, portuguese and spanish. The remaining four are app.py, play.py, podcast.py and save.py. The full list of voices and languages is not kept here at all; it lives in a VOICES.md file on the Kokoro-82M model repository, so the set of languages you can actually synthesise is defined by that file rather than by this directory.
The rest of the with_ names read as a list of the knobs the wrapper exposes. There is with_blending.py, with_continuous.py, with_session.py, with_stream.py and with_stream_save.py, plus with_log.py for logging. Each is a separate file rather than a flag, so these behaviours are shown as examples you copy rather than as arguments you pass, and the file names are the only catalogue of them.
A vendored file the linter is told not to touch
The lint configuration carries its own explanation of a constraint. The tool is ruff, pinned at 0.16.3 or newer in the development group, with required-version set to 0.16.0 or newer, concise output and fixes shown. The selected rule set is E4, E7, E9, F, I and UP, and the comment above that line says ruff 0.16 and later enable a much wider default rule set, so the previous selection is kept, meaning ruff's pre-0.16 default plus isort and pyupgrade, in order that vendored code in a file called trim.py stays untouched.
That single sentence tells you there is a third party file in the source tree being carried as a copy rather than as a dependency, and that the lint configuration was narrowed specifically to avoid rewriting it. Isort is also configured with split-on-trailing-comma set to false, which is another sign that the file's formatting is being preserved rather than normalised.
The build itself uses hatchling, the source lives under src/, and the root carries BUILDING.md and CONTRIBUTE.md next to the examples and scripts directories, so the packaging and contribution workflows are documented separately from the install walkthrough.
Two licences cover the stack you are actually shipping
The licence section is two lines long and the split matters. The kokoro-onnx code is MIT. The kokoro model is Apache 2.0. Those are different artefacts with different obligations, and a product that embeds the generated audio rather than the weights still has to account for where the weights came from.
The model side is not in this repository either. The project describes itself as text to speech with onnx runtime based on Kokoro-TTS, links to a Kokoro-TTS space, and points at Kokoro-82M for the voice list. The weights arrive from a release page in this repository, but the model they contain comes from that other project, which is where the Apache 2.0 terms apply.
So the dependency chain has three parties: this repository for the wrapper, the model project for the weights, and the espeak and phonemizer packages for turning text into phonemes. Anyone redistributing a build should read the model licence rather than assuming the MIT line at the top covers it.
Editorial conclusion
kokoro-onnx is a good fit if you want local text to speech without installing a framework stack, and if you are willing to fetch the weights once and keep them beside your code. Before you start, settle four things. Know that the GPU extra installs nothing on Apple Silicon, so an M1 Mac runs the CPU build the README calls near real time. Know that the audio writing and playback libraries live in the development group, so the plain pip path will not save a wav file until you add one. Know that the package version on PyPI and the release tags in the repository are two different version lines, with no tag for the code. And read both licences, MIT for this code and Apache 2.0 for the model.
Frequently asked questions
how to use kokoro onnx
Install it with pip install -U kokoro-onnx, or use the uv walkthrough: uv init -p 3.12, uv add kokoro-onnx soundfile, paste examples/save.py into hello.py, place kokoro-v1.0.onnx and voices-v1.0.bin downloaded from the model-files-v1.1 release in the same directory, then run uv run hello.py to produce audio.wav.
what is kokoro onnx
It is text to speech with onnx runtime based on Kokoro-TTS. The package wraps the Kokoro model in onnxruntime, supports multiple languages and multiple voices, and is described as near real time on macOS M1 at roughly 300 MB, or about 80 MB quantized.
Does kokoro-onnx use the GPU on Apple Silicon?
No. The gpu extra installs onnxruntime-gpu only where platform_machine is x86_64 and sys_platform is not darwin, with a note that onnxruntime-gpu is not available on Linux ARM or macOS. The base install uses plain onnxruntime, and examples/with_cuda.py is the CUDA variant.
Which Python versions does kokoro-onnx support?
pyproject.toml sets requires-python to >=3.10,<3.14. The setup walkthrough creates the project with uv init -p 3.12, and the repository commits both uv.lock and a .python-version file.
What licences apply to kokoro-onnx and the Kokoro model?
Two. The kokoro-onnx code is MIT and the kokoro model is Apache 2.0. The model weights are downloaded from a release in this repository but originate from the Kokoro project, which is where the Apache terms apply.
Is Kokoro Sound AI the same project as kokoro-onnx?
Nothing in this repository connects the two. kokoro-onnx describes itself as text to speech with onnx runtime based on Kokoro-TTS, and it points at a Kokoro-TTS space and at the Kokoro-82M model repository for voices. No mention of Kokoro Sound AI appears in the README, the packaging or the licence notes.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/thewh1teagle-kokoro-onnx)