AivisSpeech: a VOICEVOX-based Japanese TTS editor built around AIVMX models
AivisSpeech: AI Voice Imitation System - Text to Speech Software
At a glance
- What is it?
- AivisSpeech is an Electron desktop editor for Japanese speech synthesis, forked from the VOICEVOX editor UI and wired to the AivisSpeech Engine. It reads AIVMX model files only, which is both its main convenience and its main constraint.
- Who is it for?
- AivisSpeech suits Japanese-speaking users on Windows 10 22H2 or later, Windows 11, or macOS 13 Ventura or later who want a desktop editor and are happy to run Style-Bert-VITS2 models in AIVMX form; the README states the Engine supports Japanese synthesis only, so anyone needing English or Chinese output should look elsewhere.
- Can I use it commercially?
- Yes, with conditions. LGPL-3.0 is a weak copyleft licence: you can use it inside commercial and closed-source software, but if you distribute changes to its own files, you must publish those changes under the same licence.
- Is it still maintained?
- Yes. The repository last received commits 79 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What AivisSpeech is, and who it is actually for
AivisSpeech is a Japanese speech synthesis application built on the editor UI of VOICEVOX, with the AivisSpeech Engine embedded so that voice generation happens without a separate server the user has to manage. The repository describes it as an AI Voice Imitation System, and the primary language is TypeScript, with Vue, Electron and Vite in the topic list. The engine itself lives in a separate repository, Aivis-Project/AivisSpeech-Engine.
The audience is narrower than the name suggests. The README states that the Engine supports Japanese speech synthesis only, in the same way as VOICEVOX ENGINE. AIVM metadata can describe multilingual speakers, but the Engine will not synthesise English or Chinese even when the underlying model was trained for it. If you need English output, this project is the wrong layer.
What you get instead is a desktop application with a graphical editor, model management from the UI, and a bundled engine. The README points end users at the official site and at three documents under public/ (howtouse.md, qAndA.md, contact.md) rather than at the repository README, which is written for developers.
How the AivisSpeech Engine and AIVMX files fit together
The data flow is straightforward. The editor talks to the AivisSpeech Engine over HTTP on 127.0.0.1 port 10101, which is the host value shown in the VITE_DEFAULT_ENGINE_INFOS example in the README. The Engine loads speech synthesis models from a per-OS Models directory and exposes an API described by openapi.json in the repository root.
The model format is the interesting part. AIVM is an open file format that bundles a trained model, hyperparameters, style vectors and speaker metadata (name, description, licence, icon, voice samples) into one file. An AIVM file is Safetensors with AIVM metadata added; an AIVMX file is ONNX with the same metadata added. The Engine is a reference implementation of the AIVM specification but deliberately supports only AIVMX. The README gives the reason: dropping PyTorch shrinks the install size and lets ONNX Runtime handle CPU inference.
That choice has a cost. AIVM files, which are the more portable form for tooling that works with PyTorch, cannot be loaded directly. You convert first. The README points at AIVM Generator, a web tool for producing AIVM or AIVMX files from existing models and for editing metadata on existing files.
Supported model architectures are Style-Bert-VITS2 and Style-Bert-VITS2 (JP-Extra). Nothing else is listed.
Installing AivisSpeech and generating your first line of Japanese speech
End users do not build this from source. The README directs them to the download page at aivis-project.com/speech, and the Engine has its own releases page. The UI can add speech synthesis models, and the README recommends that route for end users over placing files by hand.
If you do place files manually, the location depends on the OS. The README lists these three paths, and notes that the Engine prints the real one at startup as Models directory:.
# Windows
C:\Users\(username)\AppData\Roaming\AivisSpeech-Engine\Models
# macOS
~/Library/Application Support/AivisSpeech-Engine/Models
# Linux
~/.local/share/AivisSpeech-Engine/ModelsDevelopers building the editor from source need Node.js 22.11.0, per the README, and the steps differ from upstream VOICEVOX. The dependency install and environment file copy are the first two commands.
npm ci
cp .env.development .envThe README notes that the copied .env does not need editing. On macOS, .env.production does: the executionFilePath value must change from "AivisSpeech-Engine/run.exe" to "../Resources/AivisSpeech-Engine/run". That setting is used when launching a product build created by npm run electron:build.
The Engine has to be running before the editor. Its development environment is built separately, and the README shows the command as poetry run task serve inside the AivisSpeech-Engine directory. Once both are up, npm run electron:serve starts the Electron build of the editor, and npm run browser:serve starts the browser build. Note that package.json pins node to ">=22.14.0 <23" and declares [email protected] as the package manager, while the README's setup instructions use npm ci; that mismatch is worth knowing before you script a CI job.
Where AivisSpeech breaks, and where it is the wrong tool
The development policy section is unusually blunt, and it describes the maintenance model rather than a bug. The project keeps changes to VOICEVOX minimal so that tracking upstream stays cheap. Refactoring is not done, because conflicts with VOICEVOX are expected and the maintainers state they are not deeply familiar with the whole codebase. Features AivisSpeech does not use, such as singing synthesis, are not deleted; they are commented out. Documentation is not updated, and the README says plainly that the documents under docs/ are inherited from VOICEVOX without modification and that their contents may not apply. Test code is not updated either, and the README states that E2E tests in particular do not work properly because the UI has changed substantially.
Practically, that means you should not expect the docs/ tree to describe AivisSpeech behaviour, and you should not treat the test suite as a signal that a given path works. It also means upstream VOICEVOX changes arrive quickly but are not re-verified here.
Platform support is another boundary. Only Windows and macOS are listed. Windows 10 must be 22H2 or later; the README warns that older and LTSC builds of Windows 10 have been reported to crash the Engine on startup. On macOS, Intel machines are not actively verified, and the README recommends Apple Silicon. The application needs at least 1.5GB of free RAM.
Finally, the language limit is a hard stop for some users. Even a model with multilingual speakers defined in its AIVM metadata will only produce Japanese through this Engine.
AivisSpeech compared with VOICEVOX and with hosted TTS APIs
The closest comparison is VOICEVOX itself, and the README frames the relationship directly: AivisSpeech is based on the VOICEVOX editor UI and follows the latest VOICEVOX release with minimal modification. The difference is the engine and the model format. VOICEVOX is the upstream project; AivisSpeech substitutes AivisSpeech Engine and the AIVM/AIVMX model format, and adds UI for adding models. If you already run VOICEVOX, the reason to switch is the AIVMX model ecosystem, not the editor, which is largely the same code.
Hosted APIs such as ElevenLabs or CoeFont are a different trade entirely. They require no local model files, no 1.5GB of free RAM and no ONNX Runtime, and they are not limited to Japanese. In exchange, synthesis happens on someone else's infrastructure, per-use costs apply, and the voice set is whatever the provider offers. AivisSpeech runs the model on your own CPU and lets you supply any AIVMX file you have the rights to use. The README does not compare AivisSpeech to these services, so treat the choice as one of deployment model rather than quality.
For Japanese-only, offline, self-hosted synthesis with Style-Bert-VITS2 voices, AivisSpeech is a reasonable fit. For multilingual output or zero local setup, it is not.
Licence and the real cost of keeping this fork current
AivisSpeech inherits only the LGPL-3.0 half of VOICEVOX's dual licence, and the repository states this explicitly. The LICENSE file at the root carries the identifier. LGPL-3.0 matters if you redistribute a modified build: the usual obligations around source availability and relinking apply. This is a description of what the repository says, not legal advice; check the licence text and your own distribution model.
Model files are a separate question. AIVM metadata includes a licence field for each speaker, so the terms attached to a voice come from the model, not from AivisSpeech. The README does not state what licence the bundled or default voices carry, so verify that per model before publishing audio.
Upgrade cost is the more interesting number. Because the project avoids refactoring and keeps edits minimal, rebasing onto a new VOICEVOX release should be small. But the README also says documentation and tests are not maintained, so the verification work after each rebase falls on you. Recent releases listed for the repository are development and preview builds (1.1.0-dev, 1.1.0-preview.3, 1.1.0-preview.4), and the last push to the repository was on 2026-07-12. There is no stable 1.1.0 release listed, so anyone deploying this should decide whether a preview build is acceptable.
Editorial conclusion
AivisSpeech suits Japanese-speaking users on Windows 10 22H2 or later, Windows 11, or macOS 13 Ventura or later who want a desktop editor and are happy to run Style-Bert-VITS2 models in AIVMX form; the README states the Engine supports Japanese synthesis only, so anyone needing English or Chinese output should look elsewhere. Before committing, confirm that your machine has at least 1.5GB of free RAM, that your CPU is Apple Silicon if you are on a Mac, and that the AIVMX models you want actually exist, since the Engine deliberately refuses plain AIVM files and depends on ONNX Runtime for CPU inference.
Frequently asked questions
Which is the best AI for speech?
The README does not rank speech AI systems, so this cannot be answered from what the repository documents. It does state that AivisSpeech supports Japanese speech synthesis only, using Style-Bert-VITS2 and Style-Bert-VITS2 (JP-Extra) models in AIVMX form.
Does speech to text count as AI?
The repository does not address speech to text at all. AivisSpeech is a text-to-speech system: it takes text and produces Japanese speech through the AivisSpeech Engine.
Which text-to-speech API is the best?
The README does not compare text-to-speech APIs. It documents that the AivisSpeech editor talks to the AivisSpeech Engine over HTTP at 127.0.0.1:10101, with openapi.json in the repository root describing that interface.
What is speech AI used for?
The README gives one concrete use: generating Japanese speech from text in a desktop editor, with speech synthesis models added from the UI or placed in the Engine's Models directory.
How does AivisSpeech compare with VOICEVOX?
AivisSpeech is based on the VOICEVOX editor UI and follows the latest VOICEVOX release with minimal modification, but it substitutes the AivisSpeech Engine and the AIVM/AIVMX model format. The editor is largely the same code; the engine and models are what differ.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/aivis-project-aivisspeech)