# Irodori-TTS: the watermarking library installs from a fork at a pinned commit, the codec installs from an unpinned branch, and every documented command carries a flag you must not drop

> Irodori-TTS is a Japanese text to speech model built on flow matching, with zero-shot voice cloning from a reference clip, a caption channel for style, and an unusual control surface where emoji in the input text change delivery. The repository ships training and inference code for four hardware backends behind mutually exclusive extras. What its own dependency files and command examples reveal is the more useful part of the read.

**Aratako/Irodori-TTS** — A Flow Matching-based Text-to-Speech Model with Emoji-driven Style Control

- Repository: https://github.com/Aratako/Irodori-TTS
- Website: https://huggingface.co/collections/Aratako/irodori-tts
- Stars: 1,382 · Forks: 169
- Language: Python
- License: MIT
- Published: 2026-09-16 · Updated: 2026-09-16 · Language: en
- Canonical page: https://hysenlabs.com/projects/aratako-irodori-tts

## Watermarking is installed from a fork at a pinned commit

One line in the dependency list deserves to be read on its own:

```
silentcipher @ git+https://github.com/SesameAILabs/silentcipher.git@d46d7d0893a583d8968ab3a6626e2289faec9152
```

The feature that needs this library is watermarking, and the readme links to the original project under a different organisation entirely. So the code that watermarks your output comes from a fork hosted by someone else, at a commit hash rather than a release tag. Pinning a commit is the safer of the two approaches, since a moving branch reference can change under you and a tag can be moved, while a hash cannot drift without the install failing loudly. But the choice of source is the thing to think about: the fork is not the upstream, nothing in the visible documentation says why, and the two are not the same project. If you publish audio from this model, that is a question to have answered before you do.

## The audio codec is installed from an unpinned branch

The other git dependency is the one everything else stands on:

```
dacvae @ git+https://github.com/facebookresearch/dacvae
```

No commit hash, no tag, no version constraint at all. The generation target for this model is continuous latents from that codec, and the readme names a specific 32 dimensional codec for 48 kHz waveform reconstruction, so a change in the codec's output space is a change in what the model was trained to produce. Every other dependency in the file is a versioned release, several of them carefully bounded on both sides. This one follows whatever the branch tip is on the day you install. It is the most likely source of an environment that worked last month and does not work today, and unlike the watermarking library there is not even a hash to tell you what you had. The same line appears in both dependency files, so there is no alternate path that avoids it.

## Every documented command carries a flag that stops the environment being rebuilt

Read the commands rather than the arguments. Every single invocation in the file begins with a no-sync flag on the environment runner:

```bash
uv run --no-sync python infer.py \
  --hf-checkpoint Aratako/Irodori-TTS-v4.1-Small \
  --text "こんにちは、私はAIです。これは音声合成のテストです。" \
  --ref-wav path/to/reference.wav \
  --output-wav outputs/sample.wav
```

The install section says why: after syncing with a backend extra, you should use the no-sync form so the environment is not re-synced without the selected extra. In other words the default behaviour of the tool you use to run things is destructive here. Drop the flag and the environment resolves again from the base dependency list, which does not carry a backend extra at all, and the carefully chosen build of the framework for your accelerator is replaced by whatever the resolver picks. That is a very sharp edge for a quick start, and it is a strong argument for wrapping the documented commands in a script or a make target the first time you set this up.

## The AMD extra exists because of one missing import path

The four extras are mutually exclusive and each one selects a different package index: a CUDA index for one accelerator, a ROCm index on Linux for another, an XPU index on Linux and Windows for a third, and a CPU index that falls back to the ordinary wheels on macOS. What makes the AMD one interesting is the note attached to it. The extra includes a second Triton package for that platform, and the reason given is that the first one on its own does not provide a particular submodule that a specific import path between two libraries needs. Then the closing sentence, which is the kind of honesty worth having in a readme: this was validated with AMD GPU inference. So the extra exists to bridge a gap in someone else's import chain, and the validation covers one vendor's hardware. Anyone on a different AMD card is extrapolating, and the readme says so rather than implying full coverage.

## The Intel extra restates its own platform in every single marker

Compare the platform conditions across two extras and a pattern appears. The AMD extra gates its two Triton packages with a Linux condition. The Intel extra gates five packages, including the framework itself, with a condition that says Linux or Windows, on every line:

```
"torch>=2.10.0,<2.11.0; sys_platform == 'linux' or sys_platform == 'win32'"
```

The documentation already says this extra is for Linux and Windows, so the markers restate what the extra's own selection implies. That is harmless, and it is also how a package ends up silently installing nothing on a platform where the extra is offered, because an unmet marker is a skip rather than an error. Note also that the Intel extra pins its Triton package exactly while the AMD extra's equivalents use floors, so two of the four backends have different tolerance for the same class of dependency. And all four extras bound the framework tightly, which is what makes a single set of backend extras the right shape here at all.

## First generation checkpoints are declared incompatible with everything after them

The compatibility statement is four bullets and one of them is a warning. The default branch tracks the current generation and its point one release, including a new sampling method and a larger model that has not shipped. The current code stays backward compatible with the released second and third generation checkpoints, both the plain and the style-conditioned families. Earlier codebase states are available through three named tags. And then: first generation checkpoints and their preprocessing are not compatible with anything that came after. Since the repository publishes no releases at all, those three tags are the only version anchors a user has, which makes the fourth bullet operationally important rather than academic. If you have audio produced by the first generation, that data is not re-synthesisable with this code, and the preprocessing that produced it is gone from the default branch.

## Both web interfaces are documented bound to every network interface

Two Gradio applications ship with the repository, one for reference audio cloning and speaker inversion on port 7860, and one for caption plus reference conditioning on port 7861. Both are launched the same way:

```bash
uv run --no-sync python gradio_app.py --server-name 0.0.0.0 --server-port 7860
```

The address is every interface, not loopback, so the documented command exposes a speech synthesis interface, with a file browser for reference audio and a player for results, to the whole local network. That is sometimes exactly what you want on a bench machine and never what you want on a laptop tethered to a cafe. Nothing in the visible documentation mentions an access token, a password or a firewall, and the readme's only network advice is that a separate repository exists for an OpenAI compatible inference server. For a development interface that is a fair default, and it is worth narrowing to the loopback address as your first edit rather than your fifth.

## Watermarking happens when available, and two files declare the same dependencies

Two loose ends. The first is the feature list's phrasing on watermarking: generated audio is watermarked with that library when available. The word when available is doing real work, because the library is a hard dependency in both dependency files, so the condition is not about installation. Nobody has written down what makes it unavailable at runtime, which means there is a documented path where your output is not watermarked and no log line announcing it. If you are relying on the watermark for provenance, that is the line to interrogate. The second is dependency duplication. A plain requirements file sits beside the packaging metadata and lists the same set, including both git references, in a slightly different order, and it omits one package that the extras introduce. Two files that must agree, one of which is what a manual install would use, is how an environment ends up one package behind the lock file.

## Conclusion

Irodori-TTS fits someone building Japanese speech synthesis who wants to fine tune a released checkpoint with low rank adapters, clone a voice from a short clip, and control delivery with something other than a prompt. It does not fit someone who needs a reproducible build without reading the dependency files. Four things to check before the first install. Which backend extra you pick, because they are mutually exclusive and each one resolves from a different package index, so choosing the wrong one is not a small mistake. That every command carries the no-sync flag, because without it the environment re-resolves and can quietly replace the backend you selected. Where your dependencies come from, because the audio codec and the watermarking library both install from git rather than an index, one of them from a fork, and one of them without a version reference at all. And which checkpoint generation you are on, since the first generation is declared incompatible with everything after it.

## FAQ

### How do I install Irodori-TTS?

Clone the repository, enter it, then sync the environment with exactly one backend extra. The extras are mutually exclusive: one for NVIDIA CUDA 12.8 on Linux or Windows, one for AMD on Linux or WSL, one for Intel on Linux or Windows, and one for CPU only or macOS.

### Which GPU backends does Irodori-TTS support?

Four, each resolving from its own package index: a CUDA 12.8 index, a ROCm index on Linux, an XPU index on Linux and Windows, and a CPU index that falls back to the standard PyPI wheels on macOS. The AMD extra adds a second Triton package because the first does not provide a submodule an import path needs, and that path was validated with AMD hardware.

### Does Irodori-TTS support zero-shot voice cloning?

Yes, from reference audio. One or more reference clips can be concatenated up to the checkpoint's 120-second limit, local checkpoint files are accepted as well as Hub checkpoints, and inference also runs without any reference audio using a flag.

### What does emoji control do in Irodori-TTS?

Emoji annotations in the input text can influence delivery and non-verbal vocal expressions, and the readme limits that to supported checkpoints. Output length is estimated automatically, so there is no manual duration argument to pass.

### Are Irodori-TTS checkpoints compatible across versions?

The default branch stays backward compatible with the released second and third generation base and style-conditioned checkpoints, earlier codebase states are available on three named tags, and first generation checkpoints and their preprocessing are explicitly not compatible with anything after them.

### Is audio generated by Irodori-TTS watermarked?

The feature list says generated audio is watermarked with the SilentCipher library when available, so the step is conditional on that library being usable at runtime. The library is installed from a fork at a pinned commit hash rather than from the upstream project it is documented against.

## Sources

- [Aratako/Irodori-TTS on GitHub](https://github.com/Aratako/Irodori-TTS)
- [Issues](https://github.com/Aratako/Irodori-TTS/issues)
- [License: MIT](https://github.com/Aratako/Irodori-TTS/blob/main/LICENSE)
- [Project website](https://huggingface.co/collections/Aratako/irodori-tts)
- [README](https://github.com/Aratako/Irodori-TTS/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/aratako-irodori-tts
