# The codec you install in T5Gemma-TTS depends on which language you speak

> T5Gemma-TTS is training and inference code for a multilingual text-to-speech model built on an encoder-decoder LLM, supporting English, Chinese, and Japanese with zero-shot voice cloning and explicit duration control. The fiddly parts are the audio codec swap, three separate inference scripts, and a requirements file that disagrees with the changelog.

**Aratako/T5Gemma-TTS** — Multilingual TTS model with voice cloning and duration control, based on T5Gemma encoder-decoder LLM

- Repository: https://github.com/Aratako/T5Gemma-TTS
- Website: https://huggingface.co/Aratako/T5Gemma-TTS-2b-2b
- Stars: 312 · Forks: 28
- Language: Python
- License: MIT
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/aratako-t5gemma-tts

## The codec swap is decided by target language, not by preference

The audio decoder is chosen per language, and the two options need different libraries.

By default a variant of XCodec2 is used for decoding, described as a model better suited to Japanese voices. For English and Chinese voices the documentation recommends the original XCodec2 model instead, and that one has an install requirement of its own: the original xcodec2 library must be installed at a pinned version with dependencies suppressed, because it must be the original library rather than the variant.

```bash
# You must use the original xcodec2 library when using the original XCodec2 model
pip install xcodec2==0.1.5 --no-deps
```

The requirements file resolves this for you in one direction. It installs the variant by pointing pip at a source archive hosted on the model hub rather than naming a package, so a fresh install gets the Japanese-leaning decoder no matter what you intend to generate.

That is the most common source of a bad first run. Generating English or Chinese with the default decoder is not an error, but it is the configuration the project recommends against, and switching means replacing the installed codec at a specific version while suppressing dependency resolution.

The decoder path is also what the low-memory switches move to CPU, which is why codec choice and memory choice are the same conversation.

## Three inference scripts because there are three model layouts

There is no single entry point. Three scripts cover three ways the model might exist on disk.

One takes a model directory in the HuggingFace format and is the path used in every quick start example. It accepts an output directory so generated audio lands where you ask. Duration control is a flag on this script, taking a target duration in seconds.

```bash
# Specify target duration in seconds
python inference_commandline_hf.py \
    --model_dir Aratako/T5Gemma-TTS-2b-2b \
    --target_text "Hello, this is a test of the text to speech system." \
    --target_duration 5.0
```

The second script works against local checkpoints instead. It takes a model root and a model name, with `trained` for a full checkpoint and `lora` for a LoRA checkpoint, so fine-tuning output is inferred through the same flag.

The third script starts a Gradio interface on a chosen port, and it is the one that carries the memory switches and the batch controls.

Voice cloning rides on the first script, adding a reference text and a reference speech path. Duration has a fallback: if the target duration is not given, the system calculates one from phoneme count and language-specific pacing rules. The documentation calls that calculation approximate, which is an argument for always passing the duration yourself.

## Only the encoder is quantized, and the size table shows the trade

Pre-quantized models are published in three forms, and the differences are small enough to be worth understanding.

The full precision model runs in bfloat16 at roughly 10.6 gigabytes of VRAM. The 8-bit encoder variant needs roughly 8.6, and the 4-bit encoder variant roughly 7.6. So going from full precision to the smallest quantized model saves about three gigabytes, which is roughly a quarter of the requirement.

The asymmetry is deliberate and stated: only the encoder is quantized, and the decoder stays in full precision to maintain audio quality. So the quantization does not touch the part of the system that produces the waveform.

There is a hard prerequisite attached to the quantized models: they require a specific quantization library and a CUDA GPU. That excludes the Apple Silicon path entirely, even though MPS is otherwise listed as a tested platform.

Usage is otherwise identical to the full model. You change the model name and nothing else.

```bash
python inference_gradio.py \
    --model_dir Aratako/T5Gemma-TTS-2b-2b-encoder-8bit \
    --low_vram \
    --port 7860
```

Combining the quantized model with the low-memory preset is the configuration that appears in the documentation, which suggests the two were introduced with the same audience in mind.

## Three switches trade GPU memory for first-run latency

The memory controls are named for what they move, and each has a stated saving and a stated cost.

A flag moves the XCodec2 tokenizer to the CPU, reducing VRAM use by roughly 3.5 gigabytes at the cost of slower audio encoding and decoding. A second moves Whisper, on the auto-transcribe path, to the CPU, saving roughly 5 gigabytes while slowing transcription. A preset enables both of those and additionally disables the graph compilation step.

The framing is important: the documentation states that these switches do not change model quality and only trade GPU memory for a bit of latency on first runs. That is a stronger guarantee than most memory options offer, because it separates the output from the resource.

The disabling of compilation in the preset is the part worth noticing, because compilation is usually a startup-time cost that pays for itself across runs. Combining both flags throws that away, which is consistent with the saving being described as applying to first runs.

Whisper appearing at all is a clue about the pipeline: there is an automatic transcription path, so audio can be turned into text somewhere in the workflow rather than only text into audio.

## Batch generation runs the encoder once and the decoder many times

Batch generation produces multiple audio variations from the same input in parallel, using different random samples, and the efficiency argument is stated precisely: the encoder runs only once, so only the decoder runs in batch.

That is the right place to parallelize, since in an encoder-decoder text-to-speech system the conditioning stage is a single pass while the decoder is the part that varies per sample. Running the encoder per sample would multiply the work for no benefit.

The control is a slider in the interface, spanning 1 to 256 samples, and results come back in an eight-column grid. The stated use is generating several candidates and picking the best one, which is a sampling-and-select workflow rather than a throughput claim.

The upper bound of 256 is a memory decision in disguise. Each sample occupies decoder state, so the largest settings are for a machine with room, and the grid layout means the interface is designed for comparison rather than for downloading a batch.

This is also why the batch feature is listed under the December update alongside quantization and the low-memory option: all three are about what fits on the device rather than about what the model can say.

## The requirements cap torch at 2.8.0 while the changelog says 2.9+ works

The dependency file and the update log disagree about the PyTorch ceiling, and that is the first thing to check in your own environment.

The requirements file caps both torch and torchaudio with a range that starts at 2.5.1 and ends at 2.8.0. The update log for December lists compatibility with PyTorch 2.9 and newer. The README's GPU install note separately pins a ceiling of 2.8.0 when installing from a wheel index.

So the strongest signal, the file a resolver actually reads, excludes the version the log claims to support. If you install from the requirements file you get 2.8.0 or lower; if you force 2.9 you are outside what the file declares.

Two more inconsistencies sit nearby. The comment at the top of the requirements file says to install torch and torchaudio with the CUDA 12.4 wheels, while the README's command points at a CUDA 12.8 index and the container defaults to CUDA 12.8. And the file installs the codec by direct URL rather than by package name, which means that dependency is unpinned in the ordinary sense.

Elsewhere the file is unusually tight: a text transformer library is pinned to an exact version, a config library is pinned exactly, a datasets library is capped below a major version, and the audio language packages are pinned to a specific variant release.

## Docker selects one of four base images from a build argument

The container setup exists mainly for Windows users, and it works by parameterizing the CUDA version rather than the base image.

The build declares a CUDA version argument with a default of 12.8, and four candidate base images are defined: a PyTorch 2.8 runtime on CUDA 12.8, and PyTorch 2.5.1 runtimes on CUDA 12.4, CUDA 12.1, and CUDA 11.8. The final stage selects one by interpolating the argument into the stage name, so the value has to match one of the four or the build fails.

A Python version argument is also declared, defaulting to 3.12, but the interpreter comes from the selected PyTorch image rather than from it, so that argument does not change anything.

The compose file adds three things the plain image does not have: a named volume for the model cache so weights are not re-downloaded, a token variable for gated hub access, and an extra-arguments passthrough on the command line. It also requests every available GPU through the device reservation block rather than a fixed count.

The Windows framing is explicit. Native Windows may show inconsistent generation times or occasional hangs, the root cause is stated as still under investigation, and WSL2 or Docker is offered as the workaround.

## Conclusion

T5Gemma-TTS suits a researcher who needs Japanese or Chinese speech with cloned voices and enough control over pacing to match a script, and who can afford roughly 8 to 11 gigabytes of VRAM or the patience to offload to CPU. It does not suit a Windows user running natively, since the documentation reports hangs with no root cause yet, or anyone expecting the dependency set to match the changelog. Before you build, pick your codec deliberately, decide whether the encoder-only quantization is enough, and check the torch version your environment will actually resolve to.

## FAQ

### Which languages does T5Gemma-TTS support?

English, Chinese, and Japanese. The default decoder is a variant of XCodec2 used for better Japanese voice support, and the documentation recommends the original XCodec2 model for English and Chinese voices, which requires installing the original xcodec2 library at a pinned version with dependencies suppressed.

### How do I control the length of audio generated by T5Gemma-TTS?

Pass a target duration in seconds to the inference script. If it is not specified, the system calculates an appropriate duration from phoneme count and language-specific pacing rules, and that calculation is described as approximate, so specifying the duration manually is recommended when the result is not what you expect.

### Does T5Gemma-TTS need a GPU and how much VRAM?

The full precision bfloat16 model needs roughly 10.6 GB, the 8-bit encoder variant roughly 8.6 GB, and the 4-bit encoder variant roughly 7.6 GB. The quantized variants require bitsandbytes and a CUDA GPU. Apple Silicon through MPS is listed as a tested platform.

### How do I reduce VRAM usage when running T5Gemma-TTS?

Three options, documented as not changing model quality. Moving the XCodec2 tokenizer to the CPU saves roughly 3.5 GB, moving Whisper to the CPU saves roughly 5 GB, and the low-VRAM preset enables both and disables torch.compile, trading GPU memory for latency on first runs.

### Why is Docker recommended for Windows users of T5Gemma-TTS?

On some native Windows environments, inference has shown inconsistent generation times and occasional hangs, observed in the developer's testing with the root cause still under investigation. WSL2 or Docker is offered as a workaround, and the compose route is the recommended one. Tested environments are Linux with CUDA, Windows with CUDA inside Docker, and Apple Silicon.

## Sources

- [Aratako/T5Gemma-TTS on GitHub](https://github.com/Aratako/T5Gemma-TTS)
- [Issues](https://github.com/Aratako/T5Gemma-TTS/issues)
- [License: MIT](https://github.com/Aratako/T5Gemma-TTS/blob/main/LICENSE)
- [Project website](https://huggingface.co/Aratako/T5Gemma-TTS-2b-2b)
- [README](https://github.com/Aratako/T5Gemma-TTS/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/aratako-t5gemma-tts
