# MisoTTS 8B is a local inference path, not the 110 ms API

> An English-only 8.2B parameter text-to-speech model from Miso Labs, shipped as inference code plus pinned dependencies. The headline latency number belongs to the hosted service, and the hardware floor is a 24 GB card.

**MisoLabsAI/MisoTTS** — Miso TTS is an 8 billion, highly emotive text-to-speech model

- Repository: https://github.com/MisoLabsAI/MisoTTS
- Stars: 3,239 · Forks: 325
- Language: Python
- License: NOASSERTION
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/misolabsai-misotts

## Two install routes, and both end at the same demo script

The quickstart assumes a Python environment manager rather than a plain virtualenv. If `uv` is not installed, the first instruction pipes an installer straight into a shell:

```bash
curl -LsSf https://astral.sh/uv/install.sh | sh
```

Then the repository is cloned and pinned to one interpreter version:

```bash
git clone https://github.com/MisoLabsAI/MisoTTS.git
cd MisoTTS
uv sync --python 3.10
source .venv/bin/activate
```

`pyproject.toml` states `requires-python = ">=3.10,<3.13"`, so 3.10 is a floor rather than a preference and the ceiling at 3.12 comes from the same constraint. The first run is one command, and it downloads the public weights from the Hugging Face repository `MisoLabs/MisoTTS` if they are not already cached:

```bash
uv run python run_misotts.py
```

The script writes `full_conversation.wav` into the repository root. A `pip` route is offered for people without `uv`, and it differs only in how the environment is built: `python3.10 -m venv .venv`, then `pip install -e .`, then `python run_misotts.py`. Both routes converge on the same entry script, so the choice of package manager buys nothing at inference time and the real cost sits in the download.

## The 110 ms figure in circulation belongs to a hosted API

The number most often quoted about this model is one the README explicitly disowns. A note in the system requirements section says that public discussions of `110 ms` latency refer to the hosted production API's time to first byte on H100-class hardware, and not to the unoptimized local inference path in this repository. It adds that startup and generation latency on consumer or workstation GPUs will be materially slower. That is an unusually direct retraction, and it sits directly above a table recommending consumer cards such as the RTX 3090 and the L4. Anyone benchmarking against the marketing figure rather than against their own card will conclude the local path is broken, when the accurate statement is only that it is slower than a served endpoint. The repository publishes no local latency measurement at all, only a pointer to somebody else's.

## An 8B backbone and a 300M decoder split the audio work

The architecture is a text-to-dialogue RVQ Transformer described as inspired by the Sesame CSM design, and it splits into two transformers. A large backbone consumes text and audio frame embeddings. A smaller decoder then predicts the higher-order audio codebooks inside each frame, autoregressively. The summary table puts the backbone at `llama-8B` and the audio decoder at `llama-300M`, for roughly 8.2B parameters counted across backbone, decoder, embeddings and heads. Audio is represented as Mimi codes drawn from 32 codebooks over a vocabulary of 2,051, against a text vocabulary of 128,256 and a maximum sequence length of 2,048. The backbone takes interleaved text and audio tokens, which is what lets a generation be conditioned on the conversation so far rather than on a single utterance:

```python
generator = load_miso_8b(
    device=device,
    model_path_or_repo_id="MisoLabs/MisoTTS",
)

audio = generator.generate(
    text="Hello from Miso.",
    speaker=0,
    context=[],
    max_audio_length_ms=10_000,
)
```

Note that the quickstart passes an empty list for `context`, so the conditioning path is available but unused by the example the project tells you to run first.

## Prompted generation is one Segment, and the example stops mid-argument

Voice conditioning is a list of `Segment` objects, each pairing a speaker number with a transcript and a waveform. You load a file with `torchaudio.load("prompt.wav")`, resample it to `generator.sample_rate`, and hand the result to a segment:

```python
context = [
    Segment(
        speaker=0,
        text="This is the transcript for the prompt audio.",
        audio=prompt_audio,
    )
]
```

The transcript is mandatory even though the audio is what carries the voice, so the caller has to supply text that matches the recording. The one flaw in this section is the last line of its own snippet, which ends at `max_audio_length_ms=10`, halfway through a value and with no closing parenthesis. The earlier example uses `10_000`, so the two readings differ by three orders of magnitude and the snippet that exists specifically to show prompted generation is the one that cannot be copied as printed.

## Weights arrive from two different silentcipher sources

The watermarking component has two identities in this project and the README never reconciles them. It says the first run downloads the SilentCipher watermarking model from `sony/silentcipher`, which is a Hugging Face model repository. The dependency list, in both `pyproject.toml` and `requirements.txt`, installs a Python package from a different organization entirely: `silentcipher @ git+https://github.com/SesameAILabs/silentcipher@d46d7d0893a583d8968ab3a6626e2289faec9152`. So a git checkout from SesameAILabs at one fixed commit, and a model download from a Sony-named repository, are both required, and they are not the same artifact. Two operational notes come with them. If the SilentCipher download times out, the instruction is to rerun the command, because the Hugging Face cache resumes from files that already completed. And the section that introduces the weights contains no download command at all: its only code block repeats the quickstart's `uv run python run_misotts.py`, so there is no supported way to prefetch the checkpoint on its own.

## Ten exact pins, written down in two files

Nothing in the dependency list is a range. `torch==2.4.0` and `torchaudio==2.4.0` are exact, as are `tokenizers==0.21.0`, `transformers==4.49.0`, `huggingface_hub==0.28.1`, `moshi==0.2.2`, `torchtune==0.4.0` and `torchao==0.9.0`, plus the git-pinned silentcipher commit. That is deliberate for a project vendoring an inference path someone else has to reproduce, but it puts the upgrade cost on your side: a patched security release or a driver that needs a newer torch has to be resolved as a code change against `pyproject.toml`, and the same list is maintained twice, once in `requirements.txt` and once in the `dependencies` array, with `uv.lock` sitting beside them as a third record. The packaging metadata names exactly five top level modules, `generator`, `models`, `moshi_compat`, `run_misotts` and `watermarking`, which matches the tree exactly. The presence of `moshi_compat.py` next to a pinned `moshi==0.2.2` dependency suggests the compatibility layer is where the Sesame-style interface is adapted, though the README does not describe that module.

## 24 GB is the bf16 floor and the CPU path has no number attached

The hardware table gives two rows. At `bfloat16` or `fp16` the weights take about 16 GB and the recommended VRAM is 24 GB, with an RTX 3090, RTX 4090, A5000 or L4 named as examples. At `float32` the weights reach about 33 GB and the recommendation is 40 GB or more, so an A100, A6000 or H100. GPU inference defaults to `torch.bfloat16`, and cards with 4 to 16 GB are called out as insufficient. Disk is the other constraint: the first run pulls 30 to 40 GB into the Hugging Face cache, covering the checkpoint, the Mimi codec, the SilentCipher watermarker and the Llama 3.2 tokenizer. The CPU row is where the documentation thins out. It states that inference runs but is slow, and asks for at least about 20 GB of RAM at bfloat16 and about 40 GB at float32, but attaches no time to first byte and no tokens per second, unlike the hosted figure it dismisses elsewhere. The relationship between the 2,048 maximum sequence length and the `max_audio_length_ms=10_000` in the example is also left unstated.

## Version 0.1.0, no releases, and a license that points at a file

There is nothing to download as a release: the repository publishes no GitHub releases at all. The version lives in one place, `version = "0.1.0"` in `pyproject.toml`, and the last push to the repository was on 2026-06-09. The license is the piece that needs a decision from you rather than from the project. The repository metadata carries no asserted license, while `pyproject.toml` declares `license = { file = "LICENSE" }` and a LICENSE file does sit at the root of the tree. Those two statements are not the same claim, and the README names no license text at all, so this article will not tell you which terms apply. Check the LICENSE file yourself before any use that matters. The safety section is short and unambiguous: no impersonation, no deceptive audio, no fraud, no harmful content, and anyone embedding this model in another application is told to supply their own private watermark key and keep it secret.

## Conclusion

MisoTTS 8B fits teams building English-only dialogue or voice-conditioned speech that already have a 24 GB GPU and roughly 40 GB of free disk. It does not fit a laptop, a 4 to 16 GB consumer card, or any requirement outside English, since the model states English only. Verify three things before building on it: that the SilentCipher watermarking download completes on your network, that the exact `torch==2.4.0` pin installs against your CUDA build, and that you hold a private watermark key before deploying anything this model generates.

## FAQ

### What languages does MisoTTS 8B support?

English only. Both the model summary table and a separate language support note state this without qualification, so a multilingual text-to-speech requirement is out of scope for this checkpoint.

### How much VRAM does MisoTTS 8B need?

24 GB at bfloat16 or fp16, where the weights are about 16 GB, and 40 GB or more at float32, where they reach about 33 GB. GPU inference defaults to torch.bfloat16, and cards with 4 to 16 GB are stated to be insufficient.

### How do I run MisoTTS on my own machine?

Clone the repository, run `uv sync --python 3.10`, activate `.venv`, then run `uv run python run_misotts.py`. The script writes `full_conversation.wav` to the repository root. Without uv you can use `python3.10 -m venv .venv`, `pip install -e .` and `python run_misotts.py` instead.

### Is audio generated by MisoTTS watermarked?

Yes, by default. The safety section states that generated audio is watermarked, and that anyone deploying the model inside another application should use their own private watermark key and keep it secret. The first run downloads the SilentCipher watermarking model from sony/silentcipher.

### Is there a MisoTTS release to download?

The repository has no GitHub releases. The version is recorded in pyproject.toml as 0.1.0, and the weights are fetched through the Hugging Face cache from the MisoLabs/MisoTTS model repository on the first run.

### Can MisoTTS reproduce a voice from a recording?

Yes, optionally. You load a wav file with torchaudio, resample it to generator.sample_rate, wrap it in a Segment carrying a speaker number and a transcript, and pass that list as context. The quickstart example runs with context set to an empty list.

## Sources

- [Issues](https://github.com/MisoLabsAI/MisoTTS/issues)
- [MisoLabsAI/MisoTTS on GitHub](https://github.com/MisoLabsAI/MisoTTS)
- [README](https://github.com/MisoLabsAI/MisoTTS/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/misolabsai-misotts
