# Fish Speech pins torch to 2.8.0, resolves a protobuf fight for you, and ships a research license with the weights

> The S2-Pro text-to-speech model is a 4B slow decoder paired with a 400M fast one over a 10-codebook RVQ codec. The engineering detail worth reading is the packaging: mutually exclusive CUDA extras, a hard protobuf override, and a license that is not the word open source implies.

**fishaudio/fish-speech** — GitHub describes it as SOTA Open Source TTS. The repository metadata lists Python as its primary language. The metadata lists the NOASSERTION license. This article stays within the project description and details documented in the GitHub repository README.

- Repository: https://github.com/fishaudio/fish-speech
- Website: https://speech.fish.audio
- Stars: 32,780 · Forks: 2,828
- Language: Python
- License: NOASSERTION
- Published: 2026-08-13 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/fishaudio-fish-speech

## The weights ship under a research license, and the README promises enforcement

The repository describes itself as an open source TTS project, but the terms are named rather than assumed. Both the README notice and the pyproject metadata point at the Fish Audio Research License for the codebase and its associated model weights, and the README adds that action will be taken against any violation. It also carries a legal disclaimer disclaiming responsibility for illegal usage and pointing readers at local DMCA and related law. The repository's own license field records NOASSERTION, so nothing in the metadata tells you the terms without opening LICENSE. The consequence is concrete: a team that reads the repository description and assumes MIT or Apache will find a different grant waiting for them, and that check belongs before you plan a product around the model, not after.

## There is no install command in the README, only five links into a hosted site

The Quick Start hands you to speech.fish.audio rather than to a shell. It points at Installation, Command Line Inference, WebUI Inference, Server Inference, and Docker Setup, and the one line aimed at automated agents says the same thing.

```
Install and configure Fish-Audio S2 by following the instructions here: https://speech.fish.audio/install/
```

For serving under SGLang or vLLM the README redirects you to those projects' own Fish Speech recipes instead of describing it here. The consequence is that the environment matrix lives outside the repository: the in-tree constraint is only that pyproject requires Python 3.10 or newer and pins torch and torchaudio at 2.8.0 with transformers held at 4.57.3 or lower. Anyone automating a setup has to fetch the hosted install page first, and a change to that page changes your instructions with no commit to review.

## The five extras are mutually exclusive, and torch is pinned in the base list too

Optional dependencies come in five flavors, stable, cpu, cu126, cu128, and cu129, and every one of them pins the same two packages.

```toml
[project.optional-dependencies]
stable = [
    "torch==2.8.0",
    "torchaudio",
]
```

```toml
conflicts = [
  [
    { extra = "cpu" },
    { extra = "cu126" },
    { extra = "cu128" },
    { extra = "cu129" },
  ],
]
```

That conflict declaration is the part that bites: asking for the CPU build together with any CUDA build is a resolver error, not a preference. Note also that torch==2.8.0 already appears in the base dependency list, so the extras are choosing a wheel index rather than a different model version. The consequence for planning is that your driver choice has to be settled before the first install, and a build that later needs a different CUDA release means a new environment, not an extra flag.

## A protobuf override exists because a transitive dependency holds the wrong floor

The packaging config carries an explicit override with the reason written next to it.

```toml
override-dependencies = [
    # descript-audiotools (transitive via descript-audio-codec) pins protobuf<3.20,
    # but fish-speech's generated proto code requires >=3.20; cap at <6 for API stability
    "protobuf>=3.20.0,<6.0.0",
]
```

So descript-audio-codec, which is in the base dependencies, pulls in a library that wants protobuf below 3.20, while the project's own generated protobuf code needs 3.20 or newer. The project resolves it by forcing the range and capping the ceiling at 6 for API stability. The consequence is that you are inheriting a compromise the maintainers already hit, and the ceiling is a deliberate bet about protobuf API drift rather than a compatibility guarantee. Anyone adding a package that wants protobuf outside 3.20 to 6.0 will meet the same wall and will need the same override in their own environment.

## Both compose services sit behind profiles, so a plain up starts nothing

The compose file declares the project as fish-speech and defines two services, both extending the app-base service in compose.base.yml. The webui service builds the webui target, passes a COMPILE flag defaulting to 0, and maps a configurable port onto 7860. The server service builds the server target and maps a configurable port onto 8080.

```yaml
  server:
    extends:
      file: compose.base.yml
      service: app-base
    build:
      target: server
    environment:
      COMPILE: ${COMPILE:-0}
    profiles: ["server"]
    ports:
      - "${API_PORT:-8080}:8080"
```

Because each service carries a profiles entry, neither starts unless you name its profile, and the two are separate long-running processes rather than one container doing both jobs. A separate compose.rocm.yml sits at the repository root for AMD hardware, and entrypoint.sh is the container entry. The consequence is that a first run that omits the profile looks like a broken install rather than an empty one.

## Dual-AR puts a 4B model on the time axis and a 400M model on every audio step

The architecture is a decoder-only transformer paired with an RVQ audio codec running 10 codebooks at roughly 21 Hz, split into a master and a slave half. The Slow AR component carries 4B parameters and operates along the time axis to predict the primary semantic codebook. The Fast AR component carries 400M parameters and generates the remaining 9 residual codebooks at each time step to reconstruct the acoustic detail. The published model table lists S2-Pro at 4B parameters as the full-featured flagship on HuggingFace. The consequence for anyone sizing hardware is that 4B is only the slow half: the fast half runs once per time step across the whole clip, so the work grows with the length of the audio you generate, and the README's own framing is that this asymmetry is what makes inference faster rather than what makes it free.

## Tag vocabulary is free-form, so nothing validates the instruction you typed

Prosody and emotion are steered by inline square-bracket tags written into the text, and the supported set is described as more than 15,000 unique tags that are not limited to fixed presets. The examples given are open-ended phrases such as [whisper in small voice], [professional broadcast tone], and [pitch up], alongside a named library of about three dozen entries covering [pause], [emphasis], [laughing], [sigh], [angry], [screaming], and similar. Control is described as sub-word level, and the model handles multi-speaker and multi-turn conversation. The consequence is that the tag language has no closed vocabulary to check against, so a description the model does not recognize is not an error you will see, it is a line read in whatever voice the model chose. The README does not document a fallback or a way to verify that a tag took effect.

## The multilingual claim covers 80 languages, the best-WER claim covers 11 of 24

The training claim is over 10 million hours of audio covering more than 80 languages, paired with reinforcement learning alignment through Group Relative Policy Optimization, where the same model suite is reused as reward models for data cleaning and annotation. The published benchmark table is the project's own: Seed-TTS Eval word error rate of 0.54% for Chinese and 0.99% for English, an Audio Turing Test posterior mean of 0.515, an EmergentTTS-Eval win rate of 81.88%, and instruction benchmark scores of 93.3% TAR and 4.51 out of 5.0 for quality. Comparisons in the same section put Qwen3-TTS at 0.77 and 1.24, MiniMax Speech-02 at 0.99 and 1.90, and Seed-TTS at 1.12 and 2.25. The catch is the multilingual row: on the 24-language MiniMax testset the model holds best WER in 11 languages and best similarity in 17. The consequence is that 80 languages of coverage and 11 languages of leading accuracy are different claims, and the tables here are the vendor's own, not an independent evaluation.

## Conclusion

Fish Speech makes sense if you want a self-hosted multilingual TTS whose prosody you steer with inline tags, and you are prepared to read the Fish Audio Research License rather than assume a permissive grant. It fits badly if you need the weights under a standard open source license, or if your GPU choice is still open, because the CUDA extras are declared mutually exclusive and torch is pinned to one version. Check the license text for your use, then pick exactly one extra matching your driver and start from the compose profiles, which do not launch by default.

## FAQ

### what is fish speech

Fish Speech is a Python text-to-speech project from Fish Audio, published as fish-speech version 2.0.0 and requiring Python 3.10 or newer. Its current model is S2-Pro, a 4B parameter model on HuggingFace that uses a Dual-Autoregressive architecture with a 400M fast component, an RVQ codec of 10 codebooks at about 21 Hz, and inline bracketed tags for emotion and prosody.

### is fish speech open source

The code and the model weights are released under the Fish Audio Research License, named in both the README notice and pyproject.toml, and the README states that action will be taken against violations of it. The repository's license field records NOASSERTION, so the terms are only in the LICENSE file itself. It also carries a legal disclaimer about illegal usage and DMCA.

### how to install fish speech

The README has no install command of its own. It sends you to https://speech.fish.audio/install/, with a Docker Setup section on the same page, and the repository ships compose.yml with a webui service on port 7860 and a server service on port 8080, both hidden behind profiles that you must name explicitly. In-tree requirements are Python 3.10 or newer and torch 2.8.0, with exactly one extra chosen from stable, cpu, cu126, cu128, or cu129.

### how to use fish speech

The project splits usage into four documented paths on speech.fish.audio: command line inference, WebUI inference, server inference, and Docker setup. If you want it behind an existing inference stack instead, the README points to the SGLang-Omni model README and to the vLLM-Omni Fish Speech S2 Pro recipe and text-to-speech user guide rather than documenting those routes itself.

### fish speech vs qwen3-tts

The only comparison the repository offers is its own benchmark table, which reports S2 at 0.54% Chinese and 0.99% English word error rate on Seed-TTS Eval against 0.77 and 1.24 for Qwen3-TTS, plus an Audio Turing Test posterior mean of 0.515. Those numbers come from the project, not from an independent evaluation, so treat them as a claim to check rather than a settled result.

## Sources

- [Official documentation](https://speech.fish.audio)
- [Official README](https://github.com/fishaudio/fish-speech#readme)
- [Project repository](https://github.com/fishaudio/fish-speech)
- [Release notes](https://github.com/fishaudio/fish-speech/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/fishaudio-fish-speech
