Fish Speech S2 Pro: Dual-AR Text-to-Speech with Inline Emotion Control
GitHub describes it as SOTA Open Source TTS. The repository metadata lists Python as its primary language. The metadata lists the NOASSERTION license. This article stays within the project description and details documented in the GitHub repository README.
At a glance
- What is it?
- Fish Speech is a multilingual text-to-speech system from Fish Audio, built around a 4-billion-parameter Dual-Autoregressive architecture trained on over 10 million hours of audio. The S2 Pro model scores best overall on Seed-TTS Eval for both Chinese and English word error rate, and supports fine-grained emotion control through in-text tags.
- Who is it for?
- Teams building multilingual text-to-speech pipelines who need fine-grained prosody control and strong word error rate performance should evaluate Fish Speech S2 Pro against their target languages. The FISH AUDIO RESEARCH LICENSE is not a standard open-source license and requires review before any commercial deployment.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 14 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 22, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Fish Speech Solves and Who Uses It
Fish Speech is a text-to-speech system designed for use cases where standard TTS systems produce speech that sounds synthetic or fails to capture prosody in multilingual or emotionally complex text. The system targets developers building voice interfaces, content pipelines, and agent-based applications that need speech generation beyond simple narration.
The S2 Pro model is the flagship variant as of the current release, version 2.0.0 in pyproject.toml. It was trained on over 10 million hours of audio data covering more than 80 languages. The README describes the benchmarks across several axes: on the Seed-TTS Eval, S2 achieves a word error rate of 0.54% in Chinese and 0.99% in English, which the README states is best overall among evaluated models including closed-source systems such as Qwen3-TTS, MiniMax Speech-02, and Seed-TTS itself. On the Audio Turing Test, S2 achieves a posterior mean of 0.515, which the README describes as surpassing Seed-TTS (0.417) by 24% and MiniMax-Speech (0.387) by 33%. On EmergentTTS-Eval, the win rate is 81.88%.
The Dual-Autoregressive Architecture
S2 Pro uses what the README calls a master-slave Dual-AR architecture consisting of two components operating together. The Slow AR is a decoder-only transformer with 4 billion parameters that operates along the time axis and predicts the primary semantic codebook. The Fast AR is a 400-million-parameter model that generates the remaining 9 residual codebooks at each time step, reconstructing acoustic detail that the Slow AR leaves abstract.
The codec underneath is an RVQ audio codec with 10 codebooks operating at approximately 21 Hz. The README describes this asymmetric design as achieving peak audio fidelity while significantly boosting inference speed, since the computationally heavy Slow AR only needs to produce one codebook and the lighter Fast AR handles the remaining nine. The system also applies reinforcement learning alignment using Group Relative Policy Optimization (GRPO). The reward models come from the same model suite used for data cleaning and annotation, which the README states resolves distribution mismatch between pre-training data and post-training objectives.
Running Fish Speech with Docker
The repository provides Docker Compose configuration for two deployment profiles. The web UI profile runs a Gradio interface at port 7860 by default, configurable through the GRADIO_PORT environment variable:
services:
webui:
ports:
- "${GRADIO_PORT:-7860}:7860"
server:
ports:
- "${API_PORT:-8080}:8080"The server profile runs an API server at port 8080 by default, configurable through API_PORT. The compose.yml file extends a base service definition in compose.base.yml for both profiles. A separate compose.rocm.yml exists for AMD ROCm GPU configurations. The repository documentation at speech.fish.audio covers the full installation process including Docker setup, command-line inference, web UI inference, and server configuration. The official documents linked from the README cover getting started for human users and a brief prompt for LLM agent users.
Inline Emotion and Prosody Control
One of the features the README highlights is sub-word level control of prosody and emotion through inline tags. The system accepts tags placed directly in the text, such as [whisper], [excited], [angry], [laughing], [pause], and [emphasis]. The README lists the supported tag set as including over 15,000 unique tags and notes that it supports free-form text descriptions rather than a fixed preset list. Examples from the README include [whisper in small voice], [professional broadcast tone], and [pitch up].
The full tag list documented in the README also covers [inhale], [chuckle], [tsk], [singing], [laughing tone], [interrupting], [chuckling], [excited tone], [volume up], [echo], [low volume], [sigh], [low voice], [screaming], [shouting], [loud], [surprised], [short pause], [exhale], [delight], [panting], [audience laughter], [with strong accent], [volume down], [clearing throat], [sad], [moaning], [shocked], and others. These tags can be placed at any position in the text and work across the supported languages. The system also natively supports multi-speaker and multi-turn conversation generation, which the README mentions as part of the S2 Pro capability set.
Benchmark Context and Multilingual Coverage
The README includes a table of benchmark results comparing S2 against other systems. On the Fish Instruction Benchmark, S2 achieves a task adherence rate of 93.3% and a quality score of 4.51 out of 5.0. On the multilingual MiniMax Testset, S2 achieves the best word error rate in 11 of 24 languages and the best speaker similarity in 17 of 24 languages.
The EmergentTTS-Eval results break down by category: paralinguistics at 91.61% win rate, questions at 84.41%, and syntactic complexity at 83.39%. These numbers come from the README's benchmark table and are included here as stated in that source. The technical report covering the system in full is at arxiv.org/abs/2603.08823, and the HuggingFace model page is at huggingface.co/fishaudio/s2-pro. A blog post describing the S2 open-source release is at fish.audio/blog/fish-audio-open-sources-s2/.
License, Limitations, and What to Verify
Fish Speech uses the FISH AUDIO RESEARCH LICENSE, not a standard OSS license such as MIT or Apache-2.0. The README states the team will take action against license violations and includes a legal disclaimer noting that illegal usage of the codebase is the user's responsibility. Teams considering commercial deployment must review the full license text before proceeding.
The pyproject.toml requires Python 3.10 or higher and pins torch to version 2.8.0. The dependencies list is long, including gradio, uvicorn, librosa, and safetensors, among others. The project is at version 2.0.0 as recorded in pyproject.toml, and the most recent GitHub release is v2.0.0-beta from 2026-03-10, with a stable 1.5.1 release from 2025-05-31. The last repository push was on 2026-09-16, and the project is not archived.
For users comparing options: the README includes benchmark comparisons against fish speech vs gpt sovits and fish speech vs f5 tts in the RELATED SEARCHES data, and the wiki at the Fish Audio blog post covers the architectural differences. The README does not document rollback to the 1.5.x model series from within the current installation, so teams relying on a specific model version should pin the release explicitly.
Editorial conclusion
Teams building multilingual text-to-speech pipelines who need fine-grained prosody control and strong word error rate performance should evaluate Fish Speech S2 Pro against their target languages. The FISH AUDIO RESEARCH LICENSE is not a standard open-source license and requires review before any commercial deployment. Teams already using GPT-SoVITS or Kokoro for TTS should review the benchmark comparisons on the Fish Audio blog before migrating, since the architectures differ. The live demo at fish.audio and the Gradio web UI launched at port 7860 provide a low-cost way to evaluate output quality before committing to the infrastructure.
Frequently asked questions
Is Fish Speech open source?
Fish Speech is published on GitHub with its model weights, but it uses the FISH AUDIO RESEARCH LICENSE rather than a standard open-source license. The README states the team will take action against license violations, so teams should review the license before commercial use.
How do you install Fish Speech?
The documentation at speech.fish.audio/install covers the installation steps. The repository also provides Docker Compose configuration for a web UI on port 7860 and an API server on port 8080. Python 3.10 or higher and torch 2.8.0 are required as listed in pyproject.toml.
What is Fish Speech?
Fish Speech is a multilingual text-to-speech system built on a 4-billion-parameter Dual-Autoregressive architecture, developed by Fish Audio. S2 Pro, the flagship model, is trained on over 10 million hours of audio covering 80+ languages and supports inline emotion control through tags like [whisper] and [excited].
Is Fish Speech free to use?
The code and model weights are publicly available, but under the FISH AUDIO RESEARCH LICENSE rather than a permissive license. The README warns that illegal usage is the user's responsibility and that violations will be acted upon. Confirm the license terms before commercial deployment.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/fishaudio-fish-speech)