Speech-AI-Forge: A Multi-Model TTS Server and Gradio WebUI
🍦 Speech-AI-Forge is a project developed around TTS generation model, implementing an API Server and a Gradio-based WebUI.
At a glance
- What is it?
- Speech-AI-Forge wraps a dozen open TTS and ASR models behind one FastAPI server and one Gradio interface. It is a consolidation layer for people who already know which voice model they want and are tired of running each one separately.
- Who is it for?
- Adopt Speech-AI-Forge if you are already running open TTS models and want one HTTP surface and one interface for all of them, and if AGPL-3.0 fits how you ship. Skip it if you need a single small, stable model in production or if you cannot accept a dependency tree that pins pydantic to 2.8.2.
- Can I use it commercially?
- Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
- Is it still maintained?
- Yes. The repository last received commits 133 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem Speech-AI-Forge solves, and for whom
Open TTS in 2026 is a pile of separate repositories. ChatTTS has its own sampling and refiner behaviour. CosyVoice, F5-TTS, FishSpeech, GPT-SoVITS, Index-TTS, Spark-TTS, FireRedTTS and Qwen3-TTS each ship their own inference script, their own checkpoint layout and their own way of accepting a reference voice. If you want to compare two of them on the same sentence, you install two environments, and if you want to serve either one over HTTP you write the wrapper yourself.
Speech-AI-Forge is a consolidation layer over that pile. The README describes it as a project built around TTS generation models that implements an API Server and a Gradio-based WebUI. The intended user is not someone who wants a hosted voice product. It is an engineer or a hobbyist who already runs these models locally, has a GPU, and wants one process that loads several of them and exposes a consistent interface. The model support table lists TTS entries including Index-TTS (v1/v1.5), Qwen3-TTS, FishSpeech 1.4, CosyVoice (v2/v3), FireRedTTS, F5-TTS (v0.6/v1), Spark-TTS, GPT-SoVITS and ChatTTS, plus a Cloud TTS row for MiniMax Cloud TTS.
That breadth is the whole pitch. It is also the source of the project's main cost, which is dependency weight: requirements.txt pulls in hydra-core, lightning, wandb, faster_whisper, stable-ts, moviepy, funasr, openai-whisper and more, because each supported model brings its own runtime.
Two entry points: webui.py and launch.py
The repository has two top-level scripts and they are deliberately different products. webui.py starts the Gradio interface. launch.py starts the API server alone, and the README frames it as the choice when you do not need the WebUI or you want higher API throughput. After starting it, the README says http://localhost:7870/docs shows which API endpoints are enabled, so the server ships with FastAPI's generated documentation rather than a hand-written reference page. The docs directory also contains api.md for more detail, and launch.py -h prints the script's arguments.
The WebUI is where most of the project's actual design lives. It groups features into TTS, SSML, voice management, ASR and tools. Under TTS the README lists speaker switching with built-in voices (27 ChatTTS and 7 CosyVoice voices plus one reference voice), custom voice upload, style control, long-text inference with automatic splitting and a Batch size setting, ChatTTS's native refiner, splitter configuration for end-of-sentence markers and thresholds, an adjuster for speed, pitch and volume with loudness equalisation, a voice enhancer model, and a history of the last three generations. The SSML tab is the more unusual part: a splitter, a Podcast mode for long multi-role audio, a From Subtitle mode that turns subtitle files into SSML scripts, and a script editor that exports and edits those scripts. ASR covers Whisper and SenseVoice transcription plus forced alignment against a transcript.
The data flow is therefore: text or subtitle file enters the WebUI, the splitter cuts it into segments sized for the selected model, the model generates audio per segment, and the adjuster and enhancer post-process the result. The API path skips the interface but uses the same model modules under modules/.
Installing Speech-AI-Forge and generating a first clip
The README points at docs/dependencies.md for environment preparation and a 模型下载 (model download) section for checkpoints, and says to install those first. It does not give a single copy-paste install command for the Python path, so the honest starting point is those two documents rather than a command I would be inventing here. What the README does give is the launch command:
python webui.pyThat starts the Gradio WebUI. The README's feature list is what you should expect to find in it: a TTS tab with speaker selection, a long-text path with automatic splitting, and an SSML tab.
For a headless setup, the API server is a separate script:
python launch.pyThe README states that after it starts, http://localhost:7870/docs lists the enabled API endpoints, and that python launch.py -h shows the script's parameters. The repository also carries examples/python/ and examples/javascript/ directories, which is where to look for client-side call shapes rather than guessing at request bodies.
Docker is the third route, and the README gives two compose files:
docker-compose -f ./docker-compose.webui.yml up -d
docker-compose -f ./docker-compose.api.yml up -dThe first brings up the WebUI, the second the API. Configuration for each comes from .env.webui and .env.api respectively. The Dockerfile itself is worth reading before you build: it starts from fabridamicelli/torchserve:0.12.0-gpu-python3.10, rewrites the Ubuntu apt sources to the Tsinghua mirror, installs ffmpeg and rubberband-cli, and installs requirements.txt from the Tsinghua PyPI mirror. Its default CMD is bash, not a server, so the compose files are doing the actual startup work.
There is also a Windows portable build distributed through Releases, described as unpack-and-run, and a Colab notebook at colab.ipynb for trying it without a local GPU.
Where Speech-AI-Forge is the wrong tool
The dependency set is the first real constraint. requirements.txt pins pydantic==2.8.2 with an explicit note that upgrading to 2.11 breaks things, and pins gradio==4.44.0, transformers==4.57.3 and numpy==1.26.4. If Speech-AI-Forge has to live inside an application that already depends on a newer pydantic or a newer Gradio, you are resolving that conflict by hand. The pyproject.toml comments say poetry was avoided because installing PyTorch on Windows produced hash errors, so the project's own packaging is deliberately plain pip plus a uv index configuration pointing at download.pytorch.org/whl/cu126 and the Tsinghua PyPI mirror. That is a pragmatic choice, and it also means you inherit a flat requirements file rather than a resolvable dependency graph.
Second, breadth costs memory. Loading several TTS models plus Whisper or SenseVoice for ASR in one process is not a small deployment. If your only need is one voice on one model, running the upstream repository directly gives you a smaller surface and fewer moving parts.
Third, the documentation is uneven. The README is feature-oriented: it tells you what the WebUI can do in detail and says relatively little about API request and response schemas, error behaviour, or concurrency limits. The README does not document rollback between model versions, and it does not describe how the server behaves under concurrent load. Anyone planning to put this behind a queue should read docs/api.md and the source before assuming throughput characteristics.
Finally, the licence. AGPL-3.0 is a strong copyleft licence, and the project is a network service by design. If you plan to expose a modified Speech-AI-Forge over a network, the licence terms are the thing to read before you build a product on it. That is a description of the licence, not legal advice.
How it compares with running one upstream model
The natural alternative is not another aggregator, it is going direct: clone ChatTTS, or F5-TTS, or CosyVoice, and run that repository's own inference script. The difference in approach is real. Upstream repositories optimise for one model's quality and one model's quirks. ChatTTS's refiner and seed-based voice sampling are first-class there. F5-TTS's reference-audio workflow is first-class there. You get the author's intended defaults, the smallest dependency set for that model, and documentation written by the people who trained it.
Speech-AI-Forge trades that depth for uniformity. You get one server, one WebUI, one voice-management layer, and the ability to switch models without reinstalling. The SSML and subtitle tooling has no equivalent in most single-model repositories: turning a subtitle file into a multi-role SSML script and then into audio is a workflow that only makes sense once several voices and a splitter live in the same interface. If your work is audiobook or podcast production across multiple voices, the aggregation is the feature. If your work is one production voice at low latency, the aggregation is overhead.
A second alternative is a cloud TTS API. Speech-AI-Forge does list MiniMax Cloud TTS support, so the two are not mutually exclusive, but the local models are the reason the project exists. Local inference means no per-character cost and no data leaving your machine, at the price of a GPU and the setup time described above.
Maintenance, releases and what the licence implies
The most recent release is portable_v0.7, published on 2026-02-02, and package.json carries version 0.8.0-rc. The last push to the default branch was on 2026-05-21. The README's breaking change log is unusually informative for judging upgrade cost: it dates each model addition, from CosyVoice in July 2024 through to CosyVoice3 on 2026-02-02, Qwen3-TTS on 2026-01-29, Index-TTS-2 on 2025-09-12, and MiniMax cloud TTS on 2026-04-02. It also records API-level changes such as the v2/tts endpoint added on 2024-11-11 and the ASR API added on 2024-08-01.
That log tells you what to expect from an upgrade: new model support arrives frequently, and it arrives as a breaking change. Pinned versions in requirements.txt mean an upgrade is not a one-line pip install; it can move transformers, gradio or numpy underneath you. Budget for testing the specific models you depend on after each pull, not just for the server starting.
The licence is AGPL-3.0. In practical terms for an engineering decision: if you modify Speech-AI-Forge and let users interact with it over a network, the AGPL's source-availability condition is the part that matters, and it is broader than the GPL's. Internal use that does not expose the software to outside users is a different situation from a hosted service. Whether your specific deployment triggers those obligations is a question for your own counsel, not something this article can settle. The relevant fact here is simply that the project is AGPL-3.0 and is designed to run as a network server, so the two facts interact.
Editorial conclusion
Adopt Speech-AI-Forge if you are already running open TTS models and want one HTTP surface and one interface for all of them, and if AGPL-3.0 fits how you ship. Skip it if you need a single small, stable model in production or if you cannot accept a dependency tree that pins pydantic to 2.8.2. Before committing, install it in a clean environment, confirm which model checkpoints your target voices actually need, and read the API docs at docs/api.md to check that the endpoint you plan to call exists in the current release.
Frequently asked questions
What is Speech-AI-Forge used for?
It generates speech from text and transcribes speech to text, through a Gradio WebUI or a FastAPI server. The README describes it as a project built around TTS generation models that implements an API Server and a WebUI, with support for ChatTTS, CosyVoice, F5-TTS, FishSpeech, GPT-SoVITS and others.
Is Speech-AI-Forge free to download and use?
The source is published on GitHub under AGPL-3.0, and Releases includes a Windows portable build described as unpack-and-run. There is no pricing information in the README. The licence's network-service conditions are the thing to check before hosting it for others.
How do I start the Speech-AI-Forge API server?
Run python launch.py. The README states that http://localhost:7870/docs then shows the enabled API endpoints, and that python launch.py -h prints the script's arguments.
Does Speech-AI-Forge run on macOS?
The README's deployment table lists a Windows portable package, a Colab notebook, local deployment and Docker. macOS is not named as a supported target in that table, and the Dockerfile is built on a CUDA GPU base image.
What is the difference between webui.py and launch.py in Speech-AI-Forge?
webui.py starts the Gradio interface, while launch.py starts only the API server. The README recommends launch.py when you do not need the WebUI or want higher API throughput.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/lenml-speech-ai-forge)