Model or dataset
devnen/Dia-TTS-Server avatar
devnen/Dia-TTS-Server

Dia TTS Server: one endpoint for Dia 1.6B and the Dia 2 family

Self-host the powerful Dia TTS model. This server offers a user-friendly Web UI, flexible API endpoints (incl. OpenAI compatible), support for SafeTensors/BF16, voice cloning, dialogue generation, and GPU/CPU execution.

355 stars66 forksPythonMIT

At a glance

What is it?
Dia TTS Server wraps Nari Labs' Dia models in a FastAPI service with an OpenAI-compatible speech endpoint, a web UI for hot-swapping between three model sizes, forty-three built-in voices, chunked long-form synthesis, and voice cloning, deployed through Docker on CUDA or CPU.
Who is it for?
Dia TTS Server suits someone who wants dialogue-style speech with speaker turns and non-verbal sounds, served behind an OpenAI-shaped endpoint so existing client code works unchanged, and who would rather self-host than call a hosted API.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.

Editorial analysis

One server fronts three Dia model sizes

The Dia models from Nari Labs come in three sizes, and this server puts all of them behind one FastAPI application. The original Dia 1.6B is supported alongside the newer Dia 2 family: Dia2-1B, a streaming model at roughly a billion parameters, and Dia2-2B, a higher quality option at roughly two billion. Switching happens from the web UI without restarting the server, using a selector dropdown that sits above the generation card with real-time status indicators. The switch does not block on the download. Model fetching and loading run on a background thread so the server stays responsive, a progress modal shows the download and loading phases, and a download in progress can be cancelled if you change your mind mid-switch. Underneath, every model is described in a single registry holding its Hugging Face repo ID, parameter count, available voice modes, and cloning method, and the engine resolves either the friendly selector name or the full repo ID from that one table.

Dia 2 conditions each speaker separately

The reason the two model families are not interchangeable shows up in how voices are attached. Dia 1 uses a single audio prompt for the whole utterance, which means one reference clip conditions the entire generation. Dia 2 replaces that with a prefix-speaker conditioning approach, so each speaker in a dialogue can carry an independent voice reference. That distinction is what makes multi-character output viable on Dia 2 rather than a limitation, and it is why migrating a two-speaker script between the two models is not just a model-name change. Dialogue is written with speaker tags in the text itself, using markers like S1 and S2 to assign turns, and the Dia models are built for exactly this shape of input, including non-verbal sounds written inline as cues such as laughs or sighs. The practical upshot is that a single request can contain a multi-character exchange with distinct timbres, which is the case the original single-prompt approach handled worst.

Forty-three built-in voices remove the cloning step

Voice cloning is the interesting path, but it is not the default, because the server ships forty-three curated voices that already work. They live in a voices directory in the repository and are selectable from a dropdown in the predefined voices mode, with the server supplying the transcripts those voices require. That last detail is what makes them reliable: a cloned voice needs a transcript matched to the reference audio, whereas a predefined voice arrives with its own pairing already known, so there is nothing to get wrong. The documentation is also candid about why this matters beyond convenience, noting that predefined voices avoid potential licensing issues that come with cloning someone else's voice. Consistency across generations is the other half of the feature, and it is handled by a fixed integer seed that can be combined with either predefined voices or cloning, so the same voice carries across separate generations and across the individual chunks of one long script.

Cloning normalises reference audio before the model sees it

When you do clone, the reference audio goes through a fixed preprocessing pipeline before inference. It is converted to mono, resampled to 44.1kHz, and truncated to roughly twenty seconds. Those three steps are not incidental: the model expects a consistent channel count and sample rate, and a fixed length window keeps a long reference from dominating the conditioning. Transcripts are handled separately and with a clear priority order. A local text file next to the reference audio is preferred and described as the accurate route, because a transcript you wrote can be corrected. If that file is missing, the server falls back to generating the transcript with Whisper, which is labelled experimental rather than dependable. The backend then prepends the transcript to the request. The reference audio directory is a normal mounted volume in the Docker setup, so swapping in a new speaker means dropping in a clip and its matching text file.

Long scripts split on sentence and speaker boundaries

Generation length is handled by chunking rather than by one enormous request. Long text is split into smaller pieces based on sentence boundaries and speaker tags, which matters because splitting on the wrong boundary would cut a sentence in half or separate a speaker turn from its tag. Each chunk is processed individually and the resulting audio is concatenated, which works around the previous generation limits. The behaviour is controlled from the interface rather than only in configuration: a toggle labelled Split text into chunks enables it, and a chunk size slider sets the granularity. Both predefined voices and the seed setting apply across that boundary, so a voice stays the same from the first chunk to the last. This is the main structural difference between this server and a single-shot synthesiser, and it is also the reason a long narration can sound slightly segmented at the joins, since the model never sees the whole script as one utterance.

A missing model package degrades instead of crashing

The Dia 1 and Dia 2 packages are imported defensively with a try and except, and the consequence is worth understanding. If one of them is not installed, that model is shown as unavailable in the interface while the server keeps working with whichever packages are present. Since the Dia 2 support is an optional install in the requirements file, a base installation that never added it still runs Dia 1.6B rather than failing at startup. The same defensive thinking shows up in the endpoints added for multi-model operation. A restart endpoint performs an asynchronous model hot-swap and returns immediately. Separate endpoints return the current model details, the full registry for populating the dropdown, and the loading progress that the interface polls. A cancel endpoint aborts an in-progress model load, and an unload endpoint frees the current model's resources, which is what you want before switching to a larger one on a constrained GPU.

Docker exposes 8003 and refuses a mounted config file

Container deployment is handled by a compose file built on an NVIDIA CUDA runtime image for Ubuntu, with ffmpeg and the libsndfile system library installed since the audio libraries need them. The service listens on port 8003 and starts by invoking the server module directly. Four directories are mounted as volumes for persistence, covering the model cache, reference audio, outputs, and voices, and the compose file carries an explicit instruction not to mount the main configuration file, because the application is meant to create that inside the container. GPU access is requested through the device mechanism with all GPUs exposed, alongside the cgroup rules some NVIDIA container toolkit versions require, and a commented legacy block offers the older reservation style as an alternative, with a warning not to enable both. Hugging Face transfers are enabled inside the container for faster model downloads. Configuration itself lives in a YAML file, with a dotenv file used only to seed that configuration on first run or a reset.

The newest tag predates the version the README describes

There is a gap worth checking before you build on this. The README's newest-features section is headed with version 2.0.0 and describes the Dia 2 multi-model work, including the hot-swapping, the registry, and the per-speaker conditioning, while the release it labels as previous is 1.4.0. The tagged releases tell a different story. The newest published tag is 1.4.0, dated 29 April 2025, with 1.0.0 before it, and the last recorded push to the main branch is 28 March 2026. That means the multi-model work described at length in the README has no matching release tag, and there have been no commits to the main branch in the roughly six months since that push. None of that says the features are absent from the default branch, only that the version history and the documentation are describing different things. If you are pinning a version rather than tracking the branch, verify which of those features that tag actually contains.

Editorial conclusion

Dia TTS Server suits someone who wants dialogue-style speech with speaker turns and non-verbal sounds, served behind an OpenAI-shaped endpoint so existing client code works unchanged, and who would rather self-host than call a hosted API. It is a weaker fit if you need long-form narration in a single pass, since long inputs are split and rejoined rather than synthesised as one utterance, and it is a poor fit if you need recent upstream model work, given the gap between the newest tagged release and the feature set the README describes. Before deploying, check the model package you actually installed, since a missing Dia 1 or Dia 2 package silently removes that model from the selector, and confirm your reference transcripts exist as text files rather than relying on the Whisper fallback.

Frequently asked questions

Which Dia models does Dia TTS Server support?

Three: the original Dia 1.6B, plus the Dia 2 family with Dia2-1B for streaming at about a billion parameters and Dia2-2B for higher quality at about two billion. They can be switched from the web UI without restarting the server.

Does Dia TTS Server have an OpenAI-compatible endpoint?

Yes. It exposes an OpenAI-compatible audio speech endpoint at /v1/audio/speech, so clients that expect OpenAI's API structure can talk to the Dia models directly.

How does Dia TTS Server keep a voice consistent across chunks?

By selecting one of 43 predefined voices or a cloned voice, optionally combined with a fixed integer seed, the same voice carries across separate generations and across the individual chunks of one long script.

Does Dia TTS Server require a GPU?

No. CUDA acceleration is detected automatically with fallback to CPU. The container image is based on an NVIDIA CUDA runtime and the compose file requests GPU access for the service.

Official sources

  1. devnen/Dia-TTS-Server on GitHub
  2. Issues
  3. License: MIT
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/devnen-dia-tts-server.svg)](https://hysenlabs.com/projects/devnen-dia-tts-server)