Open-source project
SWivid/F5-TTS avatar
SWivid/F5-TTS

F5-TTS: flow-matching voice cloning you can install with pip

Official code for "F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching"

15,316 stars2,240 forksPythonMIT

At a glance

What is it?
F5-TTS is an MIT-licensed Python text-to-speech system that clones a voice from a short reference clip and a transcript. It installs as a pip package or a Docker image, and it also ships training and finetuning entry points.
Who is it for?
Adopt F5-TTS if you need MIT-licensed voice cloning driven from Python or a CLI, and you can supply a reference clip plus its transcript. Skip it if you want a hosted API with an SLA, or if you cannot accept that the README does not document the licensing of the model checkpoints separately from the code.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 9 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap F5-TTS fills: cloning a voice from a clip and a transcript

Most text-to-speech pipelines need either a studio recording of one speaker or a paid API that owns the voice model. F5-TTS takes a third route. You give it a short reference audio file and the words spoken in that file, and it synthesises new text in that voice. The repository describes the model as a Diffusion Transformer with ConvNeXt V2, trained with flow matching, and the README points to the arXiv paper 2410.06885 for the method. A second model, E2 TTS, is included as a Flat-UNet Transformer and is described as the closest reproduction of a separate paper.

The audience is narrow and specific. This is for engineers who are comfortable creating a conda environment, installing a matched PyTorch build, and reading a source tree. The README says to take a moment to read the detailed guidance under src/f5_tts/infer before expecting good output, which is a fair signal: the default settings are a starting point, not a finished product. If you want a one-click service, the Hugging Face and ModelScope Spaces linked at the top of the README are the better entry point.

How the flow-matching pipeline is wired together

The architecture visible from the repository is a two-stage acoustic pipeline. Text and reference audio are encoded, a DiT backbone predicts a flow field, and a vocoder turns the result into a waveform. The dependency list names vocos, x_transformers, torchdiffeq and transformers, which matches that shape: x_transformers for the transformer blocks, torchdiffeq for the flow integration, vocos for the vocoder.

Sway Sampling is the part worth understanding before you tune anything. The README calls it an inference-time flow step sampling strategy, and it is applied at generation rather than training. That means the number of function evaluations and the sampling schedule change output quality without any retraining. The runtime benchmark table in the README reports results at 16 NFE, which is the knob you will be adjusting.

The repository also ships a deployment path under src/f5_tts/runtime/triton_trtllm. The README's benchmark table lists F5-TTS Base with Vocos at 253 ms average latency and an RTF of 0.0394 in client-server mode on a single L20 GPU, against an RTF of 0.1467 for offline PyTorch on the same hardware. Those are the project's own numbers, not an independent measurement, and they cover 26 prompt and target pairs at 16 NFE.

Installing F5-TTS and running your first local generation

The README starts with a separate environment and Python 3.10 or newer, then FFmpeg. The pyproject file declares Python 3 and the badges say Python 3.10, so a 3.11 or 3.12 environment is what the installation block itself shows.

bash
conda create -n f5-tts python=3.11
conda activate f5-tts
conda install ffmpeg

Next install PyTorch matched to your device. The README gives one example per accelerator family. For an NVIDIA card the current example is:

bash
pip install torch==2.8.0+cu128 torchaudio==2.8.0+cu128 --extra-index-url https://download.pytorch.org/whl/cu128

Apple Silicon uses the plain stable wheels, and AMD needs a ROCm build. The README warns that RDNA 3.5 and RDNA 4 cards such as the Radeon 8050S/8060S and RX 9060/9070 require ROCm 7.x, because those architectures are absent from ROCm 6.x, and that running ROCm 6.x on them produces HIP error: invalid device function.

For inference only, the package is on PyPI:

bash
pip install f5-tts

If you also intend to finetune, the README says to clone and install editable so the training scripts are present:

bash
git clone https://github.com/SWivid/F5-TTS.git
cd F5-TTS
pip install -e .

You then have four console scripts, declared in pyproject.toml: f5-tts_infer-gradio, f5-tts_infer-cli, f5-tts_finetune-gradio and f5-tts_finetune-cli. The quickest first run is the web interface.

bash
f5-tts_infer-gradio --port 7860 --host 0.0.0.0

You should see a Gradio app listening on port 7860. The README lists its features as basic TTS with chunk inference, multi-style and multi-speaker generation, and a voice chat mode powered by Qwen2.5-3B-Instruct. The model weights are fetched from Hugging Face on first use, which is why the Docker examples mount a volume at /root/.cache/huggingface/hub/.

If you prefer containers, the README gives both a build and a prebuilt image:

bash
docker build -t f5tts:v1 .
docker container run --rm -it --gpus=all --mount 'type=volume,source=f5-tts,target=/root/.cache/huggingface/hub/' -p 7860:7860 ghcr.io/swivid/f5-tts:main

The Dockerfile exposes 7860 and uses pytorch/pytorch:2.4.0-cuda12.4-cudnn9-devel as its base, so the container pins its own CUDA stack independently of whatever you install locally.

Where F5-TTS breaks down or is the wrong choice

The reference clip is the first failure mode. The model conditions on both the audio and its transcript, so a mismatch between what the clip says and the text you supply degrades the result. The README does not describe an automatic transcription step in the CLI path; it points you at the guidance under src/f5_tts/infer instead.

Hardware is the second. The README's runtime numbers come from an L20, and the Docker path assumes --gpus=all. The Intel GPU instructions require either the Intel Deep Learning Essentials or oneAPI Base Toolkit, or the IPEX route, which is a heavier setup than a pip install. Nothing in the README describes a CPU-only path with usable latency.

The third issue is legal rather than technical, and the repository does not resolve it. The code is MIT, but the README does not document the terms attached to the pretrained checkpoints separately, and it says nothing about consent for the voice you clone. If your use case involves a real person's voice, that gap is yours to close, not the project's.

Finally, the README is thin in places. There is no documented rollback procedure, no stated support window for old releases, and the installation section is written as a set of alternatives rather than a single supported path. That is normal for a research repository, but it means you should budget time for reading issues rather than expecting the README to cover every error.

F5-TTS against E2 TTS and against hosted TTS APIs

The most direct comparison is inside the repository. F5-TTS and E2 TTS ship together, and the README describes them as different backbones: F5-TTS is a Diffusion Transformer with ConvNeXt V2, E2 TTS is a Flat-UNet Transformer presented as the closest reproduction of a separate paper. The practical difference is that E2 TTS is the reproduction target, while F5-TTS is the one the project says has better training and inference performance in its v1 base model. If you are trying to match published numbers from the E2 paper, use E2 TTS. If you want the project's own recommended model, use F5-TTS.

The other comparison is against hosted TTS services. A hosted API gives you a stable endpoint, no GPU, and someone else's uptime. F5-TTS gives you the weights, the training code and the finetuning scripts, which means you can adapt the voice to your own data. The trade is operational: you own the CUDA version, the model download, the port and the latency. The README's Triton and TensorRT-LLM runtime exists precisely because the plain PyTorch path is slower, and the benchmark table shows roughly a 3.6x RTF gap between offline PyTorch and the client-server deployment on the same GPU.

Versioning, licence and the cost of staying current

The package is versioned through setuptools-scm, and the release cadence visible from the tags is roughly monthly, with 1.1.20 in April 2026, 1.1.21 in July 2026 and 1.1.22 later that same month. The last push to the default branch was on 2026-07-23. That is recent enough that the project is not dormant, but the README does not describe a deprecation policy, so a version bump can change inference behaviour without a migration note.

The dependency surface is the real upgrade cost. The install pulls torch, torchaudio, torchcodec, transformers, gradio, hydra-core, vocos, x_transformers, librosa, pydub, soundfile and more. Two of those are worth flagging. The numpy pin is conditional: numpy<=1.26.4 only applies on Python 3.10 and below, so a 3.11 environment resolves numpy differently and you may hit a different set of downstream constraints. And bitsandbytes is excluded on arm64 and Darwin, so Apple Silicon users get a different dependency set than Linux users.

On licensing: the repository and the pyproject metadata both state the MIT License, and the classifiers list it as OSI approved. That covers the code you install. It does not automatically cover the pretrained weights, and it does not cover the output you generate from a cloned voice. The README is silent on both, so treat those as questions for whoever owns the deployment, not as settled by the repository.

Editorial conclusion

Adopt F5-TTS if you need MIT-licensed voice cloning driven from Python or a CLI, and you can supply a reference clip plus its transcript. Skip it if you want a hosted API with an SLA, or if you cannot accept that the README does not document the licensing of the model checkpoints separately from the code. Before committing, confirm two things: that your GPU matches the PyTorch build you install, and that the reference audio you plan to clone is one you have the right to use. The code licence is MIT; what you generate with a cloned voice is a separate question the repository does not answer.

Frequently asked questions

What is F5-TTS?

It is a Python text-to-speech system built on flow matching, described in the README as a Diffusion Transformer with ConvNeXt V2. It generates speech from text and can clone a voice from a reference clip plus its transcript.

Is F5-TTS free and open source?

The repository is licensed under the MIT License, which the pyproject metadata also records, and the code is published on GitHub. The README does not state separate terms for the pretrained checkpoints.

How do I install F5-TTS locally?

Create a conda environment with Python 3.11 and install FFmpeg, install a PyTorch build matched to your GPU, then run pip install f5-tts for inference only or clone the repository and run pip install -e . if you also want training. A Docker image is also published at ghcr.io/swivid/f5-tts:main.

How do I use F5-TTS from Python or the command line?

The package installs four console scripts: f5-tts_infer-gradio, f5-tts_infer-cli, f5-tts_finetune-gradio and f5-tts_finetune-cli. Running f5-tts_infer-gradio --port 7860 --host 0.0.0.0 starts the web interface, and the CLI entry point handles the same inference without a browser.

What is the difference between F5-TTS and E2 TTS?

Both ship in this repository but use different backbones. The README describes F5-TTS as a Diffusion Transformer with ConvNeXt V2 and E2 TTS as a Flat-UNet Transformer presented as the closest reproduction of a separate paper.

Is voice cloning with F5-TTS legal?

The repository does not answer this. It states the MIT License for the code and says nothing about consent for the voice being cloned or about terms for the pretrained checkpoints, so the question sits outside what the project documents.

Official sources

  1. License: MIT
  2. Project website
  3. README
  4. Releases
  5. SWivid/F5-TTS on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/swivid-f5-tts.svg)](https://hysenlabs.com/projects/swivid-f5-tts)