# Bert-VITS2: a multilingual TTS backbone with BERT text features

> Bert-VITS2 pairs a VITS2 acoustic backbone with multilingual BERT embeddings for text. The README now points users to Fish-Speech and states the project is not maintained in the short term, so this is a review for people deciding whether to adopt a frozen codebase.

**fishaudio/Bert-VITS2** — vits2 backbone with multilingual-bert

- Repository: https://github.com/fishaudio/Bert-VITS2
- Stars: 8,800 · Forks: 1,307
- Language: Python
- License: AGPL-3.0
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/fishaudio-bert-vits2

## The text encoder problem Bert-VITS2 was built to solve

Classic VITS-style TTS conditions the acoustic model on phoneme or character sequences. That works for a single language with clean text, but it degrades when the same model has to handle Chinese, Japanese, English and Korean, and it has no natural way to carry sentence-level context into prosody. Bert-VITS2's answer, as the repository name and README state, is a VITS2 backbone with multilingual BERT features. The intended audience is people training or fine-tuning a TTS model on their own speaker data, not people looking for a hosted API. The README is explicit that this is a codebase to study: it tells advanced users to read the code themselves to learn how to train. Note also that the README carries legal restrictions on use, including a prohibition on political purposes, and the repository is licensed AGPL-3.0.

## How the BERT features reach the VITS2 backbone

The pipeline visible in the repository layout runs from raw audio to trained model in discrete stages, each with its own script. resample.py normalises sample rates, preprocess_text.py handles text normalisation and phonemisation, bert_gen.py produces the BERT feature files, spec_gen.py generates spectrograms, and train_ms.py trains the multi-speaker model. The bert/ directory holds the BERT integration, and text/ holds the per-language front ends. That split matters because the BERT step is a separate offline pass: the features are written to disk before training, so a text front-end change means regenerating them rather than just restarting training. requirements.txt shows the language coverage is assembled from separate packages rather than one unified G2P layer: jieba, pypinyin and cn2an for Chinese, fugashi, mecab-python3, unidic-lite and pyopenjtalk-prebuilt for Japanese, g2p_en and cmudict for English, and jaconv and pykakasi for Japanese text conversion. WeTextProcessing is pinned with a platform marker that excludes Windows, which tells you the text pipeline is not identical across platforms.

## Installing Bert-VITS2 and running the preprocess WebUI

The README does not give a step-by-step install section. It says that for a quick guide you should refer to webui_preprocess.py, so that file is the entry point the project itself recommends. The repository ships requirements.txt, which is the only dependency list available. Install the Python dependencies first:

```bash
pip install -r requirements.txt
```

Several of those packages are compiled or platform-specific, and WeTextProcessing is marked as not for Windows, so expect the install to be smoother on Linux than on a Windows machine. Once the dependencies resolve, start the preprocessing interface:

```bash
python webui_preprocess.py
```

Gradio is pinned at 3.50.2 in requirements.txt, so the UI you get is the one that version provides. From there the workflow follows the script names in the repository root: resample your audio, run the text preprocessing, generate BERT features with bert_gen.py, generate spectrograms with spec_gen.py, then train with train_ms.py. The repository also contains webui.py, hiyoriUI.py and infer.py for the inference side, and configs/ plus default_config.yml for model configuration. The README does not document expected output formats or error messages for any of these steps, so read the scripts before running them on a large dataset.

## Maintenance status is the first thing to check

The README opens with a project recommendation: it points to Fish-Speech, another FishAudio TTS project, describes it as currently open source SOTA in quality and under continuous maintenance, recommends it as a replacement for BV2, and states that this project will not be maintained in the short term. The repository is not archived and the last push was on 2026-09-07, but the maintainers' own stated position is that development has stopped for now. The most recent releases are from early 2024: Extra in January 2024, Extra-v2 later that month, and JP-Extra on 2024-02-01. Treat this as a frozen codebase. That is not automatically disqualifying, since the training and inference scripts still exist and the AGPL-3.0 licence permits forking, but it changes what you are buying: you are adopting a snapshot, not a dependency that will receive fixes. For a project whose failure modes often surface as environment and dependency problems, that is a real cost.

## Where Bert-VITS2 is the wrong choice

The clearest wrong choice is a new production TTS service. A frozen upstream means dependency drift is your problem: gradio is pinned to 3.50.2, and the text front end depends on a long list of packages including pyopenjtalk-prebuilt and unidic-lite, any of which can break on a newer Python. The README does not document a rollback path, a version compatibility matrix, or a supported Python version. Second, if you need a runtime this repository does not target, look elsewhere: there is an export_onnx.py and an onnx_infer.py, so ONNX export exists, but the repository has no MNN export path despite MNN being a common search term around this project. Third, if your goal is simply to generate speech from text with a pretrained model, training scripts and a preprocessing UI are the wrong shape of tool; the README's own recommendation points at Fish-Speech for that. Finally, the licence and the README's usage restrictions both constrain deployment, so a commercial closed-source product is a poor fit without legal review.

## Bert-VITS2 versus Style-Bert-VITS2

Style-Bert-VITS2 is a separate project that appears repeatedly in searches around this one, and the difference is architectural rather than cosmetic. Bert-VITS2 conditions on BERT text features to improve multilingual pronunciation and context. Style-Bert-VITS2 adds an explicit style or emotion conditioning path on top of that idea, so you can steer delivery rather than only text content. This repository does have an emotional/ directory and the README lists emotional-vits among its references, so emotion is not absent here, but the style-conditioning line of work is the other project's focus. The practical consequence: if your requirement is per-utterance style control, Bert-VITS2 is not the shortest route, and its frozen status makes it a poor base to extend. If your requirement is multilingual text handling with a VITS2 backbone and you intend to read and modify the code, the smaller surface here is easier to reason about than a superset project.

## Licence and the cost of carrying a fork

The repository is AGPL-3.0. That is a strong copyleft licence with a network-use clause, and it is materially different from permissive licences used by some TTS projects. If you modify Bert-VITS2 and let users interact with it over a network, the AGPL's source-availability expectations apply. This is not legal advice; the point is that the licence choice should be part of the adoption decision, not an afterthought discovered at release time. The README adds its own restrictions on top, prohibiting uses that violate named Chinese laws and prohibiting political use. Upgrade cost is the other half of the ledger. Because the README states the project is not maintained in the short term, there is no upgrade path to plan for: you fork, you pin your dependencies, and you maintain the text and audio pipeline yourself. Budget for that, especially the per-language front ends, which are the part most likely to need attention as your data changes.

## Conclusion

Adopt Bert-VITS2 only if you need a VITS2-style backbone with BERT text conditioning and you are prepared to own the code, because the README states the project is not maintained in the short term and points to Fish-Speech. Do not adopt it for a new production TTS service where you need upstream fixes, or if your deployment target is MNN or another runtime the repository does not ship an exporter for. Before committing, verify the AGPL-3.0 obligations against your distribution model, confirm that a pretrained checkpoint exists for your target language, and check whether for_deploy/ or onnx_infer.py covers your inference path.

## FAQ

### What is Bert-VITS2 used for?

It is a text-to-speech codebase that combines a VITS2 backbone with multilingual BERT text features, aimed at people training or fine-tuning a TTS model on their own speaker data. The README directs users to webui_preprocess.py for a quick guide.

### Is Bert-VITS2 still maintained?

The README states that the project will not be maintained in the short term and recommends Fish-Speech instead. The repository is not archived, the last push was on 2026-09-07, and the most recent releases date from early 2024.

### Does Bert-VITS2 support Korean?

The README does not list supported languages, and requirements.txt shows front-end packages for Chinese, Japanese and English only. There is no Korean-specific dependency in that file, so Korean support is not documented.

### Can Bert-VITS2 export to ONNX?

The repository contains export_onnx.py, onnx_infer.py and an onnx_modules/ directory, so an ONNX export and inference path exists in the code. The README does not document the export procedure or any runtime other than ONNX.

### What licence does Bert-VITS2 use?

The repository is licensed AGPL-3.0, and the README adds usage restrictions prohibiting purposes that violate named Chinese laws and any political use. The AGPL includes a network-use clause, so distribution and hosted use need review.

## Sources

- [fishaudio/Bert-VITS2 on GitHub](https://github.com/fishaudio/Bert-VITS2)
- [Issues](https://github.com/fishaudio/Bert-VITS2/issues)
- [License: AGPL-3.0](https://github.com/fishaudio/Bert-VITS2/blob/master/LICENSE)
- [README](https://github.com/fishaudio/Bert-VITS2/blob/master/README.md)
- [Releases](https://github.com/fishaudio/Bert-VITS2/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/fishaudio-bert-vits2
