Open-source project
babysor/MockingBird avatar
babysor/MockingBird

MockingBird and the synthesizer you have to train before anything runs

🚀Clone a voice in 5 seconds to generate arbitrary speech in real-time

36,901 stars5,173 forksPythonNOASSERTION

At a glance

What is it?
MockingBird reuses a pretrained encoder and vocoder but ships no compatible synthesizer, so the bundled demo_cli is not working. The install path is frozen at a mid-2021 PyTorch, and that choice decides which Python, CUDA and platform you can use.
Who is it for?
MockingBird is worth your time if you can train a synthesizer on Mandarin data and you want to read and modify a real-time voice cloning pipeline, because the encoder and vocoder are reusable and the web and streamlit entry points are in the tree. It is the wrong project if you wanted to run a demo, work on Apple silicon without compiling things, or rely on someone still shipping checkpoints.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Activity is slowing. The repository last received commits 7 months ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The bundled demo_cli does not run until you train a synthesizer

The model section is blunt about this. The project reuses the pretrained encoder and vocoder but not the synthesizer, because the original model is incompatible with the Chinese symbols, and the stated consequence is that the demo_cli is not working at this moment, so additional synthesizer models are required. That is the first thing a newcomer meets, and it is not a configuration problem. No flag, no environment variable and no download in this repository turns the CLI demo into a working demo. The cost lands on whoever wants a voice in the first place: you are signing up for the training steps before the inference steps, and the documented training sets are Mandarin corpora such as aidatatang_200zh, magicdata, aishell3 and data_aishell rather than a few minutes of your own recording.

Reusing the encoder is the only shortcut the project offers

The stated appeal of the design is that the result comes from only a newly trained synthesizer, because the pretrained encoder and vocoder are reused. In practice that reuses the hard part of an encoder or vocoder and still asks for the synthesizer. The tree reflects a multi-stage pipeline rather than one command: `demo_toolbox.py` for the desktop toolbox, `web.py` for the web server, `run.py`, `gen_voice.py`, `pre.py` and `train.py` for the rest, plus `monotonic_align/`, `control/`, `data/`, `models/`, `skills/` and `utils/`. The consequence for a reader is that the file list tells you the work is split, and the only stage marked optional is the encoder. The vocabularies and symbol tables the checkpoint expects still have to line up with whatever you train, which is exactly what the Chinese symbol incompatibility is about.

The install path is pinned to a PyTorch build from August 2021

The recommended environment is spelled out as repo tag 0.0.1 with PyTorch 1.9.0, Torchvision 0.10.0 and cudatoolkit 10.2, plus `requirements.txt` and `webrtcvad-wheels`, and the reason given is that `requirements.txt` was exported a few months ago so it does not work with newer versions. `pip install -r requirements.txt` is still the command, but treat the pins as binding rather than as a floor. The install failure people hit is explicit: `ERROR: Could not find a version that satisfies the requirement torch==1.9.0+cu102 (from versions: 0.1.2, 0.1.2.post1, 0.1.2.post2 )`, and the diagnosis offered is a low Python version, fixed by moving to 3.9. Note the tension with the stated minimum: the toolbox needs Python 3.7 or higher, yet the documented fix for a failed install is to leave the floor for a specific patch release.

Windows and Linux resolve to different pins, and webrtcvad is absent on one

`requirements.txt` splits by `platform_system` rather than shipping one environment. numpy is pinned to 1.19.3 when `platform_system == "Windows"` and to 1.20.3 when it is not, and webrtcvad is marked `platform_system != "Windows"`, which is why webrtcvad has to be installed separately as `pip install webrtcvad-wheels` if you need it. Other entries are unpinned, and the list runs from umap-learn, visdom, librosa, sounddevice and SoundFile through espnet, transformers, PyWavelets, torch_complex and numba into the web stack of flask, flask_cors, flask_restx, fastapi, gevent, streamlit and tensorboard. The consequence is that Windows and Linux are not the same project with different shells. Voice activity detection is missing on one platform, numpy resolves to different builds, and the pinned click at 8.0.4 plus pypinyin at 0.44.0 mean a newer dependency resolution will not simply succeed.

env.yml drops monotonic-align and hands you a CPU torch

There are two install paths and they are not equivalent. The conda route is one command:

code
conda env create -n env_name -f env.yml
mamba env create -n env_name -f env.yml

followed by `conda activate env_name`. The stated limitation of that file is that env.yml includes only the necessary dependencies to run the project, temporarily without monotonic-align, and you are told to check the official website to install the GPU version of PyTorch. So the shortcut produces an environment that is missing a package the project requires and a torch build that is not the one the model was trained against. Anyone who takes this path still has to install the GPU PyTorch separately and reconcile monotonic-align, which `requirements.txt` does pin, to `monotonic-align==1.0.0`. The quick install saves a few lines and moves the work, it does not remove it.

Apple silicon means Rosetta, a shim interpreter and a manual header path

The project runs on M1 macOS, but not natively through the GUI. The blocker named is that the PyQt5 packages used in `demo_toolbox.py` are not compatible with M1 chips, and the two offered outs are to forgo `demo_toolbox.py` entirely or to use `web.py`. The rest is a workaround. A virtual environment is created from system Python inside a Rosetta Terminal, and `PyQt5` is installed into it, and the shell then dispatches every command through `arch -x86_64`, including `pip install torch torchvision torchaudio`. The interpreter shim is a file you write yourself:

code
#!/usr/bin/env zsh
mydir=${0:a:h}
/usr/bin/arch -x86_64 $mydir/python "$@"

made executable and then used as the project interpreter, in PyCharm or on the command line as `/PathToMockingBird/venv/bin/pythonM1 demo_toolbox.py`. The cost is that you are emulating x86 for the whole pipeline, not just for the GUI.

pyworld and ctc-segmentation ship no wheels, so headers matter

Two dependencies exist for this project and are not present in the original Real-Time Voice Cloning code it grew from. Neither has a wheel, so pip falls back to compiling C code and the reported failure is that it cannot find `Python.h`. For `pyworld` the fix is `brew install python` and then pointing the compiler at the header directory:

code
export CPLUS_INCLUDE_PATH=/opt/homebrew/Frameworks/Python.framework/Headers

followed by `pip install pyworld`. The same method does not apply to `ctc-segmentation`, which has to be cloned and built by hand with `cythonize -3 ctc_segmentation/ctc_segmentation_dyn.pyx` and then `/usr/bin/arch -x86_64 python setup.py build`. The consequence for a reader is that a clean machine needs a compiler, the Python development headers and a specific include path before pip can do anything useful, and on Apple silicon the build itself has to be x86 too.

The container path reserves GPU zero and serves on port 8080

`docker-compose.yml` defines a single `server` service built from `.`, publishing 8080:8080 and reserving a specific device rather than any GPU:

code
              - driver: nvidia
                device_ids: [ '0' ]
                capabilities: [ gpu ]

State lives in two bind mounts, `./datasets` for corpora and `./synthesizer/saved_models` for checkpoints, and the training behavior is four environment variables: `DATASET_MIRROR=US`, `FORCE_RETRAIN=false`, `TRAIN_DATASETS=aidatatang_200zh magicdata aishell3 data_aishell` and `TRAIN_SKIP_EXISTING=true`. Note that the Dockerfile sets the mirror to `default` while the compose file overrides it to `US`, so the two entry points do not agree. `FORCE_RETRAIN` and `TRAIN_SKIP_EXISTING` are the difference between a restart that reuses checkpoints and one that redoes the work. The image inherits `pytorch/pytorch:latest` and installs ffmpeg, parallel and aria2 at build time.

Editorial conclusion

MockingBird is worth your time if you can train a synthesizer on Mandarin data and you want to read and modify a real-time voice cloning pipeline, because the encoder and vocoder are reusable and the web and streamlit entry points are in the tree. It is the wrong project if you wanted to run a demo, work on Apple silicon without compiling things, or rely on someone still shipping checkpoints. Before you start, confirm you can meet the PyTorch 1.9.0 and cudatoolkit 10.2 pin, budget the training time the demo_cli demands, and check whether a compiler and the Python headers are present, since pyworld and ctc-segmentation have no wheels. The last push is dated 2026-03-03, so expect to own whatever breaks.

Frequently asked questions

Why does the MockingBird demo_cli not work out of the box?

The project reuses the pretrained encoder and vocoder but not the synthesizer, because the original model is incompatible with the Chinese symbols. The stated result is that the demo_cli is not working at this moment and additional synthesizer models are required, so you have to train the synthesizer yourself.

Which Python and PyTorch version does MockingBird expect?

The toolbox needs Python 3.7 or higher, and the recommended environment is repo tag 0.0.1 with PyTorch 1.9.0, Torchvision 0.10.0 and cudatoolkit 10.2. If pip reports that it cannot find a version satisfying torch==1.9.0+cu102, the documented fix is to use Python 3.9 instead.

Does MockingBird run on Apple silicon M1 chips?

It runs on M1 macOS through a workaround rather than natively. The PyQt5 packages used in demo_toolbox.py are not compatible with M1, so you either skip demo_toolbox.py or use web.py, and the environment is built under a Rosetta Terminal with every command run through arch -x86_64.

How do you train the encoder with your own dataset in MockingBird?

Preprocess the audio and mel spectrograms with `python encoder_preprocess.py <datasets_root>`, adding `--dataset {dataset}` to choose among librispeech_other, voxceleb1 and voxceleb2, where only the train split is used and comma separates multiple datasets. Then run `python encoder_train.py my_run <datasets_root>/SV2TTS/encoder`. Training uses visdom, which you start in a separate process, and `--no_visdom` disables it.

What does the MockingBird Docker setup need from the host machine?

The compose file reserves GPU device 0 with the nvidia driver and publishes port 8080. Training behavior is controlled by DATASET_MIRROR, FORCE_RETRAIN, TRAIN_DATASETS and TRAIN_SKIP_EXISTING, and the image installs ffmpeg, parallel and aria2 on top of pytorch/pytorch:latest.

Which datasets does MockingBird list as tested?

Chinese support covers Mandarin, and the listed datasets are aidatatang_200zh, magicdata, aishell3, data_aishell and others. For the optional encoder training the possible dataset names are librispeech_other, voxceleb1 and voxceleb2.

Official sources

  1. Official README
  2. Project repository
  3. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/babysor-mockingbird.svg)](https://hysenlabs.com/projects/babysor-mockingbird)