Skip to main content
📖 The AI Tool Bible
MockingBird preview image
MockingBird logo

MockingBird

Open-source Mandarin-first voice cloning that mimics a speaker from a 5-second sample.

Free· Free, open source (MIT)AudioGE2E + Tacotron + HiFi-GAN/WaveRNN/Fre-GAN7.0 / 10

In short

MockingBird is a free, open-source Mandarin voice cloning tool that mimics speakers from 5-second samples. It runs locally on PyTorch, offering a self-hosted alternative to paid SaaS APIs.

Best for

Pick MockingBird if you need an open, self-hosted Mandarin voice cloning pipeline you can run locally without paying SaaS rates.

Skip if

Skip it if you want a maintained, plug-and-play English TTS API or a polished GUI product with vendor support.

MockingBird is an MIT-licensed voice cloning toolkit that captures the timbre of a target speaker from roughly five seconds of audio and then synthesizes arbitrary speech in that voice. It bundles a GE2E speaker encoder, a Tacotron-based synthesizer, and a choice of WaveRNN, HiFi-GAN, or Fre-GAN vocoders, with pretrained checkpoints for Mandarin and tooling to train your own on datasets like aidatatang_200zh, magicdata, and aishell3.

The project is aimed at researchers, hobbyists, and developers who want a self-hostable text-to-speech / voice-conversion pipeline without paying per-character API fees. It runs on Windows, Linux, and Apple Silicon, ships both a Qt desktop toolbox and a web.py server, and is one of the few high-quality open Mandarin TTS stacks. The original author has stepped back from active development and points commercial users to their hosted successor at noiz.ai, but the repo remains widely forked and usable.

PyTorch 1.9+ is required and you will need a GPU to train from scratch; inference works on CPU but is slow. There is no official REST API or SaaS layer, so integration means wrapping the Python code yourself. Pretrained weights are community-hosted on Google Drive and Baidu Pan, which makes setup fiddlier than a pip install.

Editor's take

MockingBird is a landmark open-source voice cloning project, especially for Mandarin, and the code still works once you wrestle the dependencies into shape. With the original author pointing commercial users to noiz.ai, treat this as a research-grade starting point rather than a production-ready tool.

— The AI Tool Bible editorial team

Pros

  • ✅ One of the strongest open-source Mandarin voice cloning stacks
  • ✅ MIT licensed, fully self-hostable with no per-call costs
  • ✅ Works on Windows, Linux, and Apple Silicon
  • ✅ Multiple vocoder choices and pretrained checkpoints included

Cons

  • ⚠️ Original author no longer actively maintains the repo
  • ⚠️ Mandarin-first; English and other languages need DIY training
  • ⚠️ Setup is fiddly: PyTorch, GPU, and external weight downloads required
  • ⚠️ No hosted API; commercial successor noiz.ai is a separate product

Use cases

voice-cloningtext-to-speechmandarin-ttsvoice-conversion

Frequently asked

How much does MockingBird cost?
MockingBird is completely free and open source under the MIT license. There are no per-character API fees or subscription costs. You can run it locally on your own hardware without paying for SaaS services.
Is MockingBird suitable for English text-to-speech?
No, MockingBird is designed primarily for Mandarin. The documentation advises skipping it if you want a maintained, plug-and-play English TTS API. It is one of the few high-quality open Mandarin TTS stacks available.
What are the system requirements for running MockingBird?
You need PyTorch 1.9+ and a GPU to train from scratch. Inference works on CPU but is slow. It supports Windows, Linux, and Apple Silicon. Pretrained weights are hosted on Google Drive and Baidu Pan, making setup slightly fiddlier than a simple pip install.
Does MockingBird offer a REST API or SaaS layer?
No, there is no official REST API or SaaS layer. Integration requires you to wrap the Python code yourself. The project ships with a Qt desktop toolbox and a web.py server, but commercial users are pointed to a hosted successor at noiz.ai.
How much audio is needed to clone a voice?
MockingBird captures the timbre of a target speaker from roughly five seconds of audio. It uses a GE2E speaker encoder and Tacotron-based synthesizer to generate new speech in that specific voice.

Explore related

Compare with similar tools

All in Audio →

Reviews