AB
AiBoss
project

MaskGCT - A large-scale speech synthesis model developed by FunWan Technology in collaboration with the Chinese University of Hong Kong.

MaskGCT is a large-scale speech synthesis model developed by FunWan Technology in collaboration with the Chinese University of Hong Kong, Shenzhen. Based on a mask generation model and decoupled encoding of speech representation, it enables applications such as voice cloning, cross-language synthesis, and voice control.

What is MaskGCT?

MaskGCT is a large-scale speech synthesis model developed by FunWan Technology in collaboration with the Chinese University of Hong Kong, Shenzhen. Based on a mask generation model and decoupled encoding of speech representation, it achieves significant results in tasks such as voice cloning, cross-language synthesis, and voice control. The model has achieved industry-leading levels on multiple TTS benchmark datasets, with some performance metrics even surpassing human capabilities. MaskGCT can quickly and realistically clone voices, flexibly adjust the duration, speed, and emotion of speech, and supports the synthesis of six languages: Chinese, English, Japanese, Korean, French, and German. The model is open-source on the Amphion system and is available to users worldwide.

Main functions of MaskGCT

  • Sound cloningIt can quickly replicate any timbre, including human and anime characters, and can completely reproduce intonation, style and emotion.
  • Cross-language synthesisIt supports speech synthesis in multiple languages, including Chinese, English, Japanese, Korean, French, and German, enabling cross-language speech generation.
  • Voice controlIt allows for flexible adjustment of the length, speed, and emotion of the generated speech, and supports editing the speech content with editable text, maintaining consistency in rhythm and timbre.
  • High-quality speech datasetsTrained on the high-quality multilingual speech dataset Emilia, it provides a wealth of speech synthesis materials.

MaskGCT technology principle

  • Speech-Semantic Representation CodecThe speech is converted into semantic tags, and a vector quantization codebook is learned using the VQ-VAE model to reconstruct the speech semantic representation from the speech self-supervised learning model.
  • Speech Acoustic CodecThe speech waveform is quantized into multi-layer discrete markers to retain all speech information. The speech waveform is compressed using the RVQ method, and the Vocos architecture is used as the decoder.
  • Text to Semantic ModelIt generates Transformers using non-autoregressive masks, without relying on text-to-speech alignment information, and predicts semantic tags based on the context learning ability of language models.
  • Semantic to acoustic modelA Transformer is generated using a non-autoregressive mask, and a multi-layer acoustic tag sequence is generated conditionally using semantic tags to reconstruct a high-quality speech waveform.

MaskGCT's project address

Application scenarios of MaskGCT

  • audiobooks and podcastsHigh-quality speech generated using MaskGCT provides natural reading voices for ebooks, audiobooks, and podcasts, enhancing the listener's auditory experience.
  • Smart assistants and chatbotsMaskGCT provides a more natural and personalized voice interaction experience in smart devices and customer service systems.
  • Video games and virtual realityIn gaming and virtual reality applications, MaskGCT generates realistic voices for characters, enhancing immersion.
  • Film and television production and dubbingIn film and television post-production, MaskGCT can quickly generate or replace character voices, improving production efficiency.
  • Language learning and educationMaskGCT generates standard or accented speech to help language learners practice pronunciation and listening comprehension.