AB
AiBoss
project

Ming-omni-tts - Ant Group's open-source unified audio generation model

Ming-omni-tts is an open-source unified audio generation model from Ant Group, employing an autoregressive architecture to jointly generate speech, music, and sound effects. The model supports natural language commands for adjusting speech rate, pitch, volume, emotion, and other parameters...

What is Ming-omni-tts?

Ming-omni-tts is an open-source unified audio generation model from Ant Group, employing an autoregressive architecture to jointly generate speech, music, and sound effects. The model supports fine-grained control over speech rate, pitch, volume, emotion, and dialect via natural language commands, achieving a 93% accuracy rate for Cantonese dialect control and a 46.7% accuracy rate for emotion control, surpassing CosyVoice3. Technically, it utilizes a unified continuous audio tokenizer and Diffusion Transformer architecture, processing multimodal audio at a 12.5Hz frame rate. A "Patch-by-Patch" compression strategy reduces the LLM inference frame rate to 3.1Hz, maintaining sound quality while reducing latency. The 16.8B parameter version achieved a WER of only 0.83% on the Seed-tts-eval Chinese test set, surpassing SeedTTS and GLM-TTS. The model includes over 100 high-quality timbres, supports zero-sample sound design, and provides Docker images and Grado demos, making it suitable for audiobooks, podcasts, and multilingual content creation.

Main functions of Ming-omni-tts

  • Unified Multimodal Audio GenerationThe industry's first autoregressive model can jointly generate speech, ambient sound, and music in a single channel, achieving an immersive auditory experience.
  • Fine-grained voice controlIt supports precise control of speech rate, tone, volume, emotion, and dialect through simple commands. The accuracy rate of Cantonese dialect control is as high as 93%, and the accuracy rate of emotion control is 46.7%.
  • Intelligent sound designIt has 100+ high-quality built-in sounds and supports zero-sample sound design through natural language description.
  • Efficient Reasoning OptimizationThe "Patch-by-Patch" compression strategy is adopted to reduce the LLM inference frame rate to 3.1Hz, significantly reducing latency.
  • Professional text normalizationIt accurately parses and reads aloud complex mathematical expressions, chemical equations, and other professional formats, with an internal test set CER of only 1.97%.
  • Multilingual supportIt supports speech synthesis and cross-language transfer in multiple languages, including Chinese and English.
  • Zero-sample TTSAny timbre can be cloned with just 3-10 seconds of reference audio, and the WER is as low as 0.83% on Seed-tts-eval.

The technical principle of Ming-omni-tts

  • Unified Continuous Audio TokenizerA VAE-based continuous tokenizer integrates speech, music, and general audio into a unified latent space at a 12.5Hz frame rate, supporting joint modeling of multimodal audio.
  • Diffusion Transformer (DiT) HeadThe diffuser architecture enhances the quality of generated audio, improving the delicacy and naturalness of the sound.
  • Patch generation strategyThe generation strategy employs a patch size of 4 and a backtracking history of 32 to achieve a balance between local acoustic details and long-term structural coherence.
  • Autoregressive Generative ArchitectureThe industry's first autoregressive model that jointly generates speech, music, and sound effects in a single channel, achieving unified audio generation.
  • "Patch-by-Patch" compression mechanismBy using a compression strategy, the LLM inference frame rate is reduced from the original frequency to 3.1Hz, significantly reducing computational latency and inference costs.
  • Instructions for fine-tuning alignmentIt enables fine-grained control over speech rate, tone, volume, emotion, and dialect through command-based fine-tuning, and supports natural language command parsing.

Ming-omni-tts's project address

  • GitHub repository:https://github.com/inclusionAI/Ming-omni-tts
  • Hugging Face Model Library:
    • https://modelscope.cn/models/inclusionAI/Ming-omni-tts-16.8B-A3B
    • https://huggingface.co/inclusionAI/Ming-omni-tts-0.5B

Application scenarios of Ming-omni-tts

  • Audiobook and podcast productionIt supports long text speech synthesis, with a CER of only 1.84% for podcast TTS tasks, making it suitable for generating audiobooks, news broadcasts, and podcast content.
  • Multilingual content creationIt supports speech synthesis in multiple languages, including Chinese and English, and cross-language voice transfer to meet the needs of global content production.
  • Game sound designIt can jointly generate voice, ambient sounds, and music to provide an immersive audio experience for game scenes.
  • Education and training sectorIt can accurately read aloud complex mathematical expressions, chemical equations, and other professional content, and is suitable for online educational courseware and academic explanations.
  • Intelligent Customer Service and AssistantIt features 100+ high-quality built-in voices, supports zero-sample sound cloning, and allows for quick customization of brand-specific voice assistants.
  • Advertising and Marketing Voiceover: Generate compelling advertising voiceovers and localized marketing content through emotional control and dialect support.