AB
AiBoss
project

TicVoice 7.0 - Mobvoi's seventh-generation voice synthesis engine

TicVoice 7.0 is Mobvoi's seventh-generation high-quality TTS (text-to-speech) engine, based on the next-generation speech generation model Spark-TTS. TicVoice 7.0 utilizes the innovative BiCodec encoding method to segment speech...

What is TicVoice 7.0?

TicVoice 7.0 is Mobvoi's seventh-generation high-quality TTS (Text-to-Speech) engine, based on the next-generation Spark-TTS speech generation model. TicVoice 7.0 utilizes an innovative BiCodec encoding method to decompose speech into Global Tokens and Semantic Tokens, achieving precise control over timbre and semantics, and maintaining a high degree of consistency with the text LLMs structure. The engine features 3-second voice cloning capability, supports multiple roles, multiple emotions, all age groups, and Chinese/English switching, delivering natural and fluent voices approaching broadcast quality. TicVoice 7.0 is already available in the "3-second voice cloning" function of Mobvoi Workshop, widely applicable to fields such as intelligent customer service, audiobooks, and film dubbing, bringing users an ultimate AI dubbing experience.

Main features of TicVoice 7.0

  •  3-second voice cloningCaptures user voiceprint in 3 seconds, accurately replicates personalized tone, and supports low-quality audio input.
  • Multiple roles and multiple emotions in the performanceIt supports simulation of various emotions such as happiness, anger, and sadness, enhancing the expressiveness of the content.
  • Voice adaptation for all agesIt covers a diverse range of sounds from children to the elderly, meeting the needs of different scenarios.
  • Flexible switching between Chinese and EnglishSupports mixed Chinese and English speech synthesis, facilitating the creation of multilingual content.
  • Broadcast-quality voiceThe synthesized speech is clear, fluent, natural, and pleasant to listen to, with strong timbre and emotional expression, approaching the level of professional broadcasting.
  • Customized exclusive voiceUsers can customize their own voice to meet their personalized dubbing needs.

Technical Principles of TicVoice 7.0

  • Innovative speech coding methodsBased on BiCodec technology, speech is decomposed into Global Tokens (global features, such as timbre) and Semantic Tokens (semantically related features, 50 tokens/second), balancing global controllability and semantic relevance. This solves the problems of traditional speech coding, such as the difficulty in accurately controlling timbre with semantic tokens and the reliance on multiple codebooks for acoustic coding.
  • Unified with text LLMs structureReusing the architecture of Qwen2.5, based on attribute labels (such as gender, fundamental frequency level) and fine-grained attribute values (such as precise fundamental frequency), it uses text + attribute labels as input to predict fine-grained attribute values → Global Tokens → Semantic Tokens in sequence. This achieves a high degree of consistency between voice token modeling and text token modeling.
  • Single-stage, single-stream generationTTS generation is achieved in a single-stage, single-stream manner using a language model (sequence monkey), without the need for additional generation models, thus improving generation efficiency and controllability.
  • Deep learning-based speech synthesisBased on deep learning technology and combined with a large amount of speech data to train the model, a natural and fluent speech synthesis effect is achieved.

TicVoice 7.0 project address

  • Project official websiteMagic Sound Workshop

Application scenarios of TicVoice 7.0

  • Intelligent Customer ServiceProvides online customer service systems with natural and fluent voice interaction capabilities, improving user experience and reducing labor costs.
  • Audiobooks and podcastsIt can quickly generate high-quality audiobooks and podcasts, support multiple characters and emotional expressions, and enhance the listener's immersion.
  • Film and television dubbing and narrationIt efficiently completes dubbing and narration work for movies and short videos, supports multi-language switching, and reduces production costs.
  • Emotional live streaming and interactionSimulating real emotions during live streams enhances interaction between streamers and viewers, thereby increasing the appeal of the content.
  • Education and TrainingIt provides vivid audio teaching content for online education platforms, supports multiple languages and multiple roles, and enhances the learning experience.