AB
AiBoss
project

Doubao Voice 2.0 - An upgraded AI voice model launched by ByteDance.

Doubao Voice 2.0 is an upgraded AI voice model launched by ByteDance, comprising three core models: Doubao Voice Recognition Model 2.0 (Doubao-Seed-ASR-2.0), Doubao Voice Synthesis Model 2.0 (Doubao-Seed-TTS 2....

What is Doubao Voice 2.0?

Doubao Voice 2.0 is an upgraded AI voice model launched by ByteDance, comprising three core models: Doubao-Seed ASR-2.0 (Doubao-Seed Speech Recognition Model 2.0), Doubao-Seed TTS 2.0 (Doubao-Seed Speech Synthesis Model 2.0), and Doubao-Seed ICL 2.0 (Doubao-Seed Voice Replication Model 2.0). The speech recognition model boasts improved reasoning capabilities, achieving accurate recognition through deep contextual understanding, with an overall 20% improvement in contextual keyword recall. It supports multimodal visual recognition, not only "understanding words" but also "understanding images," making text recognition more accurate through single and multiple image inputs. It also supports accurate recognition of 13 overseas languages, including Japanese, Korean, German, and French. The speech synthesis model 2.0 supports conversational synthesis, accurately understanding semantics and emotions, enabling the reading of complex formulas with an accuracy rate of up to 90%. The voice replication model 2.0 can replicate a voice in just 5 seconds, supports multiple languages, conveys emotions in interactions, and can portray multiple roles. Both have evolved from "speaking like" to "speaking correctly," bringing stronger comprehension and expressiveness to voice interaction, and are widely used in education, novel dubbing and other scenarios. Doubao Voice 2.0 has been officially launched on the Volcano Engine Voice Control Console Experience Center.

Main functions of Doubao Voice 2.0

  • Doubao Speech Recognition Model 2.0 (Doubao-Seed-ASR-2.0):
    • Enhanced reasoning abilityThrough the PPO reinforcement learning approach, the model can deeply understand the context and accurately identify proper nouns, polyphonic characters, etc. without relying on historical vocabulary, improving keyword recall by 20%.
    • Multimodal visual recognitionThe new image understanding capability can be combined with image content (such as single image/multiple images) to assist speech recognition and reduce errors of easily confused words (such as "slippery chicken" and "funny").
    • Multilingual supportWhile maintaining high accuracy in Chinese and English, it now accurately recognizes 13 new languages, including Japanese, Korean, German, and French.
    • Handling complex scenariosFor scenarios such as discussions of historical figures (e.g., identifying the place name "Yunzhou") and image creation (e.g., distinguishing between "horse head" and "wharf"), accuracy is improved through logical reasoning and visual analysis.
    • Technical foundationBased on the Seed hybrid expert large language model architecture, it continues the advantages of the 2 billion parameter audio encoder and focuses on the adaptation of dynamic interactive scenarios.
  • Doubao-Seed-TTS 2.0 (Doubao-Seed-TTS 2.0):
    • Dialogue SynthesisIt supports precise control of voice emotion, tone, and intonation through bracket commands, voice commands, and contextual information, understands the context of multi-turn dialogues, and achieves natural and fluent emotional expression.
    • Reading complex formulasIt is specifically optimized for educational scenarios, covering formulas for all subjects from elementary to high school, with an average accuracy rate of up to 90%, solving the reading difficulties in subject-based assistance.
    • Multi-scenario applicationsIt is widely used in educational assistance, emotional companionship, and content dubbing, making voice more interactive and human-like.
  • Doubao voice replica model 2.0 (Doubao-Seed-ICL 2.0):
    • Fast tone replicationIt can replicate a user's voice in just 5 seconds, supports multiple languages such as Chinese, English, Japanese, Spanish, and Portuguese, and easily achieves "sound similarity".
    • Emotional expressivenessThe replicated voices have stronger emotional expressiveness, can convey emotions that fit the context in interaction, and can play multiple roles.
    • Multi-scenario applicationsSuitable for scenarios such as voice interaction, novel dubbing, and podcast conversations, bringing users a vivid and natural voice experience.

Performance of Doubao Voice 2.0

Doubao Voice 2.0, through specialized optimization, overcomes the challenge of reading complex formulas and symbols in educational tutoring, increasing the average accuracy rate to 90%, significantly higher than the 50% of traditional models, providing a rigorous and efficient voice interaction experience for educational scenarios.

The project address for Doubao Voice 2.0

  • Project official websitehttps://console.volcengine.com/speech/

Application scenarios of Doubao Voice 2.0

  • Educational guidanceIt supports all subjects from elementary to high school, with an average accuracy rate of up to 90%, providing students and teachers with precise voice assistance tools.
  • Emotional companionshipIt accurately expresses emotions based on context and instructions, making voice interaction more realistic and natural, and is suitable for emotional companionship scenarios.
  • Content dubbingAdjusting tone and intonation based on text content, widely used for dubbing videos, advertisements, audiobooks, and other content.
  • Novel InterpretationIt conveys the emotions of different characters according to the context, making it suitable for dubbing novels and bringing the story to life.
  • Podcast ConversationThe model can understand the context of multi-turn dialogues, supports natural and fluent voice interaction, and is suitable for dialogue and interactive segments in podcast programs.