Tongyi Bailing - A large-scale enterprise-level voice foundation model launched by Alibaba Tongyi
Tongyi Bailing is an enterprise-level speech foundation model launched by Alibaba Tongyi Labs. It integrates Fun-ASR speech recognition and Fun-CosyVoice speech synthesis models, specifically designed for speech applications in complex environments. Through Cont...
What is Tongyi Bailing?
Tongyi Bailing is an enterprise-grade speech foundation model launched by Alibaba Tongyi Lab. It integrates Fun-ASR speech recognition and Fun-CosyVoice speech synthesis models, specifically designed for speech applications in complex environments. Through a context-enhanced architecture, it significantly reduces the illusion rate, solves the language crosstalk problem, and supports dynamic hot word injection and accurate recognition of industry terminology. The model's speech synthesis capability supports cross-language cloning, achieving leading voice similarity. Trained on massive amounts of real audio, it covers multiple industries including finance and education, enabling rapid deployment and helping enterprises efficiently implement speech applications.
Tongyi Bailing's latest upgrade features the Fun-CosyVoice3 model, which boasts a 50% reduction in first-packet latency, doubled accuracy in mixed Chinese and English speech, support for 9 common languages and 18 dialect accents, cross-language cloning and emotion control, and zero-shot voice cloning capabilities for more efficient and natural speech synthesis. Simultaneously, the Fun-ASR model has been significantly enhanced, with recognition accuracy in noisy scenarios reaching 93%, supporting free mixing of 31 languages and dialect accent coverage, and adding lyrics and rap recognition capabilities. The first-word latency for streaming recognition has been reduced to 160ms, making speech recognition more accurate and faster.
The main functions of Tongyi Bailing
-
The rate of hallucinations has dropped significantly.By using the Context Enhancement Architecture (CTC+LLM+RAG), the initial CTC screening results are used as the LLM context, reducing the hallucination rate from 78.5% to 10.7%, resulting in more stable and reliable output.
-
Completely solve the language problemCTC decodes text input LLM Prompt, greatly alleviating the "automatic translation" phenomenon, such as preventing English recordings from being output as Chinese.
-
Strong customization capabilitiesThe RAG mechanism is introduced to dynamically inject terminology into the database, supporting accurate recognition of names, brands, and industry slang (such as "ROI" and "private domain customer acquisition"), with configuration completed in 5 minutes.
-
Cross-language voice cloningBased on a multi-stage training method, a single timbre can support multiple languages, achieving industry-leading voice similarity.
-
Full coverage of industry scenariosBased on tens of millions of hours of real audio training, it covers more than 10 industries including finance, education, manufacturing, internet, and animal husbandry, and goes deep into the front line of industry.
The technical principles of Tongyi Bailing
- Fun-ASR Speech Recognition Large ModelBased on Bailing's Fun-ASR speech recognition model, an innovative Context Enhancement Architecture (CTC+LLM+RAG) is employed. CTC technology performs initial speech-to-text conversion, while LLM optimizes the generated text for context, significantly reducing the illusion rate from 78.5% to 10.7%, resulting in more stable and reliable output. A terminology database is dynamically injected based on the RAG mechanism, supporting accurate recognition of names, brands, industry jargon, etc., with configuration completed within 5 minutes to meet the personalized needs of different enterprises.
- Fun-CosyVoice Large-Scale Speech Synthesis ModelFun-CosyVoice's large-scale speech synthesis model is based on an innovative speech decoupling training method. It separates and independently trains features such as timbre, speech rate, and intonation, then combines them to generate high-quality speech, making the synthesized speech more natural and fluent. The model supports cross-language speech cloning; through a multi-stage training method, a single timbre can support multiple languages, achieving "one timbre speaks globally," with industry-leading voice similarity.
Tongyi Bailing Project Address
- Project official websiteFun-ASR, Fun-CosyVoice
Application scenarios of Tongyi Bailing
-
Financial industryIt can be used in intelligent customer service, voice transactions and risk monitoring to improve service efficiency and risk control capabilities.
-
Education industryIt supports online education platforms, intelligent tutoring systems, and voice-based homework grading, optimizing the teaching and learning experience.
-
manufacturingIt enables voice control of industrial equipment, production process monitoring, and quality inspection, thereby improving production efficiency and safety.
-
Internet industryIt supports voice search, intelligent assistants, and content creation, enhancing user experience and content diversity.
-
Livestock industryIt is applied in intelligent breeding systems, animal health monitoring, and breeding environment management to improve breeding efficiency and animal health management.