AB
AiBoss
project

Hy ASR 3.0 preview - Tencent Hunyuan's next-generation speech recognition model

Hy ASR 3.0 preview is a next-generation speech recognition model launched by Tencent Hunyuan. Based on the Hy3 large language model, it integrates high-precision speech recognition with deep semantic understanding. The model supports Mandarin, English, Cantonese, and 10 major dialect regions...

What is Hy ASR 3.0 preview?

Hy ASR 3.0 preview is a next-generation speech recognition model launched by Tencent Hunyuan. Based on the Hy3 large language model, it integrates high-precision speech recognition and deep semantic understanding. The model supports Chinese, English, Cantonese, and 10 major dialect regions, and has robustness in complex scenarios such as context-sensitive intelligent error correction, hot word injection, and high noise/whispering. The WER (Warnings Evidence) on the open-source evaluation set is controlled at around 3%, and it outperforms competitors across the board on the self-built evaluation set.

Key features of Hy ASR 3.0 preview

  • General and accurate recognition: It supports Mandarin Chinese, English, Cantonese, and mixed Chinese and English, reducing the accumulation of typos, omissions, and errors in long audio clips.
  • Contextual semantic understanding: By combining context, it accurately captures the user's context, intelligently corrects homophones, and eliminates semantic ambiguity.
  • Hot keyword injection adaptation: It supports enhancement of popular keywords such as brand names, personal names, and industry terms, reducing business access and maintenance costs.
  • Robust in complex environments: Specifically optimized for acoustic scenarios such as high noise, whispers, and private conversations, maintaining stable recognition performance.
  • Dialects are widely used: It supports 10 major dialect areas and more than 20 secondary sub-regions, including Northeastern Mandarin, Cantonese, Wu dialect, and Minnan dialect.

Technical principles of Hy ASR 3.0 preview

  • Architecture upgrade: It adopts the MoE architecture and upgrades the base model to Hy3; it has developed its own unsupervised speech encoder, which is trained on tens of millions of hours of speech data to extract high-quality acoustic representations.
  • Pre-training Scaling: The Encoder is trained in conjunction with LLM, incorporating tens of millions of hours of multi-source speech data (audiobooks, dialogues, videos/broadcasts, etc.), and acquires dialect understanding and context modeling capabilities through multi-stage capability injection.
  • Post-training enhancement: Build a high-quality SFT Recipe covering complex audio, Any Context, and general capabilities; optimize general transcription accuracy, long-range context dependency, and performance in complex long-tail scenarios through multi-stage reinforcement learning.

How to use Hy ASR 3.0 preview

You can experience it for free by accessing the Tencent Cloud API or by opening Tencent Yuanbao and holding down the button to speak.

The core advantages of Hy ASR 3.0 preview

  • Semantic understanding upgrade: Based on the Hy3 large language model, it integrates high-precision speech recognition and deep semantic understanding, evolving from word-for-word transcription to understanding context and outputting directly with one click.
  • Leading in all aspects of the evaluation: The open-source evaluation set has a WER of around 3%, while the self-built evaluation set outperforms competitors in general, dialect, contextual, and complex acoustic scenarios.
  • Context-based intelligent error correction: It combines context to accurately capture the context, intelligently corrects homophones, eliminates semantic ambiguity, and supports long-term dependency modeling.
  • Quickly adapt to trending keywords: It supports the injection of popular keywords such as brand names, personal names, and industry terms, reducing the cost of business access and maintenance in professional scenarios.
  • Coverage of dialects and complex environments: It supports 10 major dialect regions and more than 20 secondary sub-regions, maintaining stable performance in complex acoustic environments such as high noise and whispers.

Hy ASR 3.0 preview vs. competing products

Comparison Dimensions Hy ASR 3.0 preview (Tencent Hunyuan) Doubao-Seed-ASR 2.0 (ByteDance)
Base model Based on the Hy3 large language model, the MoE architecture Based on Seed Large Language Model
Voice Encoder Our self-developed unsupervised speech encoder was trained on tens of millions of hours of data. Detailed architecture not disclosed
Core competencies Contextual semantic understanding + high-precision recognition + hot word injection High-precision universal recognition
Chinese WER 3.34% 3.70%
English WER 2.62% 5.65%
Cantonese WER 3.12% 5.25%

Application scenarios of Hy ASR 3.0 preview

  • Intelligent Customer Service: In telephone and online customer service scenarios, Hy ASR 3.0 preview accurately identifies brand names and business terms through hot word injection, enabling real-time speech transcription and significantly improving service response efficiency.
  • Content creation: In the production of long audio content such as podcasts, conferences, and interviews, the model combines contextual semantic understanding to intelligently eliminate ambiguity and output coherent and readable high-quality text with one click, reducing post-editing costs.
  • Voice search: It supports users to use dialects for queries and accurately recognizes voice commands in noisy environments, effectively reducing the input threshold for users with multiple languages and accents.
  • Office collaboration: In real-time meeting minutes transcription scenarios, the model maintains stable recognition even in the face of low-volume input such as whispers and private conversations, helping teams achieve efficient communication and information retention.
  • Smart terminals: In interactions with IoT devices such as in-vehicle navigation and smart homes, the model is specifically optimized for high-noise environments to ensure robust speech recognition and user experience in complex scenarios.