AB
AiBoss
project

Qwen3-ASR-Flash - A speech recognition model launched by Alitongyi

Qwen3-ASR-Flash is the latest speech recognition model in the Tongyi Qianwen series. Based on the Qwen3 platform, it has been trained using massive amounts of multimodal and ASR data. The model supports 11 languages and multiple accents, and boasts high accuracy and high performance...

What is Qwen3-ASR-Flash?

Qwen3-ASR-Flash is the latest speech recognition model in the Tongyi Qianwen series. Based on the Qwen3 platform, it has been trained using massive amounts of multimodal and ASR data. The model supports 11 languages and various accents, boasting high-precision and robust speech recognition performance, and also supports singing voice recognition. Users can provide text context in any format to obtain customized ASR results. Qwen3-ASR-Flash performs best in multilingual benchmark tests, handling complex acoustic environments and challenging text patterns, providing robust support for speech-to-text services.

Main functions of Qwen3-ASR-Flash

  • High-precision speech recognitionIt performs exceptionally well in speech recognition of multiple languages and dialects, accurately transcribing Mandarin, Sichuanese, Minnan, Wu, Cantonese and other Chinese dialects, as well as various English accents such as British and American, and covering nine other languages including French, German, and Russian.
  • Singing RecognitionSupports singing recognition, including a cappella and full song recognition with background music, with a tested error rate of less than 8%.
  • Customized recognitionUsers can provide background text in any format, such as a list of keywords, paragraphs, or a complete document. The model can intelligently utilize contextual information to identify and match named entities and other key terms, and output customized recognition results.
  • Language recognition and non-human voice rejectionIt supports accurate language identification of speech and automatically filters out non-speech segments, including silence and background noise.
  • High robustnessIt maintains high accuracy when dealing with complex text patterns such as long and difficult sentences, language switching within sentences, and repeated words, as well as complex acoustic environments (such as vehicle noise and various types of noise).

The technical principle of Qwen3-ASR-Flash

  • Based on Qwen3 base modelQwen3-ASR-Flash is built on the Qwen3 base model. The Qwen3 base model is a powerful multimodal pre-trained model capable of processing various types of data (including text, speech, etc.).
  • Training with massive multimodal dataThe model is trained with massive amounts of multimodal data, including various types of data such as text and speech, enabling the model to understand and process information from multiple modalities.
  • Training on ASR data on a scale of tens of millions of hoursIn addition to multimodal data, Qwen3-ASR-Flash is trained using tens of millions of hours of Automatic Speech Recognition (ASR) data. The data covers a variety of languages, dialects, and accents, enabling the model to accurately recognize and transcribe speech.

Qwen3-ASR-Flash project address

  • Project official websiteLink: https://bailian.console.aliyun.com/?spm=5176.29597918.J_tAwMEW-mKC1CPxlfy227s.1.4f007b08aWhTjW&tab=model#/model-market/detail/group-qwen3-asr-flash?modelGroup=group-qwen3-asr-flash
  • Experience the demo onlinehttps://huggingface.co/spaces/Qwen/Qwen3-ASR-Demo

Application Scenarios of Qwen3-ASR-Flash

  • Meeting minutesQwen3-ASR-Flash can transcribe multilingual meeting content in real time, helping to efficiently organize meeting minutes.
  • News interviewAccurate transcription of interview audio improves the timeliness of news reporting.
  • Online EducationThe course audio explanations are transcribed into text to meet the needs of students who speak different languages.
  • Intelligent Customer ServiceIt can be integrated into the customer service system to transcribe customer inquiries in real time, thereby improving service efficiency.
  • Medical recordsAccurately transcribe doctors' voice recordings, facilitating medical record organization and data analysis.