AB
AiBoss
project

Fun-ASR1.5 - An end-to-end speech recognition model launched by Alitongyi

Fun-ASR 1.5 is a new generation version of the end-to-end speech recognition model launched by the Ali Tongyi team. A single model supports high-precision recognition of 30 languages, covering seven major Chinese dialect systems and more than twenty regional accents, with specific optimizations for classical Chinese poetry...

What is Fun-ASR1.5?

Fun-ASR 1.5 is a new generation version of the end-to-end speech recognition model launched by the Alibaba Tongyi team. A single model supports high-precision recognition of 30 languages, covering seven major Chinese dialect systems and over twenty regional accents, with specific optimizations for recognizing classical Chinese poetry recitation. The model is based on the MoE architecture to achieve automatic language switching without the need for pre-set labels. Fun-ASR 1.5 enables intelligent punctuation prediction and text normalization in post-processing, making speech-to-text transcription not just usable but also highly user-friendly.

Main functions of Fun-ASR1.5

  • Multilingual recognitionThe single model covers 30 languages including Chinese, English, Japanese, Korean, French, German, Spanish, Portuguese, Russian, and Arabic.
  • Automatic language switchingNo language tags need to be preset; it automatically recognizes and switches between multilingual mixed speech in code-switching scenarios.
  • Dialect recognitionIt covers seven major dialect systems and more than 20 local accents, with a focus on optimizing 15 high-demand dialects.
  • Ancient Poetry Recognition: Construct a corpus of ancient Chinese poetry from the pre-Qin period to modern times with voice-text alignment, supporting accurate transcription of classical Chinese readings.
  • Intelligent punctuation predictionAutomatically insert punctuation marks such as commas, periods, and question marks based on contextual semantics.
  • Text normalizationAutomatically converts spoken numbers, dates, amounts, telephone numbers, etc., into standard written format.

Technical Principles of Fun-ASR1.5

  • MoE architectureIt adopts a hybrid expert architecture, which activates only the relevant parts for processing when a specific language is heard, thereby improving the flexibility and efficiency of multilingual processing.
  • Graded and phased training: Use precise data in a graded and phased manner during the training phase to improve the ability to cope with complex real-world speech scenarios.
  • Dialect Data DrivenBased on training with hundreds of thousands of hours of real dialect speech data, the average word error rate (CER) has decreased by 56.2% compared to the previous version.
  • Ancient Chinese Poetry Corpus: Construct a corpus of real-person recitation recordings covering classic texts such as the Book of Songs, the Songs of Chu, the poetry collections of Li Bai and Du Fu, and the lyrics of Su Shi and Xin Qiji.

How to use Fun-ASR1.5

  • Alibaba Cloud Hundred Refinement PlatformVisit the Alibaba Cloud Bailian official website and enter the voice section of the Model Experience Center to call the API.
  • Moda Community:Visit https://modelscope.cn/studios/iic/FunAudio-ASR to experience it online directly.

Key information and usage requirements of Fun-ASR1.5

  • Product Positioning: End-to-end speech recognition large model.
  • Supported languages30 languages (covering major languages in Europe, East Asia, Southeast Asia, South Asia and the Middle East).
  • Dialect coverageSeven major dialect systems, with a focus on optimizing 15 high-demand dialects such as Shanghainese, Cantonese, and Sichuanese.
  • Accuracy of Classical PoetryThe internal evaluation set achieved a character-level accuracy of 97%.
  • How to useAPI calls or online experiences.
  • No preset requiredIn multilingual scenarios, there is no need to specify language tags in advance.

The core advantages of Fun-ASR1.5

  • Single Model MultilingualA single model can seamlessly switch between 30 languages, reducing the cost of deploying and maintaining multiple models.
  • Leading in dialect recognitionBased on hundreds of thousands of hours of dialect data, CER has decreased by 56.2% compared to the previous version, supporting the restoration of authentic dialect text.
  • Automatic Code SwitchingIt can handle multilingual scenarios in the same dialogue without any presets.
  • Cultural Scene OptimizationSpecialized training in reciting classical Chinese poems has achieved a character accuracy rate of 97%, contributing to cultural heritage preservation.
  • Intelligent post-processingAutomatic punctuation and text normalization significantly reduce the post-editing costs of meeting minutes, legal transcripts, and other similar documents.

Comparison of Fun-ASR1.5 with similar competing products

Dimension Fun-ASR1.5 Seed-ASR Tencent-ASR
Language coverage 30 languages, single model coverage Multilingual support Multilingual support
Dialect support Seven major dialect systems, 15 key optimizations, CER reduced by 56.2% Basic support Basic support
Code-Switching No need to preset tags, automatic recognition and switching support support
Ancient Poetry Recognition Specialized optimization, 97% character accuracy Unclear Unclear
Intelligent post-processing Automatic punctuation + text normalization (numbers/dates/amounts/telephone numbers) Basic punctuation skills Basic punctuation skills
Architectural features MoE Hybrid Expert Architecture Not disclosed Not disclosed
Open Experience Alibaba Cloud's Hundred-Refined API + Magic Community Volcano Engine Tencent Cloud

Application Scenarios of Fun-ASR1.5

  • transnational conferencesIn multinational conference scenarios, Fun-ASR1.5 can accurately transcribe multilingual mixed dialogues in real time, without requiring participants to preset languages or switch between multiple translation tools.
  • smart speakerIn smart home and in-vehicle voice interaction scenarios, Fun-ASR1.5 can accurately recognize various dialect commands, allowing smart speakers to truly "understand local accents".
  • Online EducationIn the context of online education of traditional Chinese culture, Fun-ASR1.5 supports accurate transcription of ancient poems and lyrics, contributing to the digital inheritance of traditional culture with a character-level accuracy of 97%.
  • News interviewIn news interviewing and content production scenarios, Fun-ASR1.5 can automatically add punctuation marks and normalize spoken numbers and dates into standardized formats, significantly reducing the time spent on manual editing later.