AB
AiBoss
project

Fun-ASR - A large-scale speech recognition model jointly launched by DingTalk and Tongyi

Fun-ASR is a next-generation speech recognition model jointly developed by DingTalk and the Tongyi Labs speech team. Trained on massive amounts of audio data, it can accurately recognize professional terms from more than ten industries, including internet, technology, home decoration, and animal husbandry...

What is Fun-ASR?

Fun-ASR is a next-generation speech recognition model jointly developed by DingTalk and the Tongyi Labs speech team. Trained on massive amounts of audio data, it can accurately recognize professional terms from over ten industries, including internet, technology, home decoration, and animal husbandry, and even understand jargon. For example, in the insurance industry, accuracy has improved by 18% compared to previous models, while in industries like home decoration and animal husbandry, it has achieved improvements of 15%-20%. The model can combine enterprise information within DingTalk for inference optimization, reducing illusion problems and providing more reliable transcription results. Fun-ASR supports customized training for enterprise-specific models, allowing for further algorithm optimization using real enterprise speech data to improve the recognition accuracy of specific vocabulary, and supports the import of up to 1000+ hot words.

Currently, Fun-ASR has been integrated into multiple functional modules of DingTalk, such as meeting captions, intelligent minutes, and voice assistant, providing a stable, efficient, and easily expandable speech recognition solution for enterprise-level contexts.

Tongyi Labs has comprehensively upgraded the core capabilities of Fun-ASR, improving recognition accuracy to 93% in noisy environments. It supports free mixing of 31 languages, lyric and rap recognition, and reduces the first-word latency of the streaming recognition model to 160ms, significantly improving speech recognition performance in complex environments. Furthermore, the lightweight version, Fun-ASR-Nano-0.8B, has been officially open-sourced, with the total number of parameters compressed to 0.8B, resulting in lower inference costs. It also supports local deployment and customized fine-tuning, providing developers with an efficient and flexible speech recognition solution.

Main functions of Fun-ASR

  • Multi-industry terminology recognitionFun-ASR, trained on massive amounts of audio data, can accurately identify professional terms from more than ten industries, including the internet, technology, home decoration, animal husbandry, and automobiles. In actual tests, its accuracy rate in the insurance industry has improved by 18% compared to the past, and in industries such as home decoration and animal husbandry by 15%-20%. It supports the import of up to 1000+ hot words and further optimizes the recognition of rare words.
  • Context-aware optimizationThe model can be optimized by combining enterprise information within DingTalk (such as address book, calendar, knowledge base, etc.), which can effectively alleviate the illusion problem that may occur in large models, provide more reliable transcription results, and requires enterprise authorization to take effect.
  • Customized training for enterprisesBased on an efficient end-to-end training architecture, Fun-ASR can optimize its algorithms based on real-world speech data provided by enterprises, thereby improving the recognition accuracy of proprietary words such as brand names, project codes, product names, and personal names.
  • Multi-scenario integrated applicationFun-ASR has been integrated into multiple functional modules of DingTalk, such as meeting captions and simultaneous interpretation, intelligent minutes, and voice assistant, providing a stable, efficient, and easily expandable speech recognition foundation for enterprise-level contexts and meeting the high requirements of enterprises for speech recognition.

The technical principle of Fun-ASR

  • Training with massive amounts of dataFun-ASR has been trained on hundreds of millions of hours of audio data, covering a variety of industries and scenarios, and can accurately understand professional terms in different fields.
  • Industry co-creation and optimizationBy combining real-world scenarios from multiple DingTalk clients, the model has performed exceptionally well in more than ten fields, including the internet, technology, home decoration, animal husbandry, and automobiles, significantly improving the accuracy of professional terminology recognition.
  • Contextual reasoning optimizationThe model can combine existing information within DingTalk (such as address books, calendars, knowledge bases, etc.) for reasoning optimization, effectively mitigating the illusion problem that may arise from large models and providing more reliable transcription results.
  • End-to-end training architectureBased on an efficient end-to-end training architecture, Fun-ASR can further optimize algorithms using real-world speech data provided by enterprises, improve the recognition accuracy of specific words, and support customized training of enterprise-specific models.
  • Custom hot words supportIt provides enterprises with the ability to customize hot words, supporting the import of up to 1000+ hot words, and further optimizes the recognition of rare words and proprietary terms.

Fun-ASR's project address

  • GitHub repositoryhttps://github.com/FunAudioLLM/Fun-ASR
  • HuggingFace model libraryhttps://huggingface.co/FunAudioLLM/Fun-ASR-Nano-2512

Application scenarios of Fun-ASR

  • Conference subtitles and simultaneous interpretationFun-ASR can transcribe meeting content in real time, providing accurate subtitles and simultaneous interpretation services to help participants better understand and record key points of the meeting.
  • Smart MinutesThe model can automatically generate meeting minutes, extract key information and action items, saving time for manual compilation and improving meeting efficiency.
  • voice assistantIt supports voice commands and interaction, allowing users to complete various operations through voice commands, such as querying information and scheduling, thus improving the user experience.
  • Home decoration and animal husbandry industryIn home improvement companies like KUKA Home, the model can accurately identify professional terms, such as "Belgian imported Pulse latex," providing a reliable basis for subsequent analysis of customer needs. In the livestock industry, it can also accurately identify relevant terminology, helping companies operate efficiently.
  • insurance industryFun-ASR's application in the insurance industry has significantly improved the accuracy of speech recognition, helping insurance companies better handle customer inquiries and business processes.