AB
AiBoss
project

FunAudioLLM - An open-source speech model developed by Alibaba's Tongyi team.

FunAudioLLM is an open-source speech model project launched by Alibaba Tongyi Labs, comprising two models: SenseVoice and CosyVoice. SenseVoice excels in multilingual speech recognition and emotion detection, supporting over 50 languages...

What is FunAudioLLM?

FunAudioLLM is an open-source speech model project launched by Alibaba Tongyi Labs, comprising two models: SenseVoice and CosyVoice. SenseVoice excels in multilingual speech recognition and emotion detection, supporting over 50 languages, with particularly outstanding performance in Mandarin and Cantonese. CosyVoice focuses on natural speech generation, allowing control over timbre and emotion, and supports five languages: Chinese, English, Japanese, Cantonese, and Korean. FunAudioLLM is suitable for scenarios such as multilingual translation and emotion-based voice dialogue. The related models and code are open-sourced on the Modelscope and Huggingface platforms.

Main functions of FunAudioLLM

  • SenseVoice model:
    • Focusing on high-precision speech recognition in multiple languages.
    • It supports more than 50 languages, and its recognition performance is superior to existing models, especially in Chinese and Cantonese.
    • It has emotion recognition capabilities and can identify a variety of human-computer interaction events.
    • It offers both lightweight and full-featured versions to suit different application scenarios.
  • CosyVoice model:
    • It focuses on natural speech generation and supports multiple languages, timbre, and emotion control.
    • It can quickly generate analog timbres from a small amount of raw audio, including rhythmic and emotional details.
    • It supports cross-language speech generation and fine-grained emotion control.

FunAudioLLM's project address

Application scenarios of FunAudioLLM

  • Developers and researchers: Research and development using FunAudioLLM in the fields of speech recognition, speech synthesis, and emotion analysis.
  • Enterprise usersFunAudioLLM can be applied in business scenarios such as customer service, intelligent assistants, and multilingual translation to improve efficiency and user experience.
  • Content creatorsUse FunAudioLLM to generate audiobooks or podcasts, enriching content formats and attracting more listeners.
  • EducationUsed for educational applications such as language learning and listening training to improve learning efficiency and interest.
  • People with disabilitiesIt helps visually impaired people obtain information through voice interaction, improving their quality of life.