AB
AiBoss
project

Qwen3-TTS-Flash - A text-to-speech model launched by Alitongyi

Qwen3-TTS-Flash is a flagship speech synthesis model launched by Alitongyi, supporting multiple timbres, languages, and dialects. The model boasts excellent stability in both Chinese and English pronunciation, outstanding multilingual performance, and highly expressive human-like timbres...

What is Qwen3-TTS-Flash?

Qwen3-TTS-Flash is a flagship speech synthesis model launched by Alibaba Cloud Tongyi, supporting multiple timbres, languages, and dialects. The model boasts exceptional stability in both Chinese and English pronunciation, outstanding multilingual performance, and highly expressive human-like timbres. It offers 49 high-fidelity timbres with distinct personalities (from lively and playful to calm and wise), and supports 10 mainstream languages and 9 Chinese dialects (including authentic Tianjin and Sichuan dialects), truly achieving personalized "one voice for one person" needs. Compared to its predecessor, Qwen3-TTS has made significant breakthroughs in speech naturalness, intelligently adjusting speech rate and rhythm to make the synthesized speech more "human"—both emotional fluctuations and speech rhythm are closer to real human expression. As another masterpiece in the open-source community, it is particularly suitable for scenarios such as virtual characters, content creation, and AI assistants. Users can quickly access it through the Alibaba Cloud Bailian platform API or experience its excellent speech synthesis effects online in communities such as Hugging Face.

Main functions of Qwen3-TTS-Flash

  • Highly anthropomorphicIt features naturalness, multiple timbres, and support for multiple languages/dialects, making the voice more like a real person in terms of speech speed, rhythm, and emotion.
  • Rich sound library:supply 49 high-fidelity tonesIt covers a variety of styles (such as lively, arrogant, steady, and anime), and is suitable for different scenarios (such as short videos, virtual characters, and knowledge explanations).
  • Extensive language support:cover 10 languages(Chinese, English, German, French, Spanish, Italian, Portuguese, Japanese, Korean, Russian) and 9 Chinese dialects(Such as Cantonese, Sichuanese, Tianjin dialect, etc.), the dialects are restored to be authentic and natural.
  • High expressivenessThe generated speech is natural and expressive, and can automatically adjust the tone according to the input text to make the speech more vivid.
  • High robustnessIt supports automatic processing of complex text, extracts key information, and is highly adaptable to complex and diverse text formats.
  • Quick generationIt features extremely low first-packet latency (as low as 97ms), enabling rapid speech generation and enhancing the user experience.
  • High similarity in timbreIt performs exceptionally well in multilingual speech stability and timbre similarity, surpassing other similar models.

The technical principle of Qwen3-TTS-Flash

  • Deep learning models:
    • Text encoder: Convert the input text into a semantic representation and extract key information and semantic features from the text.
    • Voice decoderIt generates speech waveforms based on the output of the text encoder, ensuring the naturalness and expressiveness of the speech.
    • Attention mechanismThrough the attention mechanism, the model can better align text and speech, improving the accuracy and fluency of the generated speech.
  • Multilingual and multi-dialect supportThe model is trained on data from multiple languages and dialects, learning the pronunciation characteristics and intonation patterns of different languages and dialects. Through timbre embedding technology, the model can generate speech with different timbres to meet diverse user needs.
  • High robustnessThe input text undergoes preprocessing, including word segmentation, part-of-speech tagging, and semantic parsing, to ensure the model can correctly understand the text content. The model has the ability to automatically process complex and erroneous text, extract key information, and generate accurate speech.
  • Technological advancementSignificantly optimized in prosody control, it can automatically adjust speech rate and intonation according to text content, and outperforms some mainstream competitors (such as MiniMax, ElevenLabs, and GPT-4o Audio Preview) in multilingual testing (WER index).

Performance of Qwen3-TTS-Flash

  • Chinese and English speech stabilityOn the seed-tts-eval test set, Qwen3-TTS-Flash achieved state-of-the-art (SOTA) performance in both Chinese and English speech stability, surpassing SeedTTS, MiniMax, and GPT-4o-Audio-Preview.
  • Multilingual speech stabilityOn the MiniMax TTS multilingual test set, Qwen3-TTS-Flash achieved state-of-the-art (SOTA) performance in WER for Chinese, English, Italian, and French, significantly lower than MiniMax, ElevenLabs, and GPT-4o-Audio-Preview.
  • Timbre SimilarityIn terms of speaker similarity in English, Italian, and French, the Qwen3-TTS-Flash outperforms MiniMax, ElevenLabs, and GPT-4o-Audio-Preview, demonstrating superior tonal performance.

Qwen3-TTS-Flash project address

  • Project official website: https://qwen.ai/blog?id=qwen3-tts-1128
  • Experience the demo onlinehttps://huggingface.co/spaces/Qwen/Qwen3-TTS-Demo

Application scenarios of Qwen3-TTS-Flash

  • Intelligent Customer ServiceProvide users with natural and fluent voice interaction to enhance the service experience, such as automatically answering common questions and guiding user operations.
  • audiobooksIt transforms text into vivid audio, allowing listeners to enjoy the pleasure of listening to books. It is suitable for various content such as novels, news, and textbooks.
  • voice assistantIn smart home devices, smart wearables, and other devices, voice interaction functions are provided to facilitate users in controlling devices and obtaining information.
  • EducationIt assists in teaching by providing students with multilingual and multi-voice audio explanations to help users learn languages and knowledge better.
  • Entertainment industryUsed in animation, games, film and television production to dub characters and create more impactful sound effects.