project
Qwen-Audio-3.0-TTS - A speech synthesis model launched by Alitongyi Qianwen
Qwen-Audio-3.0-TTS is a large-scale speech synthesis model launched by Alibaba's Tongyi Qianwen, which includes a Flash version for real-time interaction and a Plus version for high-quality generation.
What is Qwen-Audio-3.0-TTS?
Qwen-Audio-3.0-TTS is a large-scale speech synthesis model launched by Alibaba's Tongyi Qianwen, including a Flash version for real-time interaction and a Plus version for high-quality generation. The model won the championship on the Artificial Analysis global TTS leaderboard, supports fine-grained label control, natural language command-defined voice style, covers 16 languages and 20 Chinese dialects, and improves audio output to 48KHz studio-quality recording, with a maximum single synthesis time of 3 minutes.
Main functions of Qwen-Audio-3.0-TTS
-
Fine-grained label controlEmbedded in text
[gasp]、[giggles]、[angry]Structured tags allow for precise control of tone, emotion, and breathing details. -
Free-style command controlIt allows you to describe characters, emotions, scenes, and speech rates using natural language, and define your voice style without needing professional acoustic knowledge.
-
Multilingualism and DialectsIt covers 16 languages including Chinese, English, Japanese, Korean, and German, supports 20 Chinese dialects such as Cantonese, Chongqing, Northeastern, and Shanghai, and has the highest error rate (SOTA) for 10 languages.
-
Complex acoustic robustnessNative embedded speech enhancement capabilities, accurately cloning timbre even in high-noise, high-reverberation environments.
-
Premium Sound LibraryIt offers command-style timbre, dialect timbre, fine-grained control timbre, and 14+ minority language timbres, all ready to use out of the box.
-
48kHz High-Quality OutputUpgraded from 24kHz to 48kHz, with a maximum single synthesis time of 3 minutes.
Technical Principles of Qwen-Audio-3.0-TTS
- Dual-track hybrid streaming generation architectureIt adopts a Dual-Track LM architecture, which supports both streaming and non-streaming generation in a single model. It can output the first packet of audio immediately upon inputting a single character, with an end-to-end synthesis latency as low as 97ms.
- Discrete multi-codebook language modelBased on discrete multi-codebook LM, we achieve end-to-end speech modeling with full information, completely bypassing the information bottleneck and cascading errors in the traditional LM+DiT scheme, and significantly improving generalization ability and generation efficiency.
- Dual word segmenter design:
-
25Hz word segmenterA single-codebook codec that emphasizes semantic content, seamlessly integrates with Qwen-Audio, and achieves streaming waveform reconstruction through block-level DiT, making it suitable for high-quality non-streaming scenarios.
-
12Hz word segmenterThe 12.5Hz, 16-layer multi-codebook design enables extreme bit rate compression, and the lightweight causal ConvNet completes ultra-low latency streaming reconstruction, making it suitable for real-time interaction.
-
- Large-scale multilingual trainingTrained on over 5 million hours of speech data covering 10 languages, it supports ultra-fast 3-second speech cloning and text-based voice design.
- Command-driven acoustic controlBy treating voice control as a language modeling task, and using natural language commands in ChatML format, we can flexibly manipulate multi-dimensional acoustic attributes such as timbre, emotion, speech rate, and role, so as to achieve what we think is what we hear.
How to use Qwen-Audio-3.0-TTS
- Select version:
-
Plus versionPursuing the ultimate sound quality and expressiveness, it is suitable for film and television dubbing, audiobooks and other scenarios.
-
Flash versionIt aims for low-latency, real-time interaction and is suitable for scenarios such as live streaming and chatbots.
-
- Access PlatformAccess the Alibaba Cloud Refinement console, search for and activate it.
qwen-audio-3.0-tts-plusorqwen-audio-3.0-tts-flashServe. - Call API: Pass in text via the standard API, and you can attach structured tags or natural language instructions to obtain 48kHz synthesized audio.
- Select tone Choose a preset tone from the premium tone library, or upload a reference audio file to clone the tone.
The core advantages of Qwen-Audio-3.0-TTS
- Leading performance on the listIt topped the Artificial Analysis global TTS rankings, achieved unbeatable results in speaker similarity assessments across 16 languages (Plus version), and boasted state-of-the-art word error rates in 10 languages.
- Dual versions cover all scenariosThe Flash version has a 97ms latency for the first packet, meeting the requirements for real-time interaction; the Plus version has a 48KHz output, meeting the requirements for film and television quality.
- Fine-grained expressiveness:support
[angry]、[gasp]Structured tags and free-style natural language commands allow for precise control of emotions, breathing, speech rate, and role. - Multilingualism and DialectsCovering 16 languages and 20 Chinese dialects, it alleviates the problem of "weakened dialect characteristics" through specialized training, restoring the authentic flavor of native speakers.
- Acoustic robustnessIt natively embeds voice enhancement capabilities, accurately cloning timbre even in high-noise, high-reverberation environments, without requiring a quiet recording environment.
- Rapid CloningVoice cloning can be completed with just 3 seconds of reference audio, and text-based voice design is supported.
Qwen-Audio-3.0-TTS project address
- Project official website:https://funaudiollm.github.io/qwen-audio-3.0-tts/
Comparison of Qwen-Audio-3.0-TTS with similar competing products
| Dimension | Qwen-Audio-3.0-TTS-Plus | ElevenLabs v3 |
|---|---|---|
| Rankings | Artificial Analysis Part 1 | Not in top 3 |
| Sampling rate | 48KHz | Typically 44.1kHz |
| Tag control | Native support for structured tags | Indirect control via Prompt is required. |
| Dialect support | 20 Chinese dialects | limited |
| Multilingual SOTA | Optimal word error rate across 10 languages | Leading in some languages |
| Noise robustness | Native speech enhancement | Depends on clean reference audio |
| API pricing | $27.6 / 1M characters | Approximately $11 / 1M characters |
Application Scenarios of Qwen-Audio-3.0-TTS
- Film and game dubbingBy precisely controlling emotional fluctuations and breathing rhythm through structured tags, it achieves character-level performance-level voice synthesis, meeting the high expressiveness requirements of animation, film and television dramas, and game characters.
- Audiobooks and podcastsDefine narration style and pacing using natural language commands; high-quality 48kHz output up to 3 minutes per session; supports batch generation of long audio content.
- Intelligent Customer Service and AssistantThe Flash version boasts an ultra-low first-packet latency of 97ms and supports 20 dialects, enabling natural and smooth real-time voice interaction and significantly enhancing the user service experience.
- Online EducationIt clones the teacher's voice and combines it with speech rate and emotion control to generate personalized teaching voice, supporting the localization of voice production for multilingual course content.
- Live streaming and real-time interactionThe low latency feature of the Flash version is suitable for scenarios that require real-time voice feedback, such as real-time bullet screen reading, virtual anchor broadcasting, and live e-commerce.