News
Alibaba's Tongyi Qianwen launches a large-scale speech synthesis model, Qwen-Audio-3.0-TTS.
Alibaba officially launched its large-scale speech synthesis model, Qwen-Audio-3.0-TTS, which includes a Flash version for real-time interaction and a Plus version for high-quality generation. The model has achieved systematic improvements in fine-grained label control, natural language instruction style definition, multilingual and dialect coverage, and robustness to complex acoustics. It supports embedding structured tags in text to accurately control tone and emotion, and covers 16 languages and 20 dialects.