News
Qwen3-TTS full suite open source launched!
The Tongyi Qianwen team has officially open-sourced the Qwen3-TTS series of speech generation models, including 1.7B and 0.6B parameter scales, fully supporting timbre cloning, timbre creation, and anthropomorphic speech generation. Employing an innovative 12Hz multi-codebook speech encoder and a dual-track modeling architecture, it achieves efficient speech compression and high-fidelity reproduction, with a first-packet audio latency as low as 97 milliseconds. The model covers 10 mainstream languages and dialects, including Chinese, English, Japanese, and Korean, and supports precise control of timbre, emotion, and rhythm via natural language commands.