AB
AiBoss
News

Alibaba's Tongyi platform launched the full-modal large-scale model Qwen3.5-Omni.

Alibaba Cloud's Tongyi has launched the Qwen3.5-Omni full-modal large model, achieving state-of-the-art (SOTA) performance in 215 audio and video tasks, surpassing Gemini-3.1-Pro across the board. The model employs a Thinker-Talker collaborative architecture and Hybrid-MoE technology, natively supporting text, image, audio, and video input, and possessing fine-grained audio and video caption generation capabilities. New features include semantic interruption, voice cloning, and voice control for real-time interaction, supporting 256K ultra-long context, 113 language recognition, and 10 hours of audio processing.