Qwen-Audio-3.0-Realtime - A real-time voice interaction model launched by Alibaba.
Qwen-Audio-3.0-Realtime is a real-time voice interaction dialogue model developed by the Alibaba Tongyi team, offering both Plus and Flash versions. The model maintains high inference depth while achieving millisecond-level response times and supports instructionless agent...
What is Qwen-Audio-3.0-Realtime?
Qwen-Audio-3.0-Realtime is a real-time voice interaction dialogue model developed by Alibaba's Tongyi team, offering both Plus and Flash versions. The model maintains high inference depth while achieving millisecond-level response times, supporting commandless agent tool self-invocation, dynamic emotion and tone adjustment, and duplex dialogue. Qwen-Audio-3.0-Realtime leverages On-Policy Distillation and multi-teacher distillation technology to transfer the capabilities of large text models to voice models. The model is available on Alibaba Cloud's Bailian platform and is suitable for scenarios such as intelligent customer service, education and training, and emotional support.
Main functions of Qwen-Audio-3.0-Realtime
-
Dual-version architectureIt offers two versions, Plus and Flash, to suit different scenarios and needs.
-
Millisecond responseIt can directly generate responses for time-sensitive scenarios such as daily conversations, without waiting for the complete inference chain.
-
Invoke without instructionsWithout requiring the user to explicitly say "open XX", the model automatically determines and calls external tools, and the results are automatically incorporated into multi-round memory.
-
Dynamic Emotional ExpressionAdjust tone, rhythm, pitch and emotion according to context, and support paralinguistic signals such as laughter and sighs.
-
Duplex DialogueIt has a built-in multimodal perception duplex control sub-model, which enables simultaneous speaking and listening, as well as interruption and interjection at any time.
-
Voiceprint LockUpload audio samples via the audio_prompt field to pinpoint a specific speaker and focus on the conversation.
The technical principles of Qwen-Audio-3.0-Realtime
-
On-Policy DistillationThe complete reasoning capabilities of the large text model are distilled into the speech model, and the large text model corrects the speech model in real time when it generates the answer.
-
Multi-teacher distillationFour teachers were introduced: spoken language, general language, agentic language, and audio comprehension, to respectively ensure spoken expression, basic reasoning, tool usage, and audio semantic comprehension.
-
Duplex Control SubmodelIt analyzes audio signals, semantic content, and speaker voiceprint features to determine the rhythm of conversation and achieves noise interference resistance and multi-speaker switching.
-
End-to-end architectureIt integrates speech understanding, reasoning, and generation, avoiding information loss in traditional cascaded solutions.
-
FunctionCall ProtocolIt achieves seamless integration of MCP, API, and knowledge base, as well as context memory fusion, based on standard protocols.
How to use Qwen-Audio-3.0-Realtime
- Access PlatformLog in to the Alibaba Cloud Refinement console and enter the Model Plaza.
- Select ModelChoose either Qwen-Audio-3.0-Realtime-Plus or Flash version according to your needs.
- Obtaining permissionsFollow the instructions to activate the model service and obtain the API key.
- Access ApplicationIntegrate the model into your own application or agent via API or SDK.
- Configuration toolsAccess MCP, API, or knowledge base based on the FunctionCall protocol.
- Customized voiceprintUpload audio samples via the audio_prompt field to pinpoint a specific speaker.
The core advantages of Qwen-Audio-3.0-Realtime
- Achieving both reasoning and speedThe Plus version still maintains a score of 90.5 in VoiceBench's spoken language prompt, while the Flash version only experiences a 5.5-point drop in spoken language proficiency during multi-turn audio conversations, demonstrating millisecond-level response without diminishing intelligence.
- Benchmarking leadershipThe Preview version topped the Artificial Analysis charts for both speech inference and dialogue fluency, achieving scores of 97.6% and 97.8% respectively. The Plus version also tied for first place on the VStyle Chinese chart.
- Seamless tool callIt can automatically call MCP, API and knowledge base without explicit trigger words, and the call results are automatically integrated into multi-round memory, which can be directly reused in subsequent follow-up questions.
- Real-person empathy: Adjust tone, rhythm and pitch dynamically according to context, and achieve natural emotional response through paralinguistic signals such as laughter and sighs.
- Anti-interference duplex dialogueBuilt-in multimodal perception duplex control sub-model ensures uninterrupted communication even in noisy environments and locks onto the main conversation partner during multi-person discussions.
- Voiceprint focusing capabilityIt supports uploading audio samples via audio_prompt to lock onto the voiceprint of a specific speaker, enabling precise focusing in multi-speaker scenarios.
Comparison of Qwen-Audio-3.0-Realtime with similar competing products
| Dimension | Qwen-Audio-3.0-Realtime | GPT-4o Realtime |
|---|---|---|
| Inference performance | VoiceBench Spoken English score: 90.5, with a speaking attenuation of only 2.0. | Strong reasoning ability and good adaptability to conversational scenarios |
| Tool call | No explicit self-invocation instruction required; supports MCP/API. | Support for tool invocation typically requires explicit triggering or text confirmation. |
| Emotional expression | VStyle 4.22 SOTA, Dynamic Mood and Paralinguistic Signals | The voice is highly natural and expresses a wide range of emotions. |
| Duplex Interaction | Built-in multimodal duplex control, noise reduction and locking of the main object | Supports real-time full-duplex dialogue with smooth interruption handling. |
| Open Access | Alibaba Cloud's Hundred-Refined API supports third-party application integration. | The OpenAI API ecosystem supports multi-platform access. |
Application Scenarios of Qwen-Audio-3.0-Realtime
-
Intelligent Customer ServiceIt provides millisecond-level response to user inquiries and automatically invokes order inquiry and after-sales tools to complete a closed-loop service.
-
Education and TrainingReal-time spoken English practice and error correction, dynamically adjust tone to encourage learners, and support multi-round dialogue practice.
-
Emotional companionshipEmpathic responses through paralinguistic signals such as laughter and sighs provide natural emotional support and interpersonal interaction.
-
Office Meeting: Lock onto the main speaker's voiceprint, record key points in real time, and execute synchronously using calendar and email tools.
-
Entertainment and InteractionIt supports immersive scenarios such as debates and role-playing, and allows for interruptions at any time, enhancing the gaming and live streaming experience.