Wan-Streamer v0.2 - A full-modal understanding and generative model launched by Alibaba Tongyi.
Wan-Streamer v0.2 is an end-to-end, full-modal understanding and generation model for real-time duplex interaction, developed by Alibaba's Tongyi Lab. The model unifies 'listening, seeing, speaking, and acting' into a single Transformer, natively supporting text, ...
What is Wan-Streamer v0.2?
Wan-Streamer v0.2 is an end-to-end full-modal understanding and generation model for real-time duplex interaction, launched by Alibaba's Tongyi Lab. The model unifies "listening, seeing, speaking, and acting" into a single Transformer, natively supporting real-time understanding and synchronous generation of text, audio, and video. The end-to-end response latency is only 550ms, and the output resolution is 640×368@25FPS, enabling AI to listen, see, and respond like a real person, achieving truly natural face-to-face video communication.
Main features of Wan-Streamer v0.2
-
Real-time audio and video dialogueIt supports face-to-face interaction via video calls, with AI sensing the user's audio and video in real time and generating responses simultaneously.
-
Full-modal understandingIt natively supports real-time understanding of text, audio, and video input without the need for external modules.
-
Audio and video synchronization generationIt can output voice and high-definition video simultaneously, achieving perfect audio-visual synchronization.
-
Micro-expression and gesture generationIt can generate natural eye expressions, body postures, gestures, and scene details.
-
Free role-playingIt supports the real-time generation of any character to engage in dialogue using natural language descriptions.
Technical Principles of Wan-Streamer v0.2
-
Native streaming architectureIt maps user input and agent output to the same causal timeline without waiting for the entire passage to end.
-
Flow cell closed loopIt completes a full closed loop of perception, understanding, generation, and decoding every 160ms, enabling simultaneous speaking, listening, and response.
-
Thinker-Performer Dual PathThinker single card is responsible for low-latency perception and language reasoning, while Performer multi-card Ulysses parallel is responsible for high-definition video generation.
-
Overlapping schedulingThe single-card Thinker and multi-card Performer computing windows overlap, separating video generation costs from the latency-sensitive path.
-
End-to-end unified modelingIt unifies text, audio, and video input and output into a single Transformer, avoiding pipeline fragmentation.
How to use Wan-Streamer v0.2
Wan-Streamer v0.2 is currently being released as a research project. For specific access points and API integration methods, please refer to subsequent official announcements from Tongyi Labs.
The core advantages of Wan-Streamer v0.2
-
Extremely low latencyThe end-to-end processing time is 550ms, and the model-side processing time is only 200ms, which is significantly faster than mainstream real-time voice dialogue models.
-
Full-modal end-to-endIt natively supports synchronized audio and video understanding and generation, eliminating the need for external modules such as ASR, LLM, and TTS.
-
High-quality output640×368@25FPS, supports micro-expressions, gestures and scene details, no longer limited to floating heads.
-
Full-duplex interactionIt allows you to speak, listen, and respond simultaneously, eliminating the need to wait for the user to finish speaking before processing, making communication more natural.
-
Scalable architectureThe Thinker-Performer dual-path design improves image quality while maintaining extremely low latency.
The project address for Wan-Streamer v0.2
- Project official website:https://wan-streamer.com/
- arXiv technical paper:https://arxiv.org/pdf/2607.04443
Comparison of Wan-Streamer v0.2 with similar competing products
| Comparison Dimensions | Wan-Streamer v0.2 | GPT-4o Realtime |
|---|---|---|
| Research and development | Ali Tongyi Lab | OpenAI |
| End-to-end delay | Approximately 0.55s (including 350ms network latency) | Approximately 0.23–0.8s |
| Model-side delay | Approximately 200ms | Approximately 230ms |
| Video output | 640×368 @ 25FPS | Not supported |
| Video perception | support | support |
| Audio output | support | support |
| Text output | support | support |
| End-to-end architecture | Native end-to-end Transformer | Cascaded pipeline (non-end-to-end) |
| Full-duplex interaction | Fully supported | Partial support |
| External module dependencies | None (native unified modeling) | Requires ASR + LLM + TTS assembly |
| Architecture Design | Thinker-Performer dual-path + Ulysses parallel | Single path |
Application scenarios of Wan-Streamer v0.2
-
AI assistant for video callsIt allows for face-to-face communication by turning on the camera, making it suitable for scenarios that require a sense of "presence," such as oral practice, mock interviews, and psychological counseling.
-
Contextualized educationAI teachers can observe students' expressions on screen to judge their level of understanding and adjust the pace of teaching in real time.
-
Immersive game NPCsThe game characters have facial expressions, body language, and real-time reactions, allowing them to have a truly "face-to-face" conversation with the player.
-
AccessibilityIt generates real-time video responses with precise lip reading and gestures for hearing-impaired users and describes the surrounding environment for visually impaired users.
-
Virtual role-playingIt allows users to have real-time conversations with historical figures, artistic images, or custom characters, creating an immersive cultural experience.