AB
AiBoss
project

Wan-Streamer v0.2 - A full-modal understanding and generative model launched by Alibaba Tongyi.

Wan-Streamer v0.2 is an end-to-end, full-modal understanding and generation model for real-time duplex interaction, developed by Alibaba's Tongyi Lab. The model unifies 'listening, seeing, speaking, and acting' into a single Transformer, natively supporting text, ...

What is Wan-Streamer v0.2?

Wan-Streamer v0.2 is an end-to-end full-modal understanding and generation model for real-time duplex interaction, launched by Alibaba's Tongyi Lab. The model unifies "listening, seeing, speaking, and acting" into a single Transformer, natively supporting real-time understanding and synchronous generation of text, audio, and video. The end-to-end response latency is only 550ms, and the output resolution is 640×368@25FPS, enabling AI to listen, see, and respond like a real person, achieving truly natural face-to-face video communication.

Main features of Wan-Streamer v0.2

  • Real-time audio and video dialogueIt supports face-to-face interaction via video calls, with AI sensing the user's audio and video in real time and generating responses simultaneously.
  • Full-modal understandingIt natively supports real-time understanding of text, audio, and video input without the need for external modules.
  • Audio and video synchronization generationIt can output voice and high-definition video simultaneously, achieving perfect audio-visual synchronization.
  • Micro-expression and gesture generationIt can generate natural eye expressions, body postures, gestures, and scene details.
  • Free role-playingIt supports the real-time generation of any character to engage in dialogue using natural language descriptions.

Technical Principles of Wan-Streamer v0.2

  • Native streaming architectureIt maps user input and agent output to the same causal timeline without waiting for the entire passage to end.
  • Flow cell closed loopIt completes a full closed loop of perception, understanding, generation, and decoding every 160ms, enabling simultaneous speaking, listening, and response.
  • Thinker-Performer Dual PathThinker single card is responsible for low-latency perception and language reasoning, while Performer multi-card Ulysses parallel is responsible for high-definition video generation.
  • Overlapping schedulingThe single-card Thinker and multi-card Performer computing windows overlap, separating video generation costs from the latency-sensitive path.
  • End-to-end unified modelingIt unifies text, audio, and video input and output into a single Transformer, avoiding pipeline fragmentation.

How to use Wan-Streamer v0.2

Wan-Streamer v0.2 is currently being released as a research project. For specific access points and API integration methods, please refer to subsequent official announcements from Tongyi Labs.

The core advantages of Wan-Streamer v0.2

  • Extremely low latencyThe end-to-end processing time is 550ms, and the model-side processing time is only 200ms, which is significantly faster than mainstream real-time voice dialogue models.
  • Full-modal end-to-endIt natively supports synchronized audio and video understanding and generation, eliminating the need for external modules such as ASR, LLM, and TTS.
  • High-quality output640×368@25FPS, supports micro-expressions, gestures and scene details, no longer limited to floating heads.
  • Full-duplex interactionIt allows you to speak, listen, and respond simultaneously, eliminating the need to wait for the user to finish speaking before processing, making communication more natural.
  • Scalable architectureThe Thinker-Performer dual-path design improves image quality while maintaining extremely low latency.

The project address for Wan-Streamer v0.2

  • Project official website:https://wan-streamer.com/
  • arXiv technical paper:https://arxiv.org/pdf/2607.04443

Comparison of Wan-Streamer v0.2 with similar competing products

Comparison Dimensions Wan-Streamer v0.2 GPT-4o Realtime
Research and development Ali Tongyi Lab OpenAI
End-to-end delay Approximately 0.55s (including 350ms network latency) Approximately 0.23–0.8s
Model-side delay Approximately 200ms Approximately 230ms
Video output 640×368 @ 25FPS Not supported
Video perception support support
Audio output support support
Text output support support
End-to-end architecture Native end-to-end Transformer Cascaded pipeline (non-end-to-end)
Full-duplex interaction Fully supported Partial support
External module dependencies None (native unified modeling) Requires ASR + LLM + TTS assembly
Architecture Design Thinker-Performer dual-path + Ulysses parallel Single path

Application scenarios of Wan-Streamer v0.2

  • AI assistant for video callsIt allows for face-to-face communication by turning on the camera, making it suitable for scenarios that require a sense of "presence," such as oral practice, mock interviews, and psychological counseling.
  • Contextualized educationAI teachers can observe students' expressions on screen to judge their level of understanding and adjust the pace of teaching in real time.
  • Immersive game NPCsThe game characters have facial expressions, body language, and real-time reactions, allowing them to have a truly "face-to-face" conversation with the player.
  • AccessibilityIt generates real-time video responses with precise lip reading and gestures for hearing-impaired users and describes the surrounding environment for visually impaired users.
  • Virtual role-playingIt allows users to have real-time conversations with historical figures, artistic images, or custom characters, creating an immersive cultural experience.