The Wan2.5 model, based on practical testing, can generate synchronized audio and video.
The 2025 Yunqi Conference is finally here! This year's theme is "Cloud Intelligence Integration • Silicon-Carbon Symbiosis." More than 2,000 speakers from over 50 countries around the world have gathered in Hangzhou to discuss cutting-edge topics such as Agentic AI and Physical AI...
2025 Yunqi ConferenceIt's finally here! This year's theme is "Cloud-Intelligence Integration • Silicon-Carbon Symbiosis," and more than 2,000 speakers from over 50 countries around the world have gathered in Hangzhou to discuss...Agentic AIWith Physical AIThe dialogue unfolded on cutting-edge topics, creating a scene reminiscent of a science and technology Spring Festival Gala.
This morning, Alibaba was still the focus of attention.up to dateofLarge Model—General Meaning and Myriad Aspects Wan2.5-Preview Series Models.
The Wan2.5-Preview series models areMulti-sensory narrative, using nativeMultimodalArchitectureThe text, image, video, and audio processing capabilities have been comprehensively improved, and videos with synchronized audio and video can be generated directly.
These technological upgrades represent Alibaba's long-term investment in its foundational models, as well as its commitment to industrial applications and driving development.Large ModelThis is a manifestation of ecological expansion.
K-sister was also among the first to get the chance to try it out~ Next, let's take a look at the actual test results.
Wan2.5 provides two main functions: image generation and video generation.It also supports generating videos from audio combined with prompts/images..
We only need to use daily text/image/video content.Prompt wordsBy adding descriptions of human voices, ambient sound effects, and background music to the basics, you can obtain a finished video with synchronized audio and video.
The maximum video generation time is 10 seconds, and it can generate high-definition videos with a resolution of 1080p and 24fps.
Official website: https://tongyi.aliyun.com/wan/
Without further ado, let's look at a few real-world test cases to give you a feel for it:
Case 1: Variety Show Recording
Prompt: At the recording of a variety show, the stage is set up like a living room, with soft, warm lighting. Two sofas face the audience, and a coffee table in the middle holds drinks and snacks. A young male idol sits on a sofa, dressed in stylish casual wear, holding a microphone, and says, "I won't say anything charming, but I am charming in the moment." The audience erupts in laughter, and the camera cuts to other guests, who are laughing and clapping.
In this 5-second clip, Wan2.5...Prompt wordsThe adherence to the design is very high, and the details in the picture are also handled very well, such as the living room style, warm lighting, and drinks and snacks on the coffee table.
The characters' facial expressions and lip movements were very natural, especially during camera movements, where the characters would even lean towards the guests, making it seem like they were about to hand over the microphone at any moment...
Case 2 Outdoor Photography
Upload a photo of a snail
Prompt: On a rainy day, the rain pounded heavily on the grass, making a dull "shush" sound, mixed with the soft splashing of water droplets. The surrounding environment was open and damp.
Dense raindrops pounded on the snail's shell, gathering into large beads that flowed down. Wan2.5 has a pretty good understanding of the real world, judging from the scene in the picture and...Prompt wordsIt generated matching ambient sound effects, and the consistency between sound and visuals was quite good.
case 3 Concert
We uploaded an audio clip of a song.
Prompt: Close-up shot of a stunningly beautiful female singer standing center stage, performing with deep emotion. She wears an elegant gown, her long hair flowing gently in the breeze, making her even more captivating under the stage lights. She grips the microphone tightly, her voice powerful and resonant, brimming with emotion.
The video's lighting and colors are excellent, especially the hair highlights, which are very dynamic and realistic. The lip movements of the characters in the video also match the audio very well.
Wan2.5's audio-visual synchronization is not...SimpleIn addition to making the characters' mouths move, many details were added, such as the slight shaking of the head, the tense muscles in the neck when straining, and the contraction and rise and fall of the shoulders when breathing. These details make the whole picture more lifelike, as if it were really shot on location.
Case 1: Food Video
Prompt: A female college student, around 20 years old, sits in a bustling food street, picking up a small piece of braised pork with her chopsticks, chewing slowly, and leaning close to the camera, softly says, "Delicious." Her voice is sweet and her tone light. The background noise is the hustle and bustle of the food street.
The video quality generated by Wan2.5 and Veo3 is quite good, but Veo3 seems to have encountered a bug, as the entire video has no sound.
Case 2: The Evolution of Television
Prompt: Lock onto a wide-angle lens and shoot the same living room from the front, with the television centered in the frame. The image should consistently show the evolution of television over the decades, from black and white sets in the 1950s, to wooden cabinets in the 1970s, to CRT monitors in the 1990s, to flat-screen TVs in the 2000s, and finally to the 2020s.intelligent OLED TVs. Furniture, colors, and styles also change with the times: retro 70s, minimalist 90s, modern 2000s, and futuristic 2020s.
Lens: 35mm cinema lens, clear details.
Sound effects: Visual static, channel switching sounds, remote control click sounds, all synchronized with the changing times.
Mixed levels: Smooth transitions between eras
Wan2.5Prompt wordsThe degree of adherence is much higher; the television is always in the exact center of the screen, and the central composition is always used, making the theme more intuitive.
In terms of interior design style, there is not much difference between the different eras of Wan2.5, while Veo3 does a better job in this regard.
Both Wan2.5 and Veo3 showcase television styles from multiple eras and feature sound effects during transitions.
Previously, video generation often resulted in mismatched audio and video, requiring the addition of voiceovers, lip-syncing, and background noise across different platforms. Now, with minimal effort...Prompt wordsThis will generate a complete video with synchronized audio and video.
Wan2.5 makes creation directly "visual" and "audio." Creating short videos, virtual anchors, and even remote teaching no longer requires complex post-production.AI It will be able toOne-clickThis greatly lowers the barrier to entry for creation.
Wan2.5 can simultaneously align the rhythm of sound, the semantics of language, and the motion of images. This is not only an evolution in video generation, but also a step towards...MultimodalAIA key step towards mature application.
Advertising, education, film and television, and games all used to rely on manual dubbing and post-production, which was both expensive and time-consuming. Wan2.5 brings video generation to the level of production-grade tools, and low-cost, high-quality virtual content may see a full-scale explosion.