How to create synchronized audio and video videos using the Qianwen APP, with tutorials and examples.
I recently discovered that the Qianwen App has been updated again. This time, it features a comprehensive upgrade to its creative capabilities, incorporating the latest image model, Qwen-Image-Edit, for both raw and edited images, as well as Wan2.5's audio and video capabilities. A single image can generate a narrated...
Friends, I recently discovered that the Qianwen App has been updated again.
This represents a comprehensive upgrade in creative ability, incorporating...up to dateImage Model Qwen-Image-Edit's capabilities for both raw and edited images.,Wan2.5audio and video capabilitiesA single image can generate a high-definition video that can talk, sing, and has precise lip movements.
Even ordinary people who don't know editing can directly create complete short video content.
Case 1: Cute Pet Podcast
We open the Qianwen App, upload an image in the chat box, and enter...Prompt wordsYou can directly edit images. For example, you can turn the host in a picture into a cute pet.
Prompt wordsReplace the two hosts in the picture with an anthropomorphic orange cat and a white Samoyed, keeping the background the same.
Click the bottom "AI"Live video" - we can use this image to directly generate a video with synchronized audio and video.
Prompt wordsIn a podcast featuring an orange cat and a Samoyed, the orange cat excitedly complains, "He says we shed too much fur." The Samoyed, thinking for a moment, replies, "We shed fur too, he's almost bald himself!" They then look at each other and burst into laughter.
Wan2.5 generates different timbres based on the images of orange cats and Samoyeds, and the lip movements and expressions are completely synchronized. The main lines are marked with quotation marks, which makes the generation more accurate.
Let's try again, and the finished product is also straight out of the box:
Prompt wordsIn a podcast featuring an orange cat and a Samoyed, the Samoyed asked, "Did you hear your owner calling you just now?" The orange cat replied, "I pretended not to hear so she would open my food can." After saying that, the cat and dog burst into laughter.
Previously, making these kinds of short videos required first creating the visuals, then adding voiceover, and finally lip-syncing. Videos with multiple subjects in dialogue were particularly troublesome to make. Now, it only takes...Prompt wordsTo be clear, a complete video can be produced in less than 5 minutes, resulting in a significant improvement in efficiency.
Case 2 Film and Television Entrepreneurship
Prompt wordsThe character in Figure 1 changes to the pose in Figure 2.
The consistency of the figure is maintained very well. Qwen-Image-Edit not only changed the pose, but also incorporated the features of the figure in Figure 2, such as accessories and tattoos, and the background blends in very naturally without any sense of disjointedness.
We continue generating video:
Prompt wordsThe man in the picture is performing freestyle on the center of the stage, singing, "The grudges and feuds in the harem are nothing more than a pastime for me after dinner," while dancing to the rhythm.
Wan2.5 has a good understanding of Chinese lyrics and stage performance; his freestyle rap and rhythmic drum beats are both excellent. AI automaticThe generated characters' lip movements, actions, and rapping tone all match, resulting in a very coherent performance.
The microphone stand has a minor flaw, but it doesn't affect the overall look.
Case 3 Instructional Video
Prompt wordsThe main subject in the picture appears to be an English teacher explaining English words on the blackboard in a classroom. She says, "The word on the blackboard is Rabbit. It means: rabbit. Say it with me, Rabbit."
The main body's narration and speaking rhythm are completely synchronized, and the pronunciation is also very standard, so it can be used directly as teaching material.
It can also be adapted into a children's song and sung:
Prompt wordsThe main subject in the picture is like an English teacher teaching everyone to sing the Little Rabbit Nursery Rhyme in the classroom. The song goes: "Rabbit, Rabbit, little white rabbit, long ears, red eyes, white fur, soft belly."
When the subject sings, their body and ears sway naturally, and their expression is very natural.
Case 4 Singing and Dancing
Prompt wordsThe kitten twirls and leaps, singing the melody of a nursery rhyme: I am the most magical kitten.
The model is quite accurate in recognizing cartoon characters. The voice is a rather childish one, and the nursery rhyme melody is purely deduced by the model itself, which is a bit "unpleasant to listen to," but it is indeed fun to watch.
Case 5: Parody Videos
Prompt wordsThe image shows a cat performing a mechanical, frame-dropping dance, its feet rapidly tapping the ground with jerky, erratic movements. Combined with a highly rhythmic dance, it also sings in a melodic, spoken-word style with light rap: "I should have been relaxed and at ease, but instead I'm rushing around, stumbling and crawling, lying through my teeth, what are you sobbing about? What are you crying about, you're so pathetic." The overall style is bizarre, absurd, and catchy.
The kitten's movements are very rhythmic, even the rope around its neck swings in sync with the beat, adding a lot to the detail. However, Wan2.5 still cannot recognize the original melody, but it can infer the melody based on rhythm and style, making it very good at generating abstract, absurd, and funny videos.
Case 6: Terracotta Warriors Group Dance
Prompt wordsAll the characters in the picture are singing nursery rhymes while doing school radio calisthenics in unison.
The lyrics go: "One, two, three, four, stretch your arms; two, two, three, four, bend your waist; daily exercise is good for your health; let's do morning exercises together!"
The song is in the style of a nursery rhyme, with a melodySimpleIt is catchy and easy to sing, with a low vocal range, making it suitable for group singing; the rhythm is a medium-slow 4/4 time signature, with clear and stable drumbeats, leaning towards a march rhythm; each line is naturally aligned with the action commands.
The movements include raising hands, stretching, swinging arms left and right, bending over, and chest expansion exercises. The movements are standardized and mechanically consistent, giving the overall impression of orderly school calisthenics. All characters move slowly and synchronously, reminiscent of school calisthenics.
The group movements were very synchronized and corresponded to the rhythm of the nursery rhyme, resulting in a very stable overall performance.
I incorporated a clear style and rhythm.Prompt wordsAfterwards, the sense of rhythm was significantly improved, and the generated music was more in line with the current scene setting, making it quite controllable.
The core upgrade of the Qianwen App this time is:Prompt wordsWith a stronger understanding, group cohesion is maintained better.
The generated song is notSimpleInstead of using a template, AI The understanding of music allows for the generation of melodies, background music, and timbres that align with the main rhythm of the visuals, resulting in a more harmonious and complete video without the need for secondary processing.
Currently, the Qianwen App's integrated audio and video generation capabilities are quite mature.
AI The generated melody is not a tune we are familiar with, but it is precisely because of this deviation that it seems more adorable and fits the current trend of playing with abstract and meme content.
by AI Judging from the speed of iteration, today we are just working on abstraction, but after a while, Qianwen App may really become a music master.
More importantly, this upgrade of the Qianwen App brings the core capabilities of content creation from the web version to the mobile version.
Previously, creating videos with synchronized audio and video often required working on a computer. Now, you can simply pick up your phone, see an interesting scene, or have an idea, and instantly turn it into a highly polished video on the Qianwen App.
This means that creation is becoming more closely tied to everyday scenarios: commuting, spare moments, and casually jotting down inspiration can all serve as entry points for creation.
The core value of creators has shifted further from production ability to the creativity itself.