AB
AiBoss
Tutorials

How to use Tencent Hunyuan video generation model: firsthand testing

Tutorials

How to use Tencent Hunyuan video generation model: firsthand testing

Tencent has finally launched its own AI video platform – the "Hunyuan Video Model." I was recently invited to participate in the internal testing of the Hunyuan Video Model. I worked on it for two days straight over the weekend, from morning till night, and completed over 300 videos. My conclusion so far is:...

如何使用腾讯混元视频生成模型,一手实测

01 Tencent is ready

Tencent has finally launched its own...AIVideo available —Mixed-element video model

Recently, I was invited to participate in the internal testing of the Hunyuan video model. I worked on it for two days straight over the weekend, from morning till night, and produced a total of more than 300 videos.

To state the conclusion first:As the first version delivered by Tencent (Vincent Video, 5s), the overall quality is very high.Command compliance, dynamics and image stability, camera language, realistic texture, physical complianceThey performed well in various aspects.Gacha pulls are rare..

Even in someCamera transitions, motion effects, sci-fi/fantasy style, abstract understandingThere were also some pleasant surprises in other areas.

Please watch the video:

Experience path:Tencent Yuanbao APP -AIapplication-AIvideo.

02 Actual testing of 10 styles and 30 cases

To systematically test the quality of the mixed-source video model, although it's not as systematic as professional benchmarks, I've divided it into 10 sections based on styles and scenes that I personally consider important and frequently used.

These 10 styles are:Close-up, Realism, People, Animals, Science Fiction, Special Effects, Animation, Art/Abstract, Movement, Multi-person Scenes/Large-scale Scenes/Multi-camera Shots.

Design 3-5 pieces for each style.Prompt wordsLet Hunyuan produce videos for evaluation.

Prompt wordsFor part of it, I'll first come up with an idea, describe it in one sentence, and then let...AIPlease help me optimize and expand the text.AIOptimizedPrompt wordsAfter some further modifications, it should be ready to be sent to the model and run.

Prompt wordsThe framework generally relies on these templates.

  • Template 1:Prompt words=Main subject + Scene + Movement
  • Template 2:Prompt words=Subject (Subject Description) + Scene (Scene Description) + Movement (Motion Description) + (Camera Language) + (Atmosphere Description) + (Style Expression)
  • Template 3:Prompt words=Subject + Scene + Movement + (Style Expression) + (Atmosphere Description) + (Camera Movement) + (Lighting) + (Shot Size)

Key Focus Subject + Scene + Movement That's it. If you're not sure how to describe the other parts, you can also select them using the tags provided in the backend.

Without further ado, let's take a look at the running case.

P.S. All cases were tested by myself and do not include any official demos.

Realism is almost a must-test style for video models. The main focus is on how well the model generates different scenes, character expressions, movements, texture details, and lighting changes to see if they are consistent with the real world.

1) A woodpecker is pecking a hole in a tree, in a realistic style.

2) A beautiful Chinese woman wearing Hanfu (traditional Han clothing) has her hair flowing in the wind. The camera then cuts to a close-up of her face. The background is Zhangjiajie.

3) A penguin wearing a red scarf strolls through a sea of flowers, the red scarf contrasting sharply with the colors of the flowers. In the background, the sea of flowers sways gently in the wind, petals fall, and morning dew sparkles.

4) Ultra-long telephoto panning, industrial abandoned factory building, main light seeps in from the broken skylight, natural light.

Close-ups are a style that video models excel at. The key to competition among models lies in their ability to present details, such as details of object movement, details of human body movements, details of facial expressions, and details of image quality.

A good close-up shot can easily bring the audience closer to the protagonist, making them feel as if they are there.

5) A man stares in terror into the distance, with a burning and exploding city in the background. The camera focuses on the man's face, capturing his terrified expression.

6) The camera slowly zooms in. The background is a small, cozy living room, where a young woman sits on the sofa, engrossed in reading. A steaming teacup sits on the coffee table.

7) A strange and terrifying ancient creature crawls in the mud.

For human figures, the main focus is on how realistically the video model portrays skin tone, body movements, facial expressions, and clothing—these are also the easiest aspects for us humans to identify.AIThe place that's real and the place that's fake.

However, text-based videos are not particularly effective at portraying people. To achieve more stable, realistic, and consistent character portrayals, it's generally necessary to generate them using image-based videos.

8) A little boy is intently assembling building blocks.

9) A little girl is holding a balloon and slowly running forward.

10) A man was sitting on the sofa watching TV, then he put his hands on his head and looked very surprised.

Compared to human figures, the video models from various companies perform much better with animals. However, this is on the condition that your animal moves "boldly," rather than simply zooming in and out of the image.

Based on my experience running multiple cases, the Hunyuan video model performs very well in animal realism, almost like a documentary.

11) On the African savanna, a cheetah is running at high speed, chasing an antelope.

12) In the Greater Khingan Mountains, a tiger is running at high speed, with a snow-covered forest in the background.

13) A magpie is foraging for food on a tree branch in front of the red wall of the Forbidden City.

(5) Science fiction, fantasy, and xuanhuan

Science fiction, fantasy, and other fantasy genres are what attract many people to use them.AIOne of the main reasons for making videos is, of course, myself.

Fantasy style, in particular, tests the video model's dataset and generalization ability (referring to the model's ability to perform on new and unseen data), whether it can display some fantasy scenes, such as changes in light and shadow, color changes, deformation effects, motion effects, etc.

This section contains the most cases. Considering the compression during video-to-image conversion, I directly included the original video in some cases.

14) A spaceship is passing through the asteroid belt.

15) A spaceship is passing through a time tunnel, surrounded by colorful lights.

16) Two giant robots are engaged in a fierce battle in the city. Each collision generates a huge shockwave that shatters nearby buildings into pieces.

17) In a dimly lit corridor, a Marine Corps soldier is walking through an abandoned corridor.

18) In the hazy clouds, dark clouds gathered, and lightning flashed and thunder roared. Suddenly, a giant dragon burst through the clouds and came rushing towards them.

This imagination suggests that Hunyuan must have "watched" Game of Thrones many times.

Special effectsSpecial EffectsSpecial effects are the most important visual art in movies and television, with common special effects such as explosions, smoke, flames, and high speed.

Special effects shots also primarily test the generalization ability of video models, examining the model's adherence to instructions and its ability to express details.

19) In a blizzard, a steam train travels through rugged mountains, black smoke billowing straight into the sky from the engine, and the carriages leaving deep tracks in the white snow.

20) An explosion suddenly occurred inside a dilapidated warehouse.

21) On a foggy night, under the bright moonlight, a medieval sailing ship sails on the sea, filled with an eerie atmosphere.

22) Colorful jellyfish swim freely on the seabed. Their bodies are transparent in blue, purple and pink, emitting a mesmerizing glow in the water.

For animation, the main considerations are the model's support for various styles and aesthetics, such as 2D, 3D, vector, clay, ink painting, Miyazaki Hayao, Disney, etc.

One firstSoraofPrompt words.

23) Animated scene features a close-up of a short fluffy monster kneeling beside a melting red candle. The art style is 3D and realistic, with a focus on lighting and texture. The mood of the painting is one of wonder and curiosity, as the monster gazes at the flame with wide eyes and open mouth. Its pose and expression convey a sense of innocence and playfulness, as if it is exploring the world around it for the first time. The use of warm colors and dramatic lighting further enhances the cozy atmosphere of the image.

Let's take a look at Miyazaki Hayao's style.

24) A fantastical garden comes into view. The garden is filled with all sorts of exotic flowers and plants, each unique in shape and vibrant in color. A group of lively and adorable little elves, dressed in brightly colored clothes, frolic and play among the flowers and plants. The Ghibli animation style makes you feel as if you've stepped into a dream world created by Hayao Miyazaki.

(8) Art/Abstract

The art style primarily tests the video model's abstract understanding of graphics, space, color, and force changes. After testing a few cases, I was surprised that Hunyuan could also create some abstract art videos.

25) The particles rotate and converge into an abstract form.

26) Different colors form irregular shapes, which are slowly rotated.

27) A fixed camera at a 5-degree angle, with shallow depth of field and focus, interweaves purple-red neon lights with cyan holographic projections. In the center of the screen, a mechanical dancer dressed in avant-garde attire opens his arms to thank the audience.

Motion is considered the crown jewel of video modeling because it is the most challenging.

To generate videos that conform to real-world physics, the model needs a high level of technical expertise in understanding spatial relationships, handling force changes and shapes of different objects, and semantic understanding of different objects and movements. Only then can videos that follow physical rules be generated.

28) At sunset on the off-road track, a modified Ford F-150 Raptor roared past. The raised suspension allowed the massive run-flat tires to tumble wildly across the mud, splattering mud onto the roll cage and creating mottled patterns. The body decals shimmered in the golden sunlight, and the roar of the supercharger mingled with the exhaust's thunder.

29) In slow motion, a thunderstorm accompanied by lightning shows a dashing Chinese swordsman practicing his swordplay in the rain. The background is a bamboo forest.

30) An off-road vehicle was driving on the steep mountainside, and the distant Gongga Snow Mountain was slowly rising and gradually becoming clearer in the visual view.

(10) Multi-person scenes/large-scale scenes/multi-camera shots

Multi-person scenes involve coordinating the movements of multiple characters and computational power issues. Currently, many video models, including Gen3 and Keling, tend to crash. Let's see how Hunyuan performs.

31) The camera slowly rises from a close-up of the rider's gait, eventually focusing on his face, where he gazes ahead with a resolute expression. The background is a medieval battlefield, where two armies are clashing, and men and horses are falling.

32) A group of people sat around the campfire, chatting and laughing.

Now that we've tested all 10 style categories, let's summarize:

1) The mixed-element model for instructions (i.e.)Prompt words(It is relatively consistent.)In the subsequent designPrompt wordsWhen creating an image, it is recommended to have strong visual logic and clear instructions, and avoid piling up a bunch of modifiers and too many subject words.

Otherwise, it will interfere with the model's attention, which is the T in the DiT architecture of the model, Transformer, the self-attention mechanism.

2) Excellent dynamic performance and image stability.Of the 300+ videos I tested, there were definitely some failed cases, but not a single one involved zooming in or out of a PowerPoint presentation. They were all normal actions at normal speed, with very few slow-motion shots or PowerPoint animations.

3) A thorough understanding of camera language.If you specify the shot and framing, the model will strictly adhere to it. If not specified, the model will...Prompt wordsSometimes, creatively designing shots can be surprisingly effective.

For example, this one is really nice.

Prompt words:A massive wave crashes as surfers leap and perform aerial somersaults. The camera emerges from within the wave, capturing the moment sunlight filters through the water. Water splashes create perfect arcs in the air, and surfboards leave trails as they cut across the surface. The final shot freezes on the perfect moment of the surfer navigating through the curtain of water.

4) You can also switch between shots in 5-second videos.In somePrompt wordsIn scenarios (usually long)Prompt wordsEven with only 5 seconds of video, the mixed-element model can still achieve this.automaticCut the shot. After cutting the shot, it is still possible to maintain the consistency of the subject.

5) It excels in science fiction, fantasy, realistic documentary, special effects, and sports genres, and has a high success rate in producing films.The fantasy style, in particular, has a Game of Thrones feel to it, which is likely related to Tencent's own video dataset.

6) Few card draws.If the instructions are clear, sometimes a satisfactory video can be generated in one go. At worst, generating it 3-5 times will usually yield a satisfactory video.

7) Try to take care of Xiaobai as much as possible.The input box interface offers style, shot size, lighting, camera movement, and multiple modes (smooth camera movement, rich motion, director mode), making it easy even for beginners.fastGet started.

Don't underestimate these tags. During my testing, these tags greatly helped the quality of my videos, especially in terms of video style and camera movement.

Of course, some shortcomings were also found during the testing.

1) Generalization ability needs to be improved.Some unfamiliar, obscure, and untrained descriptive terms (such as subject, scene, action, etc.) cannot be recognized by the mixed-model approach, which affects the model's creativity to some extent.

2) Image quality needs improvementCurrently, it only offers 720P (it really is 720P). Although it provides a "high quality" mode, it is not enough for professional creators.

3) The understanding of local figures needs to be improved.ifPrompt wordsThe model doesn't specify "Asian," so it's usually generated based on Europeans. Of course, text-based videos aren't great at maintaining consistency in the subject matter; to improve consistency, we'll have to wait for image-based videos. Additionally, the model is slightly weaker at conveying emotions.

03 In conclusion

After three consecutive days of testing, in my opinion,As an initial model, Hunyuan's overall quality is very high, outperforming many first versions of video models.

I talked to some students at Hunyuan, and they said it stemmed from their innovation in these areas:

  • Using a new generation of language models as text encoders, it has stronger semantic understanding and image presentation capabilities;
  • The entire process uses a full attention mechanism, rather than a spatiotemporal module, which makes the transition between each frame of video smoother.
  • Using a self-developed image and video hybrid VAE (3D variational encoder), the model's ability to represent details, such as faces, fingers, and high-speed shots, is improved.

More importantly, Tencent announced that it would conduct a review of this model.open source!!

From now on, whether you are an individual or a business...All developers canHugging Faceand on GithubfreeThis model has been used.

Amazing, truly amazing! A model with 13 billion parameters...open sourceAt onceopen sourceThe complete model, including model weights, inference code, and model algorithms, is made public.

It's important to know that video models are the most technically challenging, and...open source,ableopen sourceThere are very few companies, including Movie Gen, the video model launched by Llama (the creator of "Gengoku"). They don't seem to intend to...open source.

Hunyuan video model, launched immediatelyopen sourceThat's impressive—that's the spirit, that's the vision. So far, they have already...open sourceText-to-text, text-to-image, 3D generation andup to dateThe video of Wensheng.

Tools mentioned in this article

Tencent Hunyuan Wensheng Video: