Vidu Q1 - A highly controllable large-scale video model launched by Shengshu Technology
Vidu Q1 is a highly controllable video model developed by the team of Professor Zhu Jun, Vice Dean of the Institute for Artificial Intelligence at Tsinghua University and founder and chief scientist of Shengshu Technology. It supports the generation of 1080p high-definition video with delicate image quality and rich details...
What is Vidu Q1?
Vidu Q1 is a highly controllable video model developed by the team of Professor Zhu Jun, Vice Dean of the Institute for Artificial Intelligence at Tsinghua University and founder and chief scientist of Shengshu Technology. It supports the generation of 1080p high-definition videos with fine image quality and rich details, meeting the requirements for generating 5-second videos. With upgraded first and last frame functions, only two images are needed to generate cinematic-level natural camera movements. Vidu Q1 features precise sound effect control, supporting the annotation of sound effect type and duration on the timeline, with synchronization accuracy reaching ±0.1 seconds. The model has optimized the controllability of multi-subject details, allowing users to precisely adjust the position, size, and motion trajectory of subjects in the video by uploading reference images and text commands. It can perform local super-resolution reconstruction for blurred areas, maintaining a pixelated appearance even when 4K videos are magnified 8 times. It topped the authoritative overseas video generation benchmarks VBench-1.0 and VBench-2.0 with scores of 87.41% and 60.98% respectively, surpassing models such as Runway and OpenAI Sora. In the domestic SuperCLUE image-based video rankings, Vidu Q1 also topped both lists with a score of 63.52 for anime style and 67.78 for realistic style.
Main functions of Vidu Q1
- High-definition image quality and resolutionIt supports generating 1080p high-definition videos with delicate picture quality and realistic details.
- First and last frame functionUsers only need to upload two images to generate cinematic camera movement effects, with smooth and natural transitions between the first and last frames, and a more "cinematic" visual style.
- Sound effect generationThe newly added "Generate sound effects with one sentence" function can generate background music and sound effects based on prompts. It supports fine control over the timing of each audio segment, and can be controlled in segments and freely superimposed, with sound and picture perfectly matched.
- Extreme "quality" styleThe anime style is more stable and fluid, and the characters' movements and emotional expressions are more effective.
- Video quality and semantic consistencyIn terms of video quality and semantic consistency in VBench-1.0, Vidu Q1 reaches the state of art (SOTA) level, and the generated videos perform well in terms of both surface realism and intrinsic authenticity.
- Common sense reasoning and physical understandingIn both common sense reasoning and physical law understanding dimensions of VBench-2.0, Vidu Q1 also performed well, demonstrating leading understanding and generation capabilities.
- Precisely adjust the main attributesUsers can upload reference images and text commands to select any character or object in a video and precisely adjust its position (coordinate axis positioning), size (percentage scaling), motion trajectory (custom path curve), and action details (such as "raise hand 15 degrees" or "blink frequency 2 seconds/time"). Real-world testing shows that when the same command generates 10 videos, the character offset error is less than 5 pixels, while traditional models typically exceed 200 pixels.
- Multi-subject consistencyIn multi-subject scenarios, the Vidu Q1 maintains consistency between subjects, ensuring that the actions and positions of multiple characters or objects in the video are coordinated and unified. This is crucial for producing complex multi-subject video content (such as animation, short films, etc.).
- Sound effect timeline controlUsers can mark the sound effect type and duration on the timeline, such as setting wind sound (70% intensity) from 0:00 to 0:03 seconds, and setting glass breaking sound (left channel priority) from 0:04 to 0:05 seconds. The Vidu Q1's sound effect synchronization accuracy can reach ±0.1 seconds, which greatly enhances the immersion and appeal of videos compared to traditional AI sound effect random matching.
- Local super-resolution reconstructionLocal super-resolution reconstruction is performed on blurred areas, and 4K videos remain pixelated even when magnified 8 times. Users can manually adjust light and shadow intensity, material texture, depth of field blur, etc., to further enhance the visual quality of the video.
Technical Principles of Vidu Q1
- Technical ArchitectureThe Vidu Q1 is developed based on the Diffusion Model and the U-ViT architecture. U-ViT combines the scalability of the Transformer with its ability to model long sequences, enabling it to handle 1080p videos up to 16 seconds long. The model reduces the spatial and temporal dimensions of the video through a video autoencoder, achieving efficient training and inference.
- Multimodal fusionVidu Q1 integrates information from multiple modalities, including text, images, and video, enabling consistent generation from multiple angles, subjects, and elements through flexible, diverse inputs. This allows Vidu Q1 to generate videos with high consistency and dynamism.
- Automatic generation and annotationTo address the challenge of labeling large-scale video training data, Vidu Q1 utilizes a high-performance video title generator to automatically label training videos. During inference, a re-title technique is applied to rephrase user input into a form more suitable for the model.
- Extensions of controllable video generationThe Vidu Q1 was used in other controlled video generation experiments, including edge detection-based video generation, video prediction, and subject-driven generation. These experiments demonstrated the Vidu Q1's potential in various application scenarios.
Vidu Q1 project address
- API address:platform.vidu.cn
Vidu Q1 Review Results
- Vidu Q1 topped the VBench-1.0 and VBench-2.0 lists of the authoritative overseas video generation evaluation list VBench Leaderboard, surpassing domestic and foreign video generation models such as Runway, Sora, and LumaAI with total scores of 87.41% and 60.98% respectively, and won the first place in the textual video generation track.
- It achieves SOTA (State of the Art) level in comprehensive dimensions such as video quality and semantic consistency in VBench-1.0, and common sense reasoning and physics understanding in VBench-2.0, demonstrating outstanding performance.
- In the VBench 2.0 evaluation, Vidu Q1 won first place in both common sense reasoning and physical law understanding, demonstrating leading understanding and generation capabilities.
- In the image-based video rankings released by SuperCLUE, a leading domestic benchmark for comprehensive large-scale model testing, the Vidu Q1 achieved first place in both the animation style (63.52) and realistic style (67.78) rankings, demonstrating its strong and stable image-based video capabilities in specific applications.
How to use Vidu Q1
- Registration and LoginVisit Vidu's official website and click to register or log in.
- Model selectionSelect the Vidu Q1 model in the top left corner.
- Wensheng VideoEnter text to describe the content you want to generate, and make personalized settings. You can choose to try the 1080p resolution.
- Image and videoUpload the image and a reference image for the last frame, and enter a description of the content you want to generate. You can then personalize the settings, including selecting 1080p resolution.
- Reference videoThe Vidu Q1 model is not currently supported; you can switch to the 2.0 model instead.
- Creating videosAfter setting up, click "Create" to get the generated video and make adjustments.
Application scenarios of Vidu Q1
- Film and television productionThe Vidu Q1 can quickly generate high-quality video content, significantly shortening production cycles and reducing costs. Its multi-camera generation capabilities and control over spatiotemporal consistency provide convenience for special effects production and scene editing.
- AdvertisingThe Vidu Q1 can quickly generate video ads in various styles and themes to meet the needs of different clients. It can achieve precise targeting and personalized recommendations based on user interests and behavioral data, improving ad conversion rates and effectiveness.
- Animation ProductionVidu Q1's multi-subject consistency control capability is of great value in animation production, ensuring the consistency of details of characters from different perspectives and reducing the workload of animators.