AB
AiBoss
project

Vidu Q3 - A synchronized audio-visual AI video model launched by Shengshu Technology

Vidu Q3 is the world's first 16-second AI video model with synchronized audio and video, launched by Shengshu Technology. It's specifically designed for narrative scenarios such as short dramas, comics, and commercials. A single prompt is all it takes to produce a 16-second 1080p video, perfectly capturing the visuals, dialogue, and environment...

What is Vidu Q3?

Vidu Q3, launched by Shengshu Technology, is the world's first 16-second AI video model with synchronized audio and video, specifically designed for narrative scenarios such as short dramas, comics, and commercials. It can directly output a 16-second 1080p video with a single prompt, perfectly aligned with the visuals, dialogue, ambient sound effects, and background music, requiring no post-production. The model has a built-in "director's brain," automatically or manually switching between long shots, medium shots, and close-ups to complete complex transitions; it supports direct rendering of Chinese, English, and Japanese text on the screen, with clear and readable road signs and subtitles; and it synchronizes lip movements and voices with characters during multi-person dialogues, supporting the use of three languages. Officially, it ranks first in China and second globally in the Artificial Analysis leaderboard, surpassing Runway Gen-4.5, Google Veo 3.1, and Sora 2. It is now available on the website vidu.cn and via an API platform.

Main functions of Vidu Q3

  • 16-second audio and video outputGenerates a 16-second 1080p video in one go, with the visuals, dialogue, ambient sound, and background music all fully synchronized, requiring zero post-production.
  • Director-level shotsAutomatic or manual switching between long shot/medium shot/close-up, completing multi-camera transitions in a single shot, aligning rhythm with emotion.
  • Multilingual text renderingChinese, English, and Japanese text are directly embedded in the screen, making road signs, subtitles, and product packaging clearly readable.
  • Multi-person dialogue synchronizationThe system allows for synchronized lip movements, timbre, and emotions for multiple characters, supports three languages of dialogue, and features voices that change with the character's appearance.
  • Dual-mode creationBoth text-based and image-based audio and video support any duration from 1 to 16 seconds, with customizable resolution and motion amplitude.
  • Industrial InterfaceThe website vidu.cn and API platform platform.vidu.cn are open simultaneously, with pay-as-you-go pricing and support for mass production.

Technical Principles of Vidu Q3

  • U-ViT backbone architectureBy replacing the traditional U-Net with Transformer, long skip connections are retained, and global attention can "see" the entire 16-second sequence at once. Errors will not accumulate over time, ensuring consistency between the beginning and end of the sequence.
  • Video compression and distributed trainingFirst, spatiotemporal compression is performed on 16-second high-resolution videos to reduce sequence length; then, in conjunction with a self-developed distributed framework, communication efficiency is doubled, memory usage is reduced by 80%, and training speed is increased by a cumulative 40 times, enabling end-to-end long videos to be inferred on a single card.
  • Multimodal unified diffusion: Jointly train the visual, audio and text domains in the same noise space of U-ViT to achieve "one noise - simultaneous denoising": the screen frame, dialogue waveform and ambient sound track are generated synchronously, rather than being spliced in post.
  • 3D Voice-Lips SynchronizationThe audio branch uses 3D VAST-style speech synthesis, which first predicts the character's lip shape coefficient, and then reverses to generate dialogue and sound effects with spatial orientation, ensuring that lip shape, timbre and emotion are aligned when multiple people are talking.
  • Shot scheduling algorithmDrawing inspiration from film shot composition theory, camera position labels such as "long shot-medium shot-close-up" are encoded as conditional vectors and injected into the cross-attention layer of the Transformer. The model dynamically determines the camera position of the next frame during each denoising step, achieving automatic switching within a single shot.
  • Pixel-level text rendering engineAn additional "glyph-pixel" alignment module is trained, which embeds the text vector outline as a priori mask into the diffusion process, so that Chinese/English/Japanese characters grow directly onto the surface of objects in the image and can be clearly read without post-processing textures.

How to use Vidu Q3

  • Register/LoginVisit Vidu's official website, register with your mobile phone verification code, new users receive free points, and you can earn more by checking in daily.
  • Select creation modeOn the left side of the workbench, click "AI Video" to select the mode.
    • Wensheng Audio/Video (Plain Text)
    • Image-based audio/video upload (image + text)
    • Reference video (upload 1-7 main images to identify the character).
  • Write prompt words(Key Steps): Official Structure: Scene + Subject + Action + Shot + Emotion + Sound.
  • Setting parameters
    • Duration: 4 / 8 / 16 s
    • Resolution: 540p | 720p | 1080p
    • Range of motion: Small - Medium - Large - Automatic
    • Audio: Synchronized dialogue, ambient sound, and background music can all be turned on or off individually.
  • Generate and PreviewClick "Create", wait for the output to be generated, and you can preview it online once it's complete. If you're not satisfied, simply change the prompts and run it again. A 4-second clip will be ready in about 30 seconds.
  • Post-production fine-tuningIf the image quality isn't good enough, use the "Smart Ultra HD" function to upgrade it with one click.You can change the seed for comparison, or adjust the motion amplitude and generate again..
  • Export/DownloadClick "Download" on the preview page to get the 16-second 1080p final version (including audio); you can also share it directly to social media.
  • API Batch (Optional)Developers can visit platform.vidu.cn and select REST API. The parameters are the same as those on the web version. Billing is as low as $0.07 per second.

Application scenarios of Vidu Q3

  • Short dramas and filmsIt can generate a complete 16-second clip with one click, and can preview the storyboard and check the rhythm, reducing the cost of early visualization to the level of "writing prompts"; multi-person dialogue and emotional progression can be done in one go, and it can be used directly as a "digital film set".
  • Advertising and e-commerceDuring the proposal stage, the product presentation is directly aligned with the speaker's style, and the speaker's actions, speaking speed, and selling points are synchronized; a single product image can generate multi-scenario demonstrations, improving A/B testing efficiency by 10 times.
  • Self-media account: Cat and dog talk shows, anime radio, and other "brainstorming" series. With just a reference picture and a joke, you can produce a finished product with subtitles, sound effects, and dialogue in minutes. One person is your editorial department.
  • Music VideoStatic cover image + lyrics prompts, directly generate singer playing excerpts, with synchronized lighting, lip movements, and timbre, saving bands the need to rent a studio to shoot sample videos.
  • Education and Science PopularizationThe course features a 5-second concept introduction and a 10-second summary, with automatic synchronization of audio and subtitles. Teachers can focus on writing their lecture notes while the visuals are output in batches by the model.
  • City cultural tourism promotionAerial photography combined with text banners and neon nighttime subtitles can be generated in one go, without the need to close roads or rent helicopters, allowing you to create vertical short videos of "Sydney Opera House" and "Pattaya Beach".