AB
AiBoss
project

Wan3.0 - Alibaba's latest large-scale video generation model

Wan3.0 is the latest video generation model launched by the Alibaba Cloud Wanxiang team. The model has been comprehensively upgraded in dimensions such as generation duration, versatile creation, all-around reference, and realism. It can generate 30-second videos in a single run and supports doc, xls, and other formats for the first time.

What is Wan3.0?

Wan3.0 is the latest video generation model launched by the Alibaba Cloud Wanxiang team. The model has been comprehensively upgraded in dimensions such as generation duration, universal creation capabilities, all-around reference, and realism. It can generate 30-second videos in a single run and, for the first time, supports input document formats such as doc, xls, ppt, pdf, and md, enabling video generation from virtually anything. The model excels in character portrayal, consistency maintenance, and scene reproduction, and can be widely applied in film and television production, advertising and marketing, design and creative industries, and cultural tourism communication.

Main functions of Wan3.0

  • 30-second long video generationIt supports generating 30-second videos at a time, and also supports intelligent duration recommendation and video extension functions.
  • Multimodal universal inputSupports input of text, images, audio, video, and document formats such as doc/xls/ppt/pdf/md.
  • Real-world recreationThe characters are all unique, with their facial features, skin, micro-expressions and body movements interacting naturally, and the group scenes conveying subtle emotions.
  • Universal Reference ConsistencyIt accurately replicates key dimensions such as facial features, hairstyle, clothing, props, spatial relationships, and style of the characters.
  • Video editing skillsIt supports modifications to visuals, plot, and dialogue, enabling fine-tuning of video content.

Technical Principles of Wan3.0

  • Prompt word engineering centerThe model transforms prompts into a central hub for scheduling inputs, implementing functions, and controlling expression, allowing diverse information to be transformed into coherent video expressions.
  • Multimodal unified understandingIt fully understands text, images, audio, video, and document formats such as doc/xls/ppt/pdf/md to achieve cross-modal information fusion and generation.
  • Iterative optimization architectureFrom Wan1.0 to Wan3.0, it has undergone 8 version iterations, continuously upgrading in dimensions such as generation time, consistency maintenance and realism.

How to use Wan3.0

  • Access Experience PlatformAccess Alibaba Cloud Bailian, Wanjing Yike, Wanxiang official website, Qianwen Creation PC client, IF STUDIO, Duiyou or Qianwen APP.
  • Select input modeUpload text, images, audio, video, or documents such as doc/xls/ppt/pdf/md as a reference for generation.
  • Input prompt words: Describes the desired video content, style, and camera language; the model automatically schedules multimodal inputs for generation.
  • Set generation parametersSelect 480P/720P/1080P resolution, use the smart duration function, or manually set the video length.
  • Acquiring and editing videosOnce generated, it can be downloaded directly, or the video editing function can be used to modify the visuals, plot, and dialogue.

Wan3.0's core advantages

  • Duration BreakthroughIt supports 30-second video generation per session and features intelligent duration recommendation and video extension functions to meet the needs of complete storytelling.
  • All things can be bornIt is the first to support document format input, and office materials can be directly converted into teaching courseware, product demonstrations and other video content.
  • A thousand people, a thousand facesThe portraits are depicted with keen sensitivity and detail, and the micro-expressions and body movements are naturally linked, bidding farewell to the greasy and monotonous feeling of AI characters.
  • Consistency GuaranteeThe key dimensions of the character's facial features, clothing, props, spatial relationships and style are accurately replicated in the all-around reference task.
  • Production-level pricingAPI is billed per second, with 480P/720P/1080P costing 0.3/0.6/1.2 yuan/second respectively, lowering the barrier to commercial use.

Comparison of Wan3.0 with similar competing products

Comparison Dimensions Wan3.0 Kuaishou Keling 3.0
Generation time Each video is 30 seconds long and supports intelligent duration recommendation and video extension. Supports generation of longer videos
Input mode Text/images/audio/video/documents (doc/xls/ppt/pdf/md, etc.) It primarily supports text, images, and videos.
Document to video First-ever support for direct conversion of structured office materials Direct document input generation is not supported.
consistency Characters, props, style, and spatial relationships are accurately replicated. The consistency between the character and the scene is good.
Pricing Model Billed per second, 0.3/0.6/1.2 yuan/second Billing based on points or membership subscription

Application scenarios of Wan3.0

  • Film and television creationThe model supports 30-second durations and realistic scene reproduction, enabling low-cost, high-quality production of AI short dramas, music videos, and vlogs.
  • Advertising and MarketingBreaking through the limitations of text and images, we provide creative videos for industries such as home appliances, automobiles, 3C products, and apparel, ranging from product displays to brand narratives.
  • Design CreativityIt directly transforms UI interaction demonstrations, software function animations, and data visualizations into dynamic works, eliminating the tedious animation production process.
  • Cultural tourism communication: No need for large-scale on-location shooting, low-cost digital presentation of city image films, natural scenery and cultural landscapes.
  • Office and teachingUpload PPT, PDF and other documents to generate teaching materials, business reports and product demonstration videos, improving the efficiency of information transmission.