AB
AiBoss
project

Wanxiang First and Last Frame Model - A video generation model for first and last frames open sourced by Alibaba Tongyi

Wan2.1-FLF2V-14B is an open-source 14-parameter first and last frame video generation model. Based on the first and last frame images provided by the user, the model automatically generates smooth, high-definition video transition effects and supports various...

What is the Wanxiang first and last frame model?

Wan2.1-FLF2V-14B is an open-source 14-parameter first and last frame generation video model. Based on the first and last frame images provided by the user, the model automatically generates smooth, high-definition video transitions, supporting various styles and effects. The Wan2.1-FLF2V-14B model is based on an advanced DiT architecture, combined with a high-efficiency video compression VAE model and a cross-attention mechanism, ensuring high spatiotemporal consistency in the generated videos. Users can experience it for free on the Wan2.1-FLF2V-14B official website.

The main functions of the Wanxiang first and last frame model

  • First and last frame-by-frame videoGenerates a smooth, 5-second video at 720p resolution based on the first and last frames provided by the user.
  • Supports multiple stylesIt supports generating videos in realistic, cartoon, comic, and fantasy styles.
  • Detailed replication and realistic actionIt accurately replicates the details of the input image, generating vivid and natural motion transitions.
  • Instructions followedIt controls video content based on prompts, such as camera movement, subject actions, and special effects changes.

Technical Principles of the Wanxiang First and Last Frame Model

  • DiT architectureThe core architecture is based on DiT (Diffusion in Time) and is specifically designed for video generation. It utilizes a Full Attention mechanism to accurately capture long-term spatiotemporal dependencies in the video, ensuring high temporal and spatial consistency in the generated video.
  • Video compression VAE modelThis system introduces a highly efficient video compression VAE (Variational Autoencoder) model, significantly reducing computational costs while maintaining high-quality generated videos. This makes high-definition video generation more economical and efficient, supporting large-scale video generation tasks.
  • Conditional control branchesThe user-provided first and last frames serve as control conditions, enabling smooth and precise first and last frame transformations based on additional conditional control branches. The first and last frames are concatenated with several zero-padding intermediate frames to form a control video sequence. This sequence is further concatenated with noise and a mask, serving as input to the Diffusion Transform (DiT) model.
  • Cross-attention mechanismThe CLIP semantic features of the first and last frames are extracted and injected into the DiT generation process through a cross-attention mechanism. Image stability control ensures that the generated video is highly consistent with the first and last input frames semantically and visually.
  • Training and ReasoningThe training strategy is a distributed approach combining data parallelism (DP) and full sharded data parallelism (FSDP), supporting training on 720p, 5-second video slices. Model performance is improved progressively in three stages:
    • Phase 1Hybrid training to learn the masking mechanism.
    • Phase TwoSpecialized training to optimize the ability to generate first and last frames.
    • Phase ThreeHigh-precision training improves the replication of details and the smoothness of movements.

Project address for the Wanxiang first and last frame model

Application scenarios of the Wanxiang first and last frame model

  • Creative Video ProductionQuickly generate creative videos with scene transitions or special effects changes.
  • Advertising and MarketingCreate engaging video ads to enhance visual appeal.
  • Film and television special effectsGenerate special effects shots such as the changing seasons and day-night cycles.
  • Education and DemonstrationTo create vivid animation effects to aid in teaching or demonstrations.
  • social mediaGenerate personalized videos to attract fans and increase engagement.