AB
AiBoss
project

GPDiT - A video generation model jointly developed by Tsinghua University, Peking University, and Jieyue Xingchen, among others.

GPDiT (Generative Pre-trained Autoregressive Diffusion Transformer) is a novel video generation model developed by Peking University, Tsinghua University, StepFun, and the University of Science and Technology of China. The model...

What is GPDiT?

GPDiT (Generative Pre-trained Autoregressive Diffusion Transformer) is a novel video generation model developed by Peking University, Tsinghua University, StepFun, and the University of Science and Technology of China. Combining the advantages of diffusion and autoregressive models, it predicts future potential frames using an autoregressive approach, naturally modeling motion dynamics and semantic consistency. GPDiT introduces a lightweight causal attention mechanism to reduce computational costs and employs a parameter-free rotation-based temporal conditional strategy to effectively encode temporal information. GPDiT performs exceptionally well in video generation, video representation, and few-shot learning tasks, demonstrating its versatility and adaptability across various video modeling tasks.

Main functions of GPDiT

  • High-quality video generationGenerate long video sequences with high temporal consistency and motion coherence.
  • The video represents learning.Based on autoregressive modeling and diffusion processes, it learns the semantics and dynamic representations of videos for use in downstream tasks.
  • Few-shot learningIt can quickly adapt to various video processing tasks, such as style transfer and edge detection.
  • Multi-task learningIt supports a variety of video processing tasks, such as grayscale conversion, depth estimation, and person detection.

GPDiT Technical Principles

  • Autoregressive diffusion frameworkIt predicts future potential frames based on an autoregressive approach, naturally modeling motion dynamics and semantic consistency.
  • Lightweight causal attentionA lightweight causal attention mechanism is introduced to eliminate attention computation between clean frames during training, reducing computational cost without compromising generation performance.
  • Rotational basis time conditional mechanismIntroducing a parameterless rotating base time conditional strategy that reinterprets the noise injection process as a rotation on a complex plane defined by data and noise components, removing adaLN-Zero and related parameters, and effectively encoding time information.
  • Continuous potential spaceModeling in a continuous latent space enhances the quality of generation and representation capabilities.

GPDiT project address

Application scenarios of GPDiT

  • Video creationGenerate high-quality videos for use in advertising, film and television, animation, etc.
  • Video editingIt enables style switching, color adjustment, and resolution enhancement.
  • Few-shot learningIt can quickly adapt to tasks such as person detection and edge detection.
  • Content ComprehensionAutomatically label, classify, and retrieve video content.
  • Creative generationInspire artists and designers to create art-style videos.