AB
AiBoss
project

Step-Video-T2V - Step-Star's open-source text-to-video model

Step-Video-T2V is an open-source text-to-video pre-trained model developed by the Step-Stars team. It boasts 30 billion parameters and can generate high-quality videos up to 204 frames long. The model is based on a deep-compressed variational autoencoder (Video-...

What is Step-Video-T2V?

Step-Video-T2V is an open-source text-to-video pre-trained model developed by the Step-Leap Star team. It boasts 30 billion parameters and can generate high-quality videos up to 204 frames per second. Based on a deep compression variational autoencoder (Video-VAE), the model achieves 16×16 spatial compression and 8× temporal compression, significantly improving training and inference efficiency. Step-Video-T2V is equipped with a bilingual text encoder, supporting both Chinese and English input prompts, and further enhances video quality through the Direct Preference Optimization (DPO) method. The model, based on a diffusion-based Transformer (DiT) architecture and a 3D full attention mechanism, excels in generating videos with strong motion dynamics and high aesthetic quality.

Main functions of Step-Video-T2V

  • High-quality video generationStep-Video-T2V boasts 30 billion parameters, capable of generating high-quality video up to 204 frames per second, and supports a resolution of 544×992.
  • Bilingual text supportEquipped with a bilingual text encoder, it supports direct input of Chinese and English prompts and can understand and generate videos that match the text descriptions.
  • Dynamic and aesthetic optimization: Generate videos with strong dynamic effects and high aesthetic quality through a 3D full-attention DiT architecture and Flow Matching training method.

Step-Video-T2V Technical Principles

  • Deeply compressed variational autoencoder (Video-VAE)Step-Video-T2V uses a deep compression variational autoencoder (Video-VAE) to achieve 16×16 spatial compression and 8× temporal compression. This significantly reduces the computational complexity of video generation tasks while maintaining excellent video reconstruction quality.
  • Bilingual text encoderThe model is equipped with two pre-trained bilingual text encoders that can handle both Chinese and English prompts. Step-Video-T2V can directly understand Chinese and English input and generate videos that match the text descriptions.
  • Diffusion-based Transformer (DiT) architectureStep-Video-T2V is based on a diffusion-based Transformer (DiT) architecture and incorporates a 3D full attention mechanism. Through Flow Matching training, it progressively denoises the input noise into latent frames, using text embeddings and time steps as conditional factors. It excels in generating videos with strong motion dynamics and high aesthetic quality.
  • Direct Preference Optimization (DPO)To further improve the quality of generated videos, Step-Video-T2V introduces the Video Direct Preference Optimization (Video-DPO) method. DPO fine-tunes the model using human preference data, reducing artifacts and enhancing visual effects, resulting in smoother and more realistic generated videos.
  • Cascaded training strategyThe model employs a cascaded training process, including text-to-image (T2I) pre-training, text-to-video/image (T2VI) pre-training, text-to-video (T2V) fine-tuning, and direct preference optimization (DPO) training. This accelerates model convergence and makes full use of video data of varying qualities.
  • System optimizationStep-Video-T2V has undergone system-level optimizations, including tensor parallelism, sequence parallelism, and Zero1 optimization, to achieve efficient distributed training. It introduces the high-performance communication framework StepRPC and the two-layer monitoring system StepTelemetry to optimize data transmission efficiency and identify performance bottlenecks.

Step-Video-T2V project address

Application scenarios of Step-Video-T2V

  • Video content creationStep-Video-T2V can quickly generate creative videos based on text prompts, helping creators save time and effort and lowering the barrier to video production.
  • Advertising productionIt can generate personalized video ad content for brands and advertisers, enhancing the appeal and reach of their ads.
  • Education and TrainingStep-Video-T2V can generate instructional videos to help students better understand and memorize knowledge.
  • Entertainment and FilmIt provides creative materials for film and television production, assists in generating special effects, animations, or short drama clips, and accelerates the creative process.
  • social mediaStep-Video-T2V provides users with personalized video generation tools, enriching the content ecosystem of social media platforms and enhancing user interaction. The generated videos can be used for creative content sharing on social media.