AB
AiBoss
project

VideoReward - A video generation preference dataset and reward model jointly launched by CUHK, Tsinghua University, Kuaishou, and others.

VideoReward is a video generation preference dataset and reward model jointly created by the Chinese University of Hong Kong, Tsinghua University, and Kuaishou Technology. It contains 182,000 labeled data points, covering three aspects: visual quality, motion quality, and text alignment...

What is VideoReward?

VideoReward is a video generation preference dataset and reward model jointly created by the Chinese University of Hong Kong, Tsinghua University, and Kuaishou Technology. It contains 182,000 labeled data points covering three dimensions: visual quality, motion quality, and text alignment, used to optimize video generation models. The reward model, based on human feedback, significantly improves the coherence of video generation and text alignment through multi-dimensional alignment algorithms (such as Flow-DPO and Flow-RWR) and inference-time techniques (such as Flow-NRG). Flow-NRG supports user-defined weights to meet personalized needs.

VideoReward's main functions

  • Building a large-scale preference datasetVideoReward contains 182,000 labeled data points covering three key dimensions: visual quality (VQ), motion quality (MQ), and text alignment (TA), used to capture user preferences for generated videos.
  • Multi-dimensional reward modelBased on reinforcement learning, VideoReward introduces three alignment algorithms, including training-time strategies (such as Flow-DPO and Flow-RWR) and inference-time techniques (such as Flow-NRG), to optimize video generation.
  • Personalized needs supportFlow-NRG allows users to assign custom weights to multiple targets during inference, meeting personalized video quality needs.
  • Improve video generation qualityThrough human feedback, VideoReward significantly improves the coherence of video generation and the alignment with prompt text, outperforming existing reward models.

The technical principles of VideoReward

  • Alignment AlgorithmVideoReward introduces three alignment algorithms that extend the self-diffusion model approach and are specifically designed for flow-based models:
    • Flow-DPO (Direct Preference Optimization)During the training phase, the model is directly optimized to match video pairs that human preferences are met.
    • Flow-RWR (Reward-Weighted Regression)The model is optimized by using a reward-weighted approach to make it more in line with human feedback.
    • Flow-NRG (Noisy Video Reward Guidance)During the inference phase, reward guidance is directly applied to noisy videos, allowing users to assign custom weights to multiple targets to meet personalized needs.
  • Human feedback optimizationThrough human feedback, VideoReward significantly improves the coherence of video generation and its alignment with prompt text. Experimental results show that VideoReward outperforms existing reward models, and Flow-DPO outperforms Flow-RWR and standard supervised fine-tuning methods.

VideoReward's project address

Application scenarios of VideoReward

  • Video generation quality optimizationVideoReward significantly improves the quality of video generation, particularly in visual quality, motion coherence, and text alignment, through a large-scale human preference dataset and a multi-dimensional reward model.
  • Personalized video generationVideoReward's Flow-NRG technology allows users to assign custom weights to multiple targets during inference, meeting personalized video quality needs.
  • Training and fine-tuning of video generation modelsVideoReward provides multi-dimensional reward models and alignment algorithms (such as Flow-DPO and Flow-RWR) that can be used to train and fine-tune video generation models.
  • User Preference Analysis and ResearchVideoReward's large-scale preference dataset covers multiple dimensions, including visual quality, motion quality, and text alignment.
  • Video content creation and editingIn the field of video content creation and editing, VideoReward can help generate higher-quality video footage and improve creative efficiency.