AB
AiBoss
project

Loong - A long-form video generation model jointly launched by HKU and ByteDance

Loong is a novel long-form video generation model jointly developed by the University of Hong Kong and ByteDance. It can generate minute-long videos with consistent appearance, rich dynamics, and natural scene transitions. The model is based on an autoregressive large language model (LLM)...

What is Loong?

Loong, a novel long-form video generation model jointly developed by the University of Hong Kong and ByteDance, can generate minute-long videos with consistent appearance, rich dynamics, and natural scene transitions. Based on an autoregressive large language model (LLM), the model integrates text and video information into a unified sequence. It overcomes the challenges of long-form video training using a progressive short-to-long training scheme and a loss reweighting strategy. Loong's design allows the model to learn to generate videos from text prompts during training, extending to generate videos exceeding the training length. Loong's research incorporates inference strategies, including video tag re-encoding and sampling strategies, to reduce error accumulation during inference.

Loong's main functions

  • Long video generationGenerate video content that is one minute or longer.
  • Text to video conversionGenerate video content that matches the given text prompts.
  • Content coherence: Ensure that the generated video has a high degree of consistency in appearance, dynamic changes and scene transitions.
  • Dynamic richnessCapture and represent complex dynamics and motion changes in videos.
  • Natural scene transitionAchieve smooth transitions between different scenes in the video to maintain visual continuity.

Loong's technical principles

  • Unified sequence modeling: Loong models text and video tags as a unified sequence, allowing an autoregressive large language model (LLM) to predict video tags based on text cues.
  • Progressive short-to-long training: By gradually increasing the length of training videos using a phased training strategy, the model can learn and generate more complex and coherent video content.
  • Loss reweighting: To address the loss imbalance issue in long video training, the loss of early frames is weighted to enhance the model's learning of those early frames.
  • Video tag re-encoding: During video inference, the predicted video tags are decoded into video frames in pixel space and then re-encoded to maintain the coherence and consistency of the video content.
  • Sampling strategy:Based on the Top-k sampling strategy, selection is made from the most likely labels, reducing the impact of potential errors on subsequent label generation and alleviating the problem of error accumulation.

Loong's project address

Loong's application scenarios

  • Entertainment and social mediaUsers can generate personalized long-form video content and share it on social media platforms, such as music videos, travel logs, and fun stories.
  • Film and video productionIn the initial creative stages of movie trailers, special effects production, or long-form video content, Loong quickly generates video sketches to help directors and producers explore different storylines and visual effects.
  • Advertising and MarketingBusinesses can generate engaging advertising videos to showcase their products or services in a more vivid way, increasing the appeal and memorability of their ads.
  • Education and trainingIn the field of education, L creates educational content, such as historical reenactments and scientific experiment simulations, to provide a more intuitive and interactive learning experience.
  • News and reportsNews organizations can quickly generate video summaries of news stories, improving the efficiency and appeal of their reporting.