AB
AiBoss
project

HourVideo - A benchmark dataset for long-form video understanding developed by Fei-Fei Li and Jia-Jun Wu's team.

HourVideo is a long-form video understanding benchmark dataset developed by Fei-Fei Li and Jia-Jun Wu's team at Stanford University. It contains 500 first-person perspective videos, ranging in length from 20 to 120 minutes, covering 77 daily activities, and can evaluate the performance of multimodal models...

What is HourVideo?

HourVideo is a long-form video understanding benchmark dataset developed by Fei-Fei Li and Jia-Jun Wu's team at Stanford University. It contains 500 first-person perspective videos, ranging in length from 20 to 120 minutes, covering 77 daily activities. It evaluates the ability of multimodal models to understand long videos. Based on a series of tasks, such as summarizing, perception, visual reasoning, and navigation, the dataset tests the model's ability to identify and synthesize information from multiple time segments within the video, thus advancing the development of long-form video understanding technology.

HourVideo's main functions

  • Long video comprehension assessmentBased on videos up to one hour long, HourVideo can test a model’s ability to understand long-term visual data streams.
  • Multi-task test suiteThe dataset includes various tasks such as summarizing, perception, visual reasoning, and navigation, comprehensively evaluating the model's performance in different video language understanding aspects.
  • High-quality question generation: 12,976 multiple-choice questions generated by human annotators and large language models (LLMs) to provide standardized test questions.
  • Model performance comparison: Compare with other multimodal models to evaluate the performance of different models on long video understanding tasks.

HourVideo's technical principles

  • Video dataset constructionHourVideo selected 500 first-person perspective videos from the Ego4D dataset, covering daily activities, with video lengths ranging from 20 to 120 minutes.
  • Task Suite DesignDesign a task suite containing multiple sub-tasks, each task requiring the model to understand and infer long-term dependencies in video content.
  • Problem Prototype DevelopmentFor each task, a prototype question is designed to ensure that correct answers require information identification and synthesis from multiple time segments of the video.
  • Data generation processBased on a multi-stage data generation process, including video screening, question generation, human feedback optimization, blind screening, and expert optimization, high-quality multiple-choice questions are generated.

HourVideo's project address

Application scenarios of HourVideo

  • Multimodal Artificial Intelligence ResearchResearch and development of multimodal models for understanding long-duration continuous video content.
  • Autonomous Agent and Assistant System: To help develop autonomous agents and virtual assistants that can understand long-term visual information and make decisions.
  • Augmented Reality (AR) and Virtual Reality (VR)Provides the technological foundation for creating immersive AR/VR experiences that understand and adapt to user behavior.
  • Video content analysisAnalyze and understand video content, such as surveillance videos, news reports, and educational videos, to extract key information and insights.
  • Robot VisionTo enable robots to understand long-term visual information and improve their navigation and manipulation capabilities in complex environments.