AB
AiBoss
project

MMBench-Video - A long-form video understanding benchmark jointly launched by Shanghai AI Lab and several universities.

MMBench-Video is a novel long-video, multi-question answering benchmark jointly developed by Zhejiang University, Shanghai Artificial Intelligence Laboratory, Shanghai Jiao Tong University, and the Chinese University of Hong Kong. MMBench-Video can comprehensively evaluate large-scale visual...

What is MMBench-Video?

MMBench-Video is a novel long-video, multi-question-answering benchmark jointly developed by Zhejiang University, Shanghai Artificial Intelligence Laboratory, Shanghai Jiao Tong University, and the Chinese University of Hong Kong. MMBench-Video comprehensively evaluates the video understanding capabilities of large-scale visual language models (LVLMs) by using long videos rich in content and fine-grained ability assessments, thus addressing the shortcomings of existing benchmarks in temporal understanding and complex task processing. MMBench-Video includes approximately 600 YouTube video clips covering 16 categories, with each video ranging from 30 seconds to 6 minutes in length, accompanied by high-quality question-answer pairs written by volunteers. The benchmark uses GPT-4 for automated evaluation, improving accuracy and maintaining consistency with human judgment. The launch of MMBench-Video provides researchers with a powerful tool to evaluate and improve video language models.

Main functions of MMBench-Video

  • Video understanding assessmentMMBench-Video is used to evaluate the ability of large visual language models (LVLMs) to understand long video content.
  • Multi-scenario coverageIt contains video content in 16 main categories, covering a wide range of themes and scenarios.
  • Fine-grained capability assessmentThe model's video understanding capabilities are comprehensively evaluated using 26 fine-grained capability dimensions.
  • High-quality datasetsThe video clips and question-and-answer pairs were carefully written and labeled by volunteers to ensure data quality.
  • Automated evaluationUse GPT-4 for automated assessment to improve the efficiency and accuracy of the assessment.

The technical principles of MMBench-Video

  • Long video contentMMBench-Video contains multiple long video clips collected from YouTube. These video clips are better able to test the model's temporal understanding ability than traditional short videos.
  • Manual annotationThe questions and answers are written and labeled by human volunteers to ensure high quality and reduce bias.
  • Ability classification systemWe construct a three-tiered classification system for video understanding capabilities, including two main categories: perception and reasoning, and 26 more detailed capability dimensions.
  • Temporal Reasoning ChallengeDesign problems that require temporal reasoning ability to evaluate the model's understanding of the temporal dimension of video content.
  • Automated evaluationLanguage models (such as GPT-4) automatically evaluate the semantic similarity between the model's output and the standard answer, thus assessing the model's performance.
  • Multi-model comparisonIt supports scoring and comparing multiple LVLMs to determine their strengths and weaknesses in video understanding tasks.

MMBench-Video project address

Application scenarios of MMBench-Video

  • Model Evaluation and ComparisonResearchers assessed and compared the abilities of different LVLMs in video understanding, including perception and reasoning skills.
  • Model optimization and trainingDevelopers optimized the model's architecture and training process based on the evaluation results of MMBench-Video, improving the model's ability to understand video content.
  • Academic exchange and publicationAs a tool for academic exchange, it helps researchers demonstrate the performance of their models and publish their research findings in academic conferences or journals.
  • Multimodal learning researchMMBench-Video provides a rich dataset for researching and developing multimodal learning algorithms, especially for tasks involving video and text understanding.
  • Intelligent video analytics applicationsIn fields such as intelligent video surveillance, content filtering, automatic summarization, and video recommendation, it helps developers train and test more accurate video analysis models.