MMBench-Video - A long-form video understanding benchmark jointly launched by Shanghai AI Lab and several universities.
MMBench-Video is a novel long-video, multi-question answering benchmark jointly developed by Zhejiang University, Shanghai Artificial Intelligence Laboratory, Shanghai Jiao Tong University, and the Chinese University of Hong Kong. MMBench-Video can comprehensively evaluate large-scale visual...
What is MMBench-Video?
MMBench-Video is a novel long-video, multi-question-answering benchmark jointly developed by Zhejiang University, Shanghai Artificial Intelligence Laboratory, Shanghai Jiao Tong University, and the Chinese University of Hong Kong. MMBench-Video comprehensively evaluates the video understanding capabilities of large-scale visual language models (LVLMs) by using long videos rich in content and fine-grained ability assessments, thus addressing the shortcomings of existing benchmarks in temporal understanding and complex task processing. MMBench-Video includes approximately 600 YouTube video clips covering 16 categories, with each video ranging from 30 seconds to 6 minutes in length, accompanied by high-quality question-answer pairs written by volunteers. The benchmark uses GPT-4 for automated evaluation, improving accuracy and maintaining consistency with human judgment. The launch of MMBench-Video provides researchers with a powerful tool to evaluate and improve video language models.
Main functions of MMBench-Video
- Video understanding assessmentMMBench-Video is used to evaluate the ability of large visual language models (LVLMs) to understand long video content.
- Multi-scenario coverageIt contains video content in 16 main categories, covering a wide range of themes and scenarios.
- Fine-grained capability assessmentThe model's video understanding capabilities are comprehensively evaluated using 26 fine-grained capability dimensions.
- High-quality datasetsThe video clips and question-and-answer pairs were carefully written and labeled by volunteers to ensure data quality.
- Automated evaluationUse GPT-4 for automated assessment to improve the efficiency and accuracy of the assessment.
The technical principles of MMBench-Video
- Long video contentMMBench-Video contains multiple long video clips collected from YouTube. These video clips are better able to test the model's temporal understanding ability than traditional short videos.
- Manual annotationThe questions and answers are written and labeled by human volunteers to ensure high quality and reduce bias.
- Ability classification systemWe construct a three-tiered classification system for video understanding capabilities, including two main categories: perception and reasoning, and 26 more detailed capability dimensions.
- Temporal Reasoning ChallengeDesign problems that require temporal reasoning ability to evaluate the model's understanding of the temporal dimension of video content.
- Automated evaluationLanguage models (such as GPT-4) automatically evaluate the semantic similarity between the model's output and the standard answer, thus assessing the model's performance.
- Multi-model comparisonIt supports scoring and comparing multiple LVLMs to determine their strengths and weaknesses in video understanding tasks.
MMBench-Video project address
- Project official website:mmbench-video.github.io
- GitHub repository:https://github.com/open-compass/VLMEvalKit
- HuggingFace model library:https://huggingface.co/datasets/opencompass/MMBench-Video
- arXiv technical paper:https://arxiv.org/pdf/2406.14515
Application scenarios of MMBench-Video
- Model Evaluation and ComparisonResearchers assessed and compared the abilities of different LVLMs in video understanding, including perception and reasoning skills.
- Model optimization and trainingDevelopers optimized the model's architecture and training process based on the evaluation results of MMBench-Video, improving the model's ability to understand video content.
- Academic exchange and publicationAs a tool for academic exchange, it helps researchers demonstrate the performance of their models and publish their research findings in academic conferences or journals.
- Multimodal learning researchMMBench-Video provides a rich dataset for researching and developing multimodal learning algorithms, especially for tasks involving video and text understanding.
- Intelligent video analytics applicationsIn fields such as intelligent video surveillance, content filtering, automatic summarization, and video recommendation, it helps developers train and test more accurate video analysis models.