MMSI-Video-Bench - A spatial intelligent video benchmark launched by Shanghai AI Lab
MMSI-Video-Bench is a benchmark tool for evaluating the capabilities of Multimodal Large Language Models (MLLMs) in video spatial intelligence. Developed jointly by the Shanghai Artificial Intelligence Laboratory and several universities, it comprehensively evaluates models in real-world scenarios...
What is MMSI-Video-Bench?
MMSI-Video-Bench is a benchmark tool for evaluating the capabilities of Multimodal Large Language Models (MLLMs) in video spatial intelligence. Developed jointly by the Shanghai Artificial Intelligence Laboratory and several universities, it comprehensively assesses the models' spatial understanding and reasoning abilities in the real physical world. The benchmark includes 1278 video clips from 25 public datasets and 1 self-built dataset, covering various complex scenarios such as indoor scenes, outdoor street scenes, and robot operations. The problems were meticulously designed by 11 3D vision researchers to ensure high challenge and accuracy. MMSI-Video-Bench employs a multi-level task design, encompassing spatial perception, motion understanding, planning, prediction, and cross-video reasoning capabilities, comprehensively examining the model's video understanding and decision-making abilities.
Main functions of MMSI-Video-Bench
-
Multimodal capability assessmentIt is a benchmark tool specifically designed to evaluate the performance of multimodal large language models (MLLMs) in video spatial intelligence, and can comprehensively measure the model's understanding and reasoning ability of video content.
-
Diverse datasetsIt contains 1,278 video clips from 25 public datasets and 140 anonymized internal videos, covering a variety of complex scenarios such as indoor scenes, outdoor street scenes, and robot operations, ensuring the diversity and richness of the data.
-
High-quality annotationAll questions were designed and annotated by 3D vision experts, and each question is accompanied by detailed explanatory reasons to ensure the accuracy and high quality of the annotations.
-
Comprehensive task designThrough a multi-layered task framework, covering capabilities such as spatial perception, motion understanding, planning, prediction, and cross-video reasoning, the model's performance in video spatial intelligence is comprehensively examined.
-
Model performance measurementIt provides detailed evaluation results for 25 open-source and proprietary MLLMs, helping researchers and developers understand the strengths and weaknesses of the models and guiding model improvement and optimization.
MMSI-Video-Bench Technical Principles
-
Real-world scenario drivenIt uses dynamic video data from the real physical world, eliminating the reliance on template generation and building a test environment full of uncertainty and diversity.
-
Multimodal fusionThe model integrates visual, linguistic, and other modal information from videos, requiring it to accurately capture the occurrence nodes and spatial relationships of key events in the spatiotemporal dimensions.
-
Task DesignBased on a four-level framework of perception, planning, prediction, and cross-video reasoning, a multi-dimensional reasoning task covering cross-time, cross-viewpoint, and cross-object was designed.
-
Expert annotationEach question is carefully designed and reviewed by 3D vision experts to ensure its accuracy and unambiguity.
-
Dynamic testing environmentBy introducing natural behaviors and physical laws from real-world scenarios to generate problems, the model is forced to deeply understand the spatial relationships, motion trajectories, and underlying causal logic of objects in the video.
-
Fine-grained annotation systemA fine-grained annotation system was established, covering multi-level cognitive tasks ranging from basic spatial relationships to high-order causal reasoning.
MMSI-Video-Bench project address
- Project official website: https://rbler1234.github.io/MMSI-VIdeo-Bench.github.io/
- Github repositoryhttps://github.com/InternRobotics/MMSI-Video-Bench
- Huggingface model libraryhttps://huggingface.co/datasets/rbler/MMSI-Video-Bench
- arXiv technical paper: https://arxiv.org/pdf/2512.10863
Application scenarios of MMSI-Video-Bench
-
Model performance evaluationIt is used to comprehensively evaluate the performance of multimodal large language models (MLLMs) in video understanding tasks, helping researchers and developers understand the strengths and weaknesses of the models.
-
academic researchIt provides the academic community with a standardized testing platform for researching and improving the performance of multimodal models in video spatial intelligence.
-
Technology DevelopmentIt helps developers optimize and improve multimodal models, especially in key capabilities such as spatial perception, motion understanding, planning, and prediction.
-
Industry application testingIt is suitable for fields such as autonomous driving, robot navigation, and intelligent monitoring, and is used to test the performance of models in real-world application scenarios.
-
Education and TrainingAs a teaching resource, it helps students and researchers better understand and practice multimodal video understanding technology.
-
Model Comparison AnalysisIt provides a unified test benchmark for different multimodal models, facilitating cross-sectional comparisons and performance analysis.