AB
AiBoss
project

WorldSense - A new benchmark for comprehensive multimodal evaluation launched by Xiaohongshu in collaboration with Shanghai Jiao Tong University.

WorldSense, developed by Xiaohongshu and Shanghai Jiao Tong University, is a benchmark test used to evaluate the comprehensive understanding capabilities of multimodal large language models (MLLMs) of visual, auditory, and text input in real-world scenarios. WorldSense...

What is WorldSense?

WorldSense, developed by Xiaohongshu and Shanghai Jiao Tong University, is a benchmark test used to evaluate the comprehensive understanding capabilities of multimodal large language models (MLLMs) of visual, auditory, and text input in real-world scenarios. WorldSense includes 1662 diverse audio-video synchronized videos covering 8 major domains and 67 subcategories, as well as 3172 multiple-choice question-answer pairs, involving 26 different cognitive tasks. WorldSense emphasizes the tight coupling of audio and video information; all questions require both modalities to arrive at the correct answer. WorldSense's high-quality annotations are manually completed by 80 expert annotators and undergo multiple rounds of verification to ensure accuracy and reliability.

WorldSense's main functions

  • Multimodal collaborative assessmentThis approach emphasizes the tight coupling of audio and video information, designing questions that require both visual and auditory information to answer correctly. It rigorously tests the model's understanding capabilities under multimodal inputs to ensure the model can effectively integrate information from different modalities for accurate comprehension.
  • Diverse video and task coverageWorldSense includes 1,662 diverse audio-video synchronized videos covering 8 main domains and 67 subcategories, as well as 3,172 multiple-choice question-and-answer pairs covering 26 different cognitive tasks.
  • High-quality annotation and verificationAll question-and-answer pairs were manually annotated by 80 expert annotators and underwent multiple rounds of verification, including manual review and automatic model verification, to ensure the accuracy and reliability of the annotations.

The technical principles of WorldSense

  • Multimodal input processingWorldSense requires its models to process video, audio, and text input simultaneously. Synchronization of video and audio ensures the model can capture the correlation between visual and auditory information, leading to a more comprehensive understanding of the scene. Multimodal input processing capability is key to evaluating whether a model can handle complex environments like a human.
  • Task design and annotation:Based on carefully designed question-answer pairs, each question requires the integration of multimodal information to arrive at the correct answer. The annotation process involves multiple rounds of manual review and automated verification to ensure the reasonableness of the questions and the accuracy of the annotations.
  • Multimodal fusion and reasoningBased on diverse task designs, this study evaluates the model's multimodal understanding capabilities at different levels, including basic perception (such as detection of audio and visual elements), comprehension (grasping multimodal relationships), and reasoning (such as causal inference and abstract thinking). This multi-level evaluation method comprehensively tests the model's multimodal fusion and reasoning abilities.
  • Data collection and filteringWorldSense's data collection process includes filtering video clips with strong audio-visual correlations from a large-scale video dataset, ensuring the quality and diversity of video content through human review, and ensuring that benchmarks cover a wide range of real-world scenarios.

WorldSense project address

Application Scenarios of WorldSense

  • autonomous drivingIt helps autonomous driving systems better understand visual and auditory information in the traffic environment, improving the accuracy of decision-making.
  • Smart EducationTo assess and improve the ability of educational tools to understand instructional video content and to support personalized learning.
  • Intelligent monitoringTo enhance the monitoring system's ability to perceive and understand visual and audio information in videos, thereby improving the effectiveness of security detection.
  • Intelligent Customer Service: Evaluate the intelligent customer service system's ability to understand users' voice, facial expressions, and text input, and optimize the interactive experience.
  • Content creationIt helps multimedia content creation and analysis systems understand video content more intelligently, improving creation and recommendation efficiency.