AB
AiBoss
project

MV-MATH - A benchmark dataset launched by the Chinese Academy of Sciences to evaluate the mathematical reasoning ability of models in processing multi-visual information.

MV-MATH is a new benchmark dataset proposed by the Institute of Automation, Chinese Academy of Sciences, to evaluate the mathematical reasoning capabilities of multimodal large language models (MLLMs) in multi-visual scenes. The dataset contains 2009 high-quality mathematical problems, each...

What is MV-MATH?

MV-MATH is a new benchmark dataset proposed by the Institute of Automation, Chinese Academy of Sciences, to evaluate the mathematical reasoning ability of multimodal large language models (MLLMs) in multi-visual scenes. The dataset contains 2009 high-quality mathematical problems, each combining multiple images and text to form a multi-visual scene with interwoven images and text. The problems are divided into three types: multiple choice, fill-in-the-blank, and multi-step question-and-answer, covering 11 mathematical domains, including analytic geometry, algebra, metric geometry, combinatorics, transformation geometry, logic, solid geometry, arithmetic, combinatorial geometry, descriptive geometry, and statistics, and are divided into three difficulty levels.

Main functions of MV-MATH

  • Multi-view scene reasoningEach problem contains multiple images (2-8 images), which are interwoven with text to form complex scenes that are closer to real-world mathematical problems, allowing for a comprehensive evaluation of the model's reasoning ability to process multi-visual information.
  • Diverse mathematical fieldsIt covers 11 mathematical fields (such as analytic geometry, algebra, solid geometry, etc.) and 3 difficulty levels, and can comprehensively evaluate the reasoning performance of the model in different fields.
  • Image correlation analysisThis paper introduces image relevance labels for the first time, dividing the dataset into interdependent sets (MD) and independent sets (ID), which can be used to evaluate the model's reasoning ability when processing related and independent images respectively.
  • Educational applicationsBased on real-world K-12 education scenarios, it can be used to develop intelligent tutoring systems that help students solve complex math problems through a combination of text and graphics.
  • Research toolsIt provides standardized evaluation tools for multimodal learning research, helping researchers identify and improve the performance gaps of models in mathematical reasoning.
  • High-quality annotationEach sample is cross-validated by at least two annotators and includes questions, answers, detailed analysis, and image correlation annotations to provide comprehensive information for model evaluation.
  • Collection of real questionsThe questions are all derived from real-world scenarios, ensuring the usability and reliability of the dataset.

MV-MATH technical principles

  • Mutually Dependent Set (MD)Images are interconnected; understanding one image requires referring to other images.
  • Independent Set (ID)The images are independent of each other and can be interpreted individually.

MV-MATH project address

Application scenarios of MV-MATH

  • Intelligent tutoring systemThe MV-MATH dataset can be used to develop intelligent tutoring systems that help students solve complex mathematical problems through a combination of text and graphics.
  • Multimodal learning researchMV-MATH provides a standardized evaluation tool for multimodal learning research. Researchers can use datasets to evaluate the mathematical reasoning capabilities of multimodal large language models (MLLMs) in multi-visual scenarios, thus advancing the development of multimodal learning technologies.
  • Performance gap analysisThrough extensive experiments, researchers can identify and improve the performance gaps of models in mathematical reasoning.
  • Multi-graph reasoning taskThe dataset can be used to develop and optimize solutions for multi-graph reasoning tasks, handling multiple images and text information in complex mathematical problems.
  • Automated evaluation systemData sets can be used to evaluate and optimize automated testing systems, ensuring their accuracy and reliability when handling multimodal inputs.