AB
AiBoss
project

QVQ-72B-Preview - Alibaba Tongyi's open-source multimodal inference model

QVQ-72B-Preview is an open-source multimodal reasoning model from Alibaba Cloud's Tongyi Qianwen team, focusing on improving visual reasoning capabilities. The model performs exceptionally well in multiple benchmark tests, demonstrating powerful capabilities in multimodal understanding and reasoning tasks...

What is QVQ-72B-Preview?

QVQ-72B-Preview is an open-source multimodal reasoning model from Alibaba Cloud's Tongyi Qianwen team, focusing on improving visual reasoning capabilities. The model performs exceptionally well in multiple benchmark tests, demonstrating powerful capabilities in multimodal understanding and reasoning tasks. It can accurately understand image content, perform complex step-by-step reasoning, support inferring specific information such as object height and quantity from images, and recognize the deeper meaning of images, such as the connotations of memes.

Main functions of QVQ-72B-Preview

  • Powerful visual reasoning abilityThe QVQ-72B-Preview can accurately understand image content and perform complex step-by-step reasoning. It supports inferring specific information such as the height and quantity of objects from images and can recognize the deeper meaning of images, such as the connotations of "memes".
  • Multimodal processingThe model can process both image and text information simultaneously for deep reasoning. It can seamlessly integrate linguistic and visual information, making the AI's reasoning process more efficient.
  • Scientific-level reasoning performanceThe QVQ-72B-Preview excels at handling complex scientific problems, thinking like a scientist and providing accurate answers. It delivers more reliable and intelligent results by questioning hypotheses and optimizing reasoning steps.

QVQ-72B-Preview Performance Review

QVQ-72B-Preview was evaluated on the following four datasets:

  • MMMUA university-level multidisciplinary and multimodal assessment dataset that evaluates the model’s comprehensive understanding and reasoning abilities related to vision. The visual reasoning score is 70.3, which is at the university level.
  • MathVistaA math-centric visual reasoning test suite that evaluates logical reasoning using jigsaw puzzles, algebraic reasoning using function graphs, and scientific reasoning using numbers from academic papers. It surpasses OpenAI o1 and demonstrates powerful mathematical and graphical reasoning capabilities.
  • MathVisionA high-quality multimodal mathematical reasoning test set derived from real mathematical competitions, with greater problem diversity and subject breadth compared to MathVista, outperforming GPT-4o and Claude 3.5.
  • OlympiadBenchAn Olympiad-level bilingual, multimodal science benchmark set containing 8476 questions from Olympiad mathematics and physics competitions (including the Chinese National College Entrance Examination), outperforming GPT-4o and Claude 3.5.

QVQ-72B-Preview project address

Application scenarios of QVQ-72B-Preview

  • EducationIn the context of knowledge transmission and learning, QVQ-72B-Preview can help teachers and students solve difficult problems such as the derivation of complex mathematical formulas and the analysis of scientific experimental principles.
  • Scientific explorationWhen faced with challenging scientific problems requiring in-depth study, such as interpreting quantum mechanical phenomena in physics or constructing models of galaxy evolution in astronomy, QVQ-72B-Preview can help scientists uncover the truth hidden behind data and phenomena.
  • Multimodal interactionIn intelligent customer service responses to users' inquiries with both text and images, or in the precise classification and management of massive amounts of text and image information on social media platforms, the QVQ-72B-Preview can perfectly integrate image and text information to provide an ideal response that meets user needs.