BabyVision - A multimodal understanding evaluation dataset launched by the UniPat AI team.
BabyVision is a multimodal understanding evaluation dataset launched by the UniPat AI team, evaluating the performance of multimodal language models (MLLMs) and image generation models on visual reasoning tasks. It includes two main tracks: MLLM evaluation and generation...
What is BabyVision?
BabyVision is a multimodal understanding evaluation dataset launched by the UniPat AI team, evaluating the performance of multimodal language models (MLLMs) and image generation models on visual reasoning tasks. It includes two main tracks: MLLM evaluation and generation evaluation. The evaluation dataset is designed with four main visual ability categories: fine discrimination, visual tracking, spatial perception, and visual pattern recognition, comprising 22 sub-tasks and a total of 388 questions. These tasks strictly control for language dependencies to realistically reflect the model's visual understanding capabilities.
BabyVision's main functions
-
Evaluate the visual reasoning ability of multimodal modelsBy designing rigorous visual tasks, we tested the performance of multimodal language models (MLLMs) and image generation models in purely visual scenarios, revealing the shortcomings of these models in visual understanding.
-
Two evaluation tracks are provided.One is an MLLM evaluation for multimodal language models, and the other is a generative evaluation for image generation models, comprehensively covering different types of multimodal models.
-
Covering four major categories of visual abilitiesIt includes fine discrimination, visual tracking, spatial perception, and visual pattern recognition. Through diverse task design, it comprehensively evaluates the model's reasoning ability in different visual scenarios.
-
Strictly control language dependenciesThis ensures that tasks cannot be solved through language prompts, thus truly reflecting the model's visual understanding capabilities and preventing the model from relying on language prompts to complete tasks.
-
Provides detailed evaluation results and rankingsIt demonstrates the performance of different models through metrics such as accuracy and compares them with human baselines, providing researchers with an intuitive reference.
-
Supports fast startup and flexible configurationIt provides complete datasets, evaluation scripts, and detailed documentation, making it easy for researchers to get started quickly and flexibly configure evaluation parameters through environmental variables and other means.
-
Promote the development of multimodal technologyBy revealing the shortcomings of current models, this study provides direction for future technological optimization and innovation, and promotes further improvement of multimodal models in vision tasks.
BabyVision's review results
-
Human baseline performance is excellentThe average accuracy rate of human testers was as high as 94.1%, demonstrating the powerful ability of humans in visual reasoning tasks.
-
The performance of closed-source models varies.Gemini3-Pro-Preview leads with an accuracy rate of 49.7%, GPT-5.2 has 34.4%, and Doubao-Seed-1.8 has 30.2%, but overall they are still far below human levels.
-
The gap between open source models is obvious.The accuracy of Qwen3-VL-Plus is only 19.2%, and most open-source models perform poorly, significantly lagging behind human baselines and some closed-source models.
-
The model has shortcomings in visual tasks.Regardless of whether the model is closed-source or open-source, it generally performs poorly on visual tasks that require continuous tracking, spatial imagination, and geometric induction, exposing the shortcomings of current multimodal models in terms of basic visual capabilities.
-
Generative evaluation results are not idealIn generative tasks, although some models exhibit "more human-like" behavior, the overall model still lacks the ability to consistently achieve a completely correct solution.
-
Evaluation results drive technological improvementsBy clearly pointing out the shortcomings of the model, BabyVision provides an important reference direction for the optimization and technological innovation of future multimodal models.
BabyVision's project address
- Github repositoryhttps://github.com/UniPat-AI/BabyVision
Application scenarios of BabyVision
-
Multimodal model evaluationIt is used to systematically evaluate the performance of multimodal language models and image generation models in visual reasoning tasks, helping researchers understand the visual understanding capabilities of the models.
-
Technology Research and DevelopmentIt provides AI researchers with a standardized testing platform for developing and optimizing multimodal models, thereby advancing visual reasoning technology.
-
Model performance comparisonBy using unified evaluation criteria, we can compare the performance of different models on visual tasks, providing a reference for model selection and improvement.
-
Education and learning toolsThis provides educators and students with a tool to understand multimodal AI vision capabilities for use in teaching and research activities.
-
Industry Application ReferenceIt provides a reference for model performance for industries that require multimodal visual reasoning capabilities (such as autonomous driving and medical image analysis), and helps the development and optimization of industry applications.
-
Academic research and publicationIt provides data support for academic research, helps researchers publish relevant research results, and promotes academic development in the field of multimodal AI.