AB
AiBoss
project

OmniAlign-V - A high-quality dataset jointly launched by Shanghai Jiao Tong University and Shanghai AI Lab, among others.

OmniAlign-V is a high-quality dataset jointly developed by Shanghai Jiao Tong University, Shanghai AI Lab, Nanjing University, Fudan University, and Zhejiang University. It is specifically designed to improve the alignment between multimodal large language models (MLLMs) and human preferences...

What is OmniAlign-V?

OmniAlign-V is a high-quality dataset jointly developed by Shanghai Jiao Tong University, Shanghai AI Lab, Nanjing University, Fudan University, and Zhejiang University. It is designed to improve the alignment ability of multimodal large language models (MLLMs) with human preferences. OmniAlign-V contains approximately 200,000 multimodal training samples, covering natural images and infographics, combined with open-ended, knowledge-rich question-answer pairs. OmniAlign-V's design emphasizes task diversity, including knowledge-based question answering, reasoning tasks, and creative tasks, enhancing the model's alignment ability based on complex questions and diverse answer formats. OmniAlign-V introduces an image selection strategy to ensure that semantically rich and complex images are used for data generation.

Main functions of OmniAlign-V

  • Provide high-quality multimodal training dataIt contains approximately 200,000 multimodal training samples, covering natural images and infographics (such as posters, charts, etc.), combined with complex questions and diverse answer formats, to help the model better understand human preferences and needs.
  • Enhance the model's open-ended question answering capabilities.The dataset design emphasizes open-ended questions, interdisciplinary knowledge, and comprehensive answers, enabling the model to generate responses that better align with human preferences.
  • Enhance the model's reasoning and creative abilities.The goal is to train models to engage in more complex thinking and creation, thereby improving their performance in multimodal interactions.
  • Optimize multimodal instruction tuningBased on high-quality instruction tuning data, it helps the model better follow human instructions and maintain basic capabilities (such as object recognition, OCR, etc.).
  • Support continuous optimization of multimodal modelsOmniAlign-V is used for supervised fine-tuning (SFT), and combined with direct preference optimization (DPO) to further improve the model's alignment capability.

OmniAlign-V's technical principles

  • Image filtering and classificationBased on Image Complexity (IC) scoring and Object Category (OC) filtering, semantically rich and complex images are selected. Images are classified into natural images and infographics, and different tasks are designed for different types of images.
  • Task Design and Data GenerationNatural image tasks include knowledge-based question answering, reasoning, and creative tasks, enhancing the model's ability to understand and generate data from real-world scenes. Infographic tasks are designed for specific tasks involving charts, posters, etc., requiring the model to understand and interpret complex information. High-quality question-answer pairs are generated using advanced models such as GPT-4o, and data quality is optimized through post-processing.
  • Post-processing optimizationThe generated question-and-answer pairs are post-processed, including instruction enhancement, reasoning enhancement, and infographic answer refinement, to ensure data diversity and high quality.
  • Multimodal training and optimizationThe alignment capability of the model is improved by using Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO). The dataset design emphasizes diversity and complexity, allowing the model to better understand human preferences in multimodal interactions.
  • Benchmarking and Evaluation: Introduce the MM-AlignBench benchmark test to evaluate the performance of MLLMs in human preference alignment and ensure the applicability of the model in real-world scenarios.

OmniAlign-V project address

Application scenarios of OmniAlign-V

  • Multimodal dialogue systemTo improve the quality of interaction between intelligent assistants and users, and to make answers more in line with human preferences.
  • Image-assisted question answeringIt combines image information to provide more comprehensive and accurate question-and-answer services, applicable to fields such as education and tourism.
  • Creative content generationIt helps users quickly generate high-quality creative text, such as advertising copy and story creation.
  • Education and learning supportTo provide students with richer learning materials to help them understand complex charts and illustrations.
  • Infographic InterpretationIt helps users interpret complex charts, provides background knowledge and reasoning results, and improves data understanding capabilities.