AB
AiBoss
project

Ovis 1.6 - A multimodal large model launched by Alibaba's international AI team, surpassing the closed-source GPT-4o-mini.

Ovis 1.6 is a multimodal large-scale model launched by Alibaba's international AI team. It has achieved excellent results on OpenCompass, a leading comprehensive benchmark for multimodal computing, especially ranking first in overall score among models with fewer than 3 billion parameters...

What is Ovis 1.6?

Ovis 1.6 is a multimodal large-scale model launched by Alibaba's international AI team. It has achieved excellent results on OpenCompass, a leading multimodal benchmark, ranking first in overall score among models with fewer than 3 billion parameters, surpassing other mainstream models. The Ovis 1.6 model performs exceptionally well in multiple tasks, including mathematical reasoning and visual understanding, even outperforming the closed-source GPT-4o-mini model. Ovis 1.6 can handle various data inputs, including text and images, and possesses powerful multimodal task processing capabilities, including visual perception reasoning, mathematical and scientific problem solving, and understanding everyday scenarios.

Main features of Ovis 1.6

  • Mathematical Reasoning Questions and AnswersTo accurately answer a variety of mathematical questions, including complex mathematical formulas and logical reasoning.
  • Object recognitionThis ability to identify different objects, such as flower varieties, demonstrates its capacity for image recognition.
  • Text ExtractionSupports text extraction in multiple languages; Ovis 1.6 can identify and extract text information from various documents.
  • Complex task decisionIt processes and understands various types of data inputs to perform complex decision-making tasks, such as the comprehensive analysis of images and text.
  • Image understandingIt achieves state-of-the-art (SOTA) level performance in image understanding tasks, capable of handling high-resolution and extremely aspect ratio images.

Technical Principles of Ovis 1.6

  • Innovative architecture designOvis 1.6 is based on an architecture that combines a visual tokenizer, a visual embedding table, and a large language model. The design introduces a learnable visual embedding table, which converts continuous visual features into probabilistic visual tokens. Structured visual embeddings are then obtained through multiple indexing and weighting of the visual embedding table, improving performance on multimodal tasks.
  • High-resolution image processingOvis 1.6 supports processing images with extreme aspect ratios and is compatible with high-resolution images, enabling the model to demonstrate excellent capabilities in image understanding tasks.
  • Comprehensive data optimizationOvis 1.6 uses various datasets during training, including Caption, VQA, OCR, Table, Chart, etc., providing comprehensive data coverage and significantly improving the model's performance on tasks such as multimodal question answering and instruction following.
  • Superior model performanceOn OpenCompass, a leading multimodal benchmarking platform, Ovis1.6-Gemma2-9B achieved the top ranking among models with fewer than 30 parameters, demonstrating its excellent performance.

Project address for Ovis 1.6

Application scenarios of Ovis 1.6

  • Education and learning supportOvis 1.6 can accurately answer math questions, identify and interpret mathematical formulas, and as an educational tool, it can help students learn and understand complex concepts.
  • Agricultural and plant identificationWith its object recognition capabilities, Ovis 1.6 helps identify different plant varieties, playing an important role in fields such as agricultural research and plant protection.
  • Language translation and text processingIt supports text extraction and translation in multiple languages, making it suitable for cross-language communication, international business, and multilingual content creation.
  • Image recognition and analysisIt can recognize handwritten fonts and complex images, and is suitable for image content review, security monitoring and art analysis.
  • autonomous drivingIntegrating visual data improves the environmental perception and decision-making capabilities of autonomous vehicles, thereby enhancing driving safety.
  • Medical diagnosisIt assists doctors in analyzing medical images, improving the accuracy and efficiency of disease diagnosis.