AB
AiBoss
project

OmniVision - A minimal parameter multimodal model optimized for edge devices.

OmniVision is a compact multimodal model with 968M parameters, optimized for edge devices. OmniVision can handle both visual and text input, and based on improvements to the LLaVA architecture, it significantly reduces the number of image tokens, lowers latency, and...

What is OmniVision?

OmniVision is a compact multimodal model with 968M parameters, optimized for edge devices. It handles both visual and text input, and based on improvements to the LLaVA architecture, it significantly reduces the number of image tokens, lowering latency and computational costs. Trained on trusted data using DPO, OmniVision delivers more reliable results, making it suitable for tasks such as visual question answering and image captioning.

OmniVision's main functions

  • Visual Question AnsweringOmniVision can understand image content and provide accurate answers to questions posed by images.
  • Image CaptioningThe model can generate text describing the content of an image.
  • End-to-end visual language understandingBased on an integrated visual encoder and language model, OmniVision achieves seamless conversion from images to text, understanding image content and expressing it in natural language.
  • Optimize edge deploymentOptimized for edge devices, reducing the demand for computing resources, allowing the model to run in resource-constrained environments.

OmniVision's technical principles

  • Compact multimodal architectureOmniVision combines the basic language model Qwen2.5-0.5B-Instruct and the visual encoder SigLIP-400M, and uses an MLP projection layer to align the image embedding with the text tag space to achieve end-to-end visual language understanding.
  • Efficient Token ProcessingBased on technological innovation, OmniVision significantly reduces the number of image tokens, thereby lowering the computational cost and latency of the model while maintaining model performance.
  • Precise training strategiesBased on a three-stage training process, including pre-training, supervised fine-tuning, and direct preference optimization, the accuracy of the model's understanding and response to vision and language is improved.

OmniVision's project address

OmniVision application scenarios

  • Visual Question AnsweringWhen users ask questions about the content of an image, OmniVision can understand the questions and provide accurate answers based on the image content.
  • Image CaptioningThe model can automatically generate descriptive text for images, making it suitable for fields such as social media, content management, and image archiving.
  • Content moderationUsing its visual and text understanding capabilities, OmniVision can assist in content moderation of images and text, identifying inappropriate content.
  • Assisted visual searchOn e-commerce platforms or in image databases, users search for specific images based on descriptions, and OmniVision can understand the descriptions and match relevant images.
  • Smart assistants and chatbotsWhen integrated into chatbots, OmniVision can understand the images and text messages sent by users, providing a richer and more accurate interactive experience.