AB
AiBoss
project

Ivy-VL - A lightweight multimodal model open-sourced by AI Safeguard in collaboration with Carnegie Mellon and Stanford.

Ivy-VL is a lightweight multimodal AI model developed by AI Safeguard in collaboration with Carnegie Mellon University and Stanford University, specifically designed for mobile and edge devices. The model boasts 3 billion parameters, significantly reducing the size compared to other large multimodal models...

What is Ivy-VL?

Ivy-VL is a lightweight multimodal AI model developed by AI Safeguard in collaboration with Carnegie Mellon University and Stanford University, specifically designed for mobile and edge devices. With only 3 bytes of parameters, it significantly reduces computational resource requirements compared to other large multimodal models, enabling efficient operation on resource-constrained devices such as AI glasses and smartphones. Ivy-VL demonstrates outstanding performance in multimodal tasks such as visual question answering, image captioning, and complex reasoning, achieving the best score among models with fewer than 4 bytes in the OpenCompass benchmark.

Main functions of Ivy-VL

  • Visual Q&A: To understand and answer questions related to the content of images.
  • Image DescriptionThe model can generate text describing the content of an image.
  • Complex Reasoning: Handling visual tasks involving multi-step reasoning.
  • Multimodal data processingIn smart home and Internet of Things (IoT) devices, it processes and understands data from different modalities, such as vision and language.
  • Augmented Reality (AR) ExperienceIn smart wearable devices, real-time visual question answering is supported to enhance the AR experience.

Ivy-VL Technical Principles

  • Lightweight designIvy-VL has only 3B parameters, making it more efficient on resource-constrained devices.
  • Multimodal fusion technologyIvy-VL combines an advanced visual encoder and a powerful language model to achieve effective information fusion between different modalities.
  • Visual encoderUse Googlegoogle/siglip-so400m-patch14-384Visual encoders process and understand image information.
  • Language Model: combinationQwen2.5-3B-InstructLanguage models understand and generate text information.
  • Optimized dataset training: Improve model performance on multimodal tasks by training on carefully selected and optimized datasets.

Ivy-VL project address

Application scenarios of Ivy-VL

  • Smart wearable devicesIt provides real-time visual question-and-answer functionality to help users obtain information in augmented reality (AR) environments.
  • Mobile Smart AssistantIt provides more intelligent multimodal interaction capabilities, such as image recognition and voice interaction, to enhance the user experience.
  • Internet of Things (IoT) devicesEnables efficient multimodal data processing in smart home and IoT scenarios, such as controlling home devices with images and voice.
  • Mobile Education and EntertainmentEnhance image understanding and interaction capabilities in educational software to drive mobile learning and immersive entertainment experiences.
  • Visual question answering systemIn places like museums and exhibition centers, users can ask questions by taking photos, and the system will provide relevant information.