AB
AiBoss
project

Aya Vision - Cohere launches multimodal, multilingual visual models

Aya Vision, developed by Cohere, is a multimodal, multilingual vision model that enhances multilingual and multimodal communication capabilities globally. It supports 23 languages and can perform image description generation, visual question answering, text translation, and more.

What is Aya Vision?

Aya Vision, developed by Cohere, is a multimodal, multilingual vision model that enhances multilingual and multimodal communication capabilities globally. It supports 23 languages and can perform tasks such as image caption generation, visual question answering, text translation, and multilingual summarization. Aya Vision is available in two versions: Aya Vision 32B and Aya Vision 8B, each with its own advantages in performance and computational efficiency. The model is trained using synthetic annotation and multilingual data augmentation techniques, enabling high efficiency even with limited resources.

Main functions of Aya Vision

  • Image description generationAya Vision can generate accurate and detailed descriptive text based on the input image, helping users quickly understand the image content. It is suitable for visually impaired people or scenarios that require rapid extraction of image information.
  • Visual Question Answering (VQA)Users can upload images and ask questions related to them. Aya Vision combines visual information and language understanding to provide accurate answers.
  • Multilingual supportAya Vision supports 23 major languages and can handle multilingual text input and output. It can generate image descriptions, answer questions, or translate text in different language environments, breaking down language barriers.
  • Text translation and summarizationAya Vision can translate text content and generate concise summaries, helping users quickly obtain key information.
  • Cross-modal understanding and generationAya Vision combines visual and linguistic information to enable cross-modal interaction. For example, it can convert image content into text descriptions or text instructions into visual search results.

Aya Vision's technical principles

  • Multimodal architectureAya Vision employs a modular architecture, comprising a visual encoder, a visual language connector, and a language model decoder. The visual encoder, based on SigLIP2-patch14-384, is responsible for extracting image features; the visual language connector maps image features to the embedding space of the language model; and the decoder generates text output.
  • Synthetic annotation and data augmentationTo improve multilingual performance, Aya Vision is trained using synthetic annotations (AI-generated annotations). These annotations are processed through translation and restatement, enhancing the quality of multilingual data. The model employs dynamic image resolution processing and pixel shuffling downsampling techniques to improve computational efficiency.
  • Two-stage training processAya Vision's training consists of two phases: visual-language alignment and supervised fine-tuning. The first phase aligns the visual and language representations, while the second phase jointly trains the connector and language model on a multimodal task.
  • High-performance computingAya Vision has a smaller parameter size (8B and 32B), but its performance outperforms larger models, such as Llama-3.2 90B Vision, in multiple benchmark tests. This is due to its efficient training strategy and optimization of computational resources.

Aya Vision's project address

Application Scenarios of Aya Vision

  • EducationAya Vision helps students and teachers better understand visual content. For example, through image description features, students can quickly understand the style and origin of artworks.
  • Content creationAya Vision can generate image descriptions for multilingual websites, enhancing the user experience. It can also be used to generate creative content such as news reports, stories, or poems.
  • auxiliary toolsAya Vision can be used as an assistive tool to help visually impaired people understand their surroundings through image descriptions.
  • Multilingual translation and communicationAya Vision supports text translation and summarization in 23 languages, helping users communicate across language barriers.
  • Research and DevelopmentResearchers can explore new application scenarios based on its efficiency and multilingual support capabilities.