Maya - an open-source, multilingual, multimodal model capable of handling and understanding eight different languages.
Maya is an open-source, multilingual, multimodal model that uses instruction-based fine-tuning to extend its capabilities across multiple languages and cultural contexts. Maya is built on the LLaVA framework and includes a newly created pre-trained dataset containing eight languages, improving visual...
What is Maya?
Maya is an open-source, multilingual, multimodal model that uses instruction-based fine-tuning to extend its capabilities across multiple languages and cultural contexts. Built on the LLaVA framework, Maya includes newly created pre-trained datasets in eight languages, improving cultural and language understanding in vision-language tasks. Maya employs toxicity analysis and dataset filtering to ensure the safety and quality of training data, supporting multiple languages including Chinese, French, Spanish, Russian, Hindi, Japanese, and Arabic, and is dedicated to improving the quality of AI content generation for low-resource languages.
Maya's main functions
- Multilingual supportMaya can handle and understand eight different languages, including Chinese, French, Spanish, Russian, Hindi, Japanese, Arabic, and English, with enhanced support for low-resource languages.
- Multimodal capabilitiesBy combining image and text data, machines can understand the visual world through natural language and perform tasks such as image description and answering visual questions.
- Command fine-tuningBased on instruction fine-tuning, it can better understand and respond to natural language instructions, improving performance and adaptability in practical applications.
- Dataset creation and toxicity filteringCreate a multilingual image-text pre-trained dataset, perform toxicity analysis and filtering to ensure data security and quality.
- Cross-cultural understandingBased on multilingual and multimodal data, we can better understand and process visual and linguistic information from different cultural backgrounds.
Maya's technical principles
- Model ArchitectureBased on the LLaVA 1.5 architecture, using the Aya-23 8B model as the multilingual language model (LLM) and SigLIP as the visual encoder, it supports multilingual and multimodal input.
- pre-trained datasetCreate a multilingual image-text pre-trained dataset containing 558,000 images in eight languages to support the development of multilingual visual language models.
- Toxicity analysisWe used LLaVAGuard 7B and Toxic-BERT to perform toxicity analysis on images and text in the dataset, identifying and filtering out unsafe or harmful content.
- Pre-training and fine-tuning:
- Pre-trainingThe image features are converted into language features using the projection matrix W, and pre-trained based on multi-turn dialogue data to optimize the alignment of images and text.
- Fine-tuningFine-tuning was performed on the PALO 150K instruction fine-tuning dataset to further improve the model's understanding and response to instructions.
- Cross-modal alignmentBased on the projection matrix and training strategy, the alignment between image features and language features is optimized to improve the model's performance in vision-language tasks.
Maya's project address
- GitHub repository:https://github.com/nahidalam/maya
- HuggingFace model library:https://huggingface.co/maya-multimodal/maya
- arXiv technical paper:https://arxiv.org/pdf/2412.07112
Maya's application scenarios
- Cross-language content understandingIt helps users understand image content in different languages, such as recognizing and interpreting road signs, advertisements, menus, etc. in a multilingual environment.
- Image and video analysisIn fields such as security monitoring and content moderation, it analyzes images and videos to identify and filter inappropriate content.
- Education and LearningProvides image and text analysis of multilingual learning materials for non-native language learners, enhancing their language learning experience.
- Tourism and NavigationIt helps tourists identify and translate street signs, maps, and cultural landmarks in different countries.
- e-commerceOn multilingual e-commerce platforms, this helps users understand product descriptions and images, enhancing the shopping experience.