Pangea - Carnegie Mellon University's open-source multilingual and multimodal large language model
Pangea is a multilingual, multimodal, large-scale language model (LLM) developed by a team at Carnegie Mellon University. It aims to improve the coverage of global language and cultural diversity. The model includes a diverse dataset of 6 million instructions and supports 39 languages...
What is Pangea?
Pangea, a multilingual, multimodal large-scale language model (LLM) developed by a team at Carnegie Mellon University, aims to enhance global coverage of linguistic and cultural diversity. The model includes a diverse dataset of 6 million instructions, supports 39 languages, and features high-quality English instructions, machine translation instructions, and culture-related tasks. Pangea's performance was evaluated using the PangeaABench evaluation suite, which includes 14 datasets covering 47 languages. Pangea outperforms existing open-source models (such as Llava-1.5-7B and Llava-Next-7B) in multilingual and cultural contexts. Research found that the proportion of English data, language popularity, and the number of multimodal training samples significantly impact performance.
Pangea's main functions
- Multilingual supportIt can understand and generate text in 39 different languages, which is very useful in multilingual communication and processing.
- Multimodal understandingIn addition to text, it can process and understand images, and performs well in tasks such as image description and visual question answering.
- Cross-cultural coverageIncluding culture-related multimodal tasks in training helps the model better understand and adapt to different cultural backgrounds.
- High-quality instructions followPangea uses high-quality English instructions and carefully machine-translated instructions during training to ensure the model's accuracy and consistency across different languages.
Pangea's technical principles
- Dataset ConstructionBased on the Pangea dataset, a multilingual dataset containing 6 million instructions covering 39 languages.
- Machine translationTo address the scarcity of multilingual data, machine translation technology is used to translate high-quality English instructions into other languages.
- Culture-related tasksIncorporate culture-related multimodal tasks into the training to improve the model's understanding and adaptability to cultural differences.
- Evaluation kitPangeaABench is an evaluation suite containing 14 datasets covering 47 languages, used to comprehensively evaluate model performance on multilingual and multimodal tasks.
- Model ArchitectureBased on the LLaVA-Next architecture, Qwen2-7B-Instruct is used as the backbone of the language model, providing the model with powerful language understanding and generation capabilities.
Pangea's project address
- Project official website:neulab.github.io/Pangea
- GitHub repository:https://github.com/neulab/Pangea
- HuggingFace model library:https://huggingface.co/collections/neulab/pangea-6713c3b0d78a453906eb2ed8
- arXiv technical paper:https://arxiv.org/pdf/2410.16153
- Experience the demo online:https://huggingface.co/spaces/neulab/Pangea
Pangea's application scenarios
- Multilingual customer serviceIn a global company, we provide multilingual customer support and services to help solve problems for customers who speak different languages.
- Education and LearningAs an educational tool, it helps learners access multilingual learning materials or provides assistance in language teaching.
- Cross-cultural communicationTo promote communication and understanding among people from different cultural backgrounds within international or non-governmental organizations.
- Social media and content creationPangea helps content creators generate multilingual content or interact with users of different languages on social media.
- Tourism and NavigationIn the tourism industry, providing multilingual travel information and navigation services helps tourists overcome language barriers.