AB
AiBoss
project

Pangea - Carnegie Mellon University's open-source multilingual and multimodal large language model

Pangea is a multilingual, multimodal, large-scale language model (LLM) developed by a team at Carnegie Mellon University. It aims to improve the coverage of global language and cultural diversity. The model includes a diverse dataset of 6 million instructions and supports 39 languages...

What is Pangea?

Pangea, a multilingual, multimodal large-scale language model (LLM) developed by a team at Carnegie Mellon University, aims to enhance global coverage of linguistic and cultural diversity. The model includes a diverse dataset of 6 million instructions, supports 39 languages, and features high-quality English instructions, machine translation instructions, and culture-related tasks. Pangea's performance was evaluated using the PangeaABench evaluation suite, which includes 14 datasets covering 47 languages. Pangea outperforms existing open-source models (such as Llava-1.5-7B and Llava-Next-7B) in multilingual and cultural contexts. Research found that the proportion of English data, language popularity, and the number of multimodal training samples significantly impact performance.

Pangea's main functions

  • Multilingual supportIt can understand and generate text in 39 different languages, which is very useful in multilingual communication and processing.
  • Multimodal understandingIn addition to text, it can process and understand images, and performs well in tasks such as image description and visual question answering.
  • Cross-cultural coverageIncluding culture-related multimodal tasks in training helps the model better understand and adapt to different cultural backgrounds.
  • High-quality instructions followPangea uses high-quality English instructions and carefully machine-translated instructions during training to ensure the model's accuracy and consistency across different languages.

Pangea's technical principles

  • Dataset ConstructionBased on the Pangea dataset, a multilingual dataset containing 6 million instructions covering 39 languages.
  • Machine translationTo address the scarcity of multilingual data, machine translation technology is used to translate high-quality English instructions into other languages.
  • Culture-related tasksIncorporate culture-related multimodal tasks into the training to improve the model's understanding and adaptability to cultural differences.
  • Evaluation kitPangeaABench is an evaluation suite containing 14 datasets covering 47 languages, used to comprehensively evaluate model performance on multilingual and multimodal tasks.
  • Model ArchitectureBased on the LLaVA-Next architecture, Qwen2-7B-Instruct is used as the backbone of the language model, providing the model with powerful language understanding and generation capabilities.

Pangea's project address

Pangea's application scenarios

  • Multilingual customer serviceIn a global company, we provide multilingual customer support and services to help solve problems for customers who speak different languages.
  • Education and LearningAs an educational tool, it helps learners access multilingual learning materials or provides assistance in language teaching.
  • Cross-cultural communicationTo promote communication and understanding among people from different cultural backgrounds within international or non-governmental organizations.
  • Social media and content creationPangea helps content creators generate multilingual content or interact with users of different languages on social media.
  • Tourism and NavigationIn the tourism industry, providing multilingual travel information and navigation services helps tourists overcome language barriers.