AB
AiBoss
project

LEOPARD - A visual language model developed by Tencent AI Lab in Seattle

LEOPARD is a visual language model developed by Tencent AI Lab in Seattle, specifically designed for understanding and processing multi-image tasks containing large amounts of text. LEOPARD is based on two main technological innovations: firstly, it incorporates approximately one million specially designed...

What is LEOPARD?

LEOPARD, developed by Tencent AI Lab in Seattle, is a visual language model designed for understanding and processing multi-image tasks containing large amounts of text. LEOPARD is based on two main technological innovations: first, a high-quality multimodal instruction dataset of approximately one million lines specifically optimized for text-rich, multi-image scenarios; and second, the development of an adaptive high-resolution multi-image encoding module that dynamically optimizes visual sequence length allocation. LEOPARD demonstrates superior performance in multiple benchmark tests, excelling in complex tasks requiring the understanding of single image content and reasoning across multiple visual inputs.

LEOPARD's main functions

  • Processing text-rich multi-image tasksIt is used to understand and process multi-image scenarios containing a large amount of text information, such as slides, scanned documents, and web page screenshots.
  • Cross-image reasoningThe model can understand the content of a single image and perform logical reasoning and establish relationships between multiple images.
  • High-resolution image processingBased on an adaptive high-resolution multi-image encoding module, it can effectively process high-resolution images while maintaining the clarity of text and details.
  • Dynamic visual sequence length optimizationThe visual sequence length is dynamically optimized based on the original aspect ratio and resolution of the input image, balancing image detail and model processing capabilities.
  • Multimodal instruction tuning: Optimize datasets using large-scale multimodal instructions, enabling optimization for complex visual language tasks.

LEOPARD's technical principles

  • Multimodal Large Language Model (MLLM)Based on the MLLM architecture, it integrates a visual encoder, a visual language connector, and a language model to process visual and textual information.
  • Dataset ConstructionWe constructed the LEOPARD-INSTRUCT dataset, which contains approximately one million instructions for text-rich, multi-image scenes, and used them for model training and optimization.
  • Adaptive high-resolution codingBased on an adaptive strategy, the visual feature sequence is dynamically adjusted according to the characteristics of the input image to adapt to the sequence length limit of the model.
  • Pixel shuffling technologyThe pixel shuffling operation is used to losslessly compress long visual feature sequences into shorter sequences, making it easier for the model to process more high-resolution images.
  • Image segmentationThe high-resolution image is segmented into multiple sub-images for independent processing and detail preservation. The visual features are then fed into the language model along with the textual information.

LEOPARD project address

Application scenarios of LEOPARD

  • Automated document understandingIt can process multi-page documents, such as contracts, reports, and academic papers, and automatically extract key information and data.
  • Education and academic researchEducational aids, such as e-learning materials and academic presentations, provide an interactive learning experience.
  • Business intelligence and data analyticsAnalyze business charts and tables to provide market trend forecasting and decision support.
  • Web page content analysisIt understands and extracts web page content for use in search engine optimization (SEO) and content recommendation systems.
  • Customer service and supportBased on the analysis of user-uploaded images and text, we provide more accurate customer service and technical support.