AB
AiBoss
project

FG-CLIP 2 - 360 Open Source Bilingual Fine-Grained Visual Language Alignment Model

FG-CLIP 2 is an open-source, fine-grained bilingual visual-language alignment model launched by 360, designed specifically to solve the problem of accurate alignment between vision and language. It has achieved significant breakthroughs in the field of visual-language understanding, especially in Chinese-English bilingual tasks...

What is FG-CLIP 2?

FG-CLIP 2, an open-source bilingual fine-grained visual-language alignment model launched by 360, is designed to solve the problem of accurate alignment between vision and language. It has achieved significant breakthroughs in visual-language understanding, particularly excelling in Chinese-English bilingual tasks. The model employs a hierarchical alignment architecture, progressively improving its ability to understand image details through global semantic alignment and fine-grained visual-language learning. It introduces a dynamic attention mechanism that intelligently focuses on key regions of the image, better handling complex visual-language tasks. FG-CLIP 2 has surpassed existing top models, such as Google's SigLIP 2 and Meta's MetaCLIP 2, in multiple authoritative benchmark tests, becoming one of the world's strongest visual-language models.

Main functions of FG-CLIP 2

  • Fine-grained visual language understandingIt can accurately understand the details in an image, including the attributes of objects and spatial relationships, thus solving the shortcomings of traditional models in fine-grained recognition.
  • Bilingual supportThe model performs well on both Chinese and English tasks, achieving true native bilingual support.
  • Hierarchical alignment architectureIt adopts a hierarchical alignment architecture, which simultaneously grasps both macroscopic scenes and microscopic details, thereby improving the model's ability to understand image details.
  • Dynamic attention mechanismIt features a dynamic attention mechanism that can intelligently focus on key areas of an image, enabling it to better handle complex visual language tasks.
  • Optimize bilingual collaborative strategiesTo address the imbalance in understanding between Chinese and English, and improve the overall performance of the model in bilingual tasks.
  • Powerful performanceIn 29 authoritative and publicly available benchmark tests, it surpassed Google's SigLIP 2 and Meta's MetaCLIP 2, becoming the world's strongest visual language model.
  • High concurrency response speedIt adopts an explicit dual-tower structure, and image and text features can be pre-computed and cached to ensure millisecond-level response speed in high-concurrency scenarios.
  • Adaptive input sizeThe dynamic resolution mechanism allows the model to adaptively handle inputs of different sizes, improving the model's flexibility and adaptability.
  • Abundant open source resourcesIt provides code, model weights, and detailed training datasets, greatly facilitating researchers and developers.

Technical Principles of FG-CLIP 2

  • Hierarchical alignment architectureBy using global semantic alignment and fine-grained visual language learning, the model's ability to understand image details is gradually improved.
  • Dynamic attention mechanismIt intelligently focuses on key areas of an image, enabling better handling of complex visual language tasks.
  • Bilingual Collaborative StrategiesOptimize the balance between Chinese and English comprehension to improve the overall performance of bilingual tasks.
  • Multimodal data training: Training with large-scale Chinese and English image-text pairs enhances the model's bilingual generalization ability.
  • Fine-grained supervised learningIntroducing supervision signals such as region-text matching and long description modeling improves fine-grained visual language understanding capabilities.
  • Intra-text modal comparisonBy using intra-text modal contrastive loss, we can better distinguish semantically similar descriptions.
  • Training with difficult samplesIntroducing "hard negative samples" generated from large models further improves model performance.
  • Dynamic resolution mechanismIt can adaptively handle inputs of different sizes, improving the flexibility and adaptability of the model.

Project address for FG-CLIP 2

  • Project official website: https://360cvgroup.github.io/FG-CLIP/
  • Github repository: https://github.com/360CVGroup/FG-CLIP
  • arXiv technical paper: https://arxiv.org/pdf/2510.10921

Application scenarios of FG-CLIP 2

  • Home robotsIt can accurately understand and execute complex household commands, such as "pick up the phone with a cracked screen on the coffee table," enhancing the robot's practicality in the home environment.
  • Security monitoringIt can quickly locate and identify targets, such as "finding suspicious persons wearing black baseball caps," improving the efficiency and accuracy of security systems.
  • e-commerce sectorAccurately understand product descriptions, improve the accuracy of "text-based image search", reduce the cost of multilingual annotation and adaptation, and optimize user experience.
  • autonomous drivingAccurately identify objects and scenes in the road environment, such as "identifying whether there are obstacles in the lane ahead", to improve the safety of autonomous driving systems.
  • Medical imagingIt assists doctors in image diagnosis, such as "identifying abnormal areas in X-rays," improving the accuracy and efficiency of diagnosis.
  • EducationUsed in intelligent educational tools, such as "identifying objects in images and providing related knowledge," to enrich teaching content and formats.