FG-CLIP 2 - 360 Open Source Bilingual Fine-Grained Visual Language Alignment Model
FG-CLIP 2 is an open-source, fine-grained bilingual visual-language alignment model launched by 360, designed specifically to solve the problem of accurate alignment between vision and language. It has achieved significant breakthroughs in the field of visual-language understanding, especially in Chinese-English bilingual tasks...
What is FG-CLIP 2?
FG-CLIP 2, an open-source bilingual fine-grained visual-language alignment model launched by 360, is designed to solve the problem of accurate alignment between vision and language. It has achieved significant breakthroughs in visual-language understanding, particularly excelling in Chinese-English bilingual tasks. The model employs a hierarchical alignment architecture, progressively improving its ability to understand image details through global semantic alignment and fine-grained visual-language learning. It introduces a dynamic attention mechanism that intelligently focuses on key regions of the image, better handling complex visual-language tasks. FG-CLIP 2 has surpassed existing top models, such as Google's SigLIP 2 and Meta's MetaCLIP 2, in multiple authoritative benchmark tests, becoming one of the world's strongest visual-language models.
Main functions of FG-CLIP 2
-
Fine-grained visual language understandingIt can accurately understand the details in an image, including the attributes of objects and spatial relationships, thus solving the shortcomings of traditional models in fine-grained recognition.
-
Bilingual supportThe model performs well on both Chinese and English tasks, achieving true native bilingual support.
-
Hierarchical alignment architectureIt adopts a hierarchical alignment architecture, which simultaneously grasps both macroscopic scenes and microscopic details, thereby improving the model's ability to understand image details.
-
Dynamic attention mechanismIt features a dynamic attention mechanism that can intelligently focus on key areas of an image, enabling it to better handle complex visual language tasks.
-
Optimize bilingual collaborative strategiesTo address the imbalance in understanding between Chinese and English, and improve the overall performance of the model in bilingual tasks.
-
Powerful performanceIn 29 authoritative and publicly available benchmark tests, it surpassed Google's SigLIP 2 and Meta's MetaCLIP 2, becoming the world's strongest visual language model.
-
High concurrency response speedIt adopts an explicit dual-tower structure, and image and text features can be pre-computed and cached to ensure millisecond-level response speed in high-concurrency scenarios.
-
Adaptive input sizeThe dynamic resolution mechanism allows the model to adaptively handle inputs of different sizes, improving the model's flexibility and adaptability.
-
Abundant open source resourcesIt provides code, model weights, and detailed training datasets, greatly facilitating researchers and developers.
Technical Principles of FG-CLIP 2
-
Hierarchical alignment architectureBy using global semantic alignment and fine-grained visual language learning, the model's ability to understand image details is gradually improved.
-
Dynamic attention mechanismIt intelligently focuses on key areas of an image, enabling better handling of complex visual language tasks.
-
Bilingual Collaborative StrategiesOptimize the balance between Chinese and English comprehension to improve the overall performance of bilingual tasks.
-
Multimodal data training: Training with large-scale Chinese and English image-text pairs enhances the model's bilingual generalization ability.
-
Fine-grained supervised learningIntroducing supervision signals such as region-text matching and long description modeling improves fine-grained visual language understanding capabilities.
-
Intra-text modal comparisonBy using intra-text modal contrastive loss, we can better distinguish semantically similar descriptions.
-
Training with difficult samplesIntroducing "hard negative samples" generated from large models further improves model performance.
-
Dynamic resolution mechanismIt can adaptively handle inputs of different sizes, improving the flexibility and adaptability of the model.
Project address for FG-CLIP 2
- Project official website: https://360cvgroup.github.io/FG-CLIP/
- Github repository: https://github.com/360CVGroup/FG-CLIP
- arXiv technical paper: https://arxiv.org/pdf/2510.10921
Application scenarios of FG-CLIP 2
-
Home robotsIt can accurately understand and execute complex household commands, such as "pick up the phone with a cracked screen on the coffee table," enhancing the robot's practicality in the home environment.
-
Security monitoringIt can quickly locate and identify targets, such as "finding suspicious persons wearing black baseball caps," improving the efficiency and accuracy of security systems.
-
e-commerce sectorAccurately understand product descriptions, improve the accuracy of "text-based image search", reduce the cost of multilingual annotation and adaptation, and optimize user experience.
-
autonomous drivingAccurately identify objects and scenes in the road environment, such as "identifying whether there are obstacles in the lane ahead", to improve the safety of autonomous driving systems.
-
Medical imagingIt assists doctors in image diagnosis, such as "identifying abnormal areas in X-rays," improving the accuracy and efficiency of diagnosis.
-
EducationUsed in intelligent educational tools, such as "identifying objects in images and providing related knowledge," to enrich teaching content and formats.