News
360 Releases: FG-CLIP2 Tops the List of World's Strongest Cross-Modal Image-Text Model
360's FG-CLIP2 model has achieved a major breakthrough in the field of cross-modal image-text communication. The model outperformed Google and Meta in all eight task categories and 29 tests, becoming the most powerful cross-modal image-text VLM model to date. FG-CLIP2 achieves pixel-level image understanding, accurately identifying details such as hair, spots, and colors, and possesses powerful fine-grained understanding capabilities in both Chinese and English.