AB
AiBoss
project

GLM-4.5V - The latest generation visual reasoning model from Zhipu Open Source.

GLM-4.5V is the latest generation visual inference model from Zhipu Open Source. Built with 106B parameters and possessing 12B activation capability, it is currently a leading visual language model (VLM). The model is based on GLM-4.1V-Thinking...

What is GLM-4.5V?

GLM-4.5V is the latest generation visual reasoning model launched by Zhipu. Built with 106B parameters and possessing 12B activation capability, it is currently a leading visual language model (VLM). The model is an upgrade from GLM-4.1V-Thinking, inheriting its excellent architecture and trained in conjunction with the next-generation text-based model GLM-4.5-Air. The model demonstrates outstanding performance in visual understanding and reasoning capabilities, suitable for scenarios such as web front-end replication, grounding, graph-finding games, and video understanding, and is expected to further promote the development of multimodal applications. To help developers intuitively experience the powerful capabilities of GLM-4.5V and create their own multimodal applications, the team has open-sourced a desktop assistant application that can take screenshots and record screens in real time, leveraging the GLM-4.5V model to handle various visual tasks such as code assistance, video analysis, game solving, and document interpretation.

Main functions of GLM-4.5V

  • Visual understanding and reasoningIt can understand and analyze visual content such as images and videos, and perform complex visual reasoning tasks, such as recognizing objects, scenes, and relationships between people.
  • Multimodal interactionIt supports the fusion of text and visual content, such as generating images from text descriptions or generating text descriptions from images.
  • Web front-end replicationGenerate front-end code based on web page design mockups to enable rapid web page development.
  • Image Search GameIt supports image-based search and matching tasks, such as finding specific targets in complex scenes.
  • Video UnderstandingIt supports analyzing video content, extracting key information, and performing tasks such as video summarization and event detection.
  • Cross-modal generationIt supports generating text from visual content or generating visual content from text, enabling seamless conversion of multimodal content.

Technical Principles of GLM-4.5V

  • Large-scale pre-trainingThe model is based on a 106-B pre-trained architecture and is trained with massive amounts of text and visual data to learn a joint representation of language and vision.
  • Visual language fusionIt adopts the Transformer architecture to fuse text and visual features and realizes the interaction between text and visual information based on the cross-attention mechanism.
  • Activation mechanismThe model is designed with 12B activation parameters, which are used to dynamically activate relevant parameter subsets during inference, thereby improving computational efficiency and inference performance.
  • Structural Inheritance and OptimizationInheriting the excellent structure of GLM-4.1V-Thinking, and trained in conjunction with the new generation text pedestal model GLM-4.5-Air, performance is further improved.
  • Multimodal task adaptationBased on fine-tuning and optimization, the model can adapt to a variety of multimodal tasks, such as visual question answering, image description generation, and video understanding.

GLM-4.5V performance

  • General VQAThe GLM-4.5V performs best in general visual question answering tasks, especially scoring as high as 88.2 in the MMBench v1.1 benchmark test.
  • STEMThe GLM-4.5V also leads in science, technology, engineering, and mathematics (STEM) tasks, achieving a high score of 84.6 on the MathVista test.
  • Long Document OCR & ChartIn the OCR Bench test, which handles long documents and charts, the GLM-4.5V demonstrated excellent performance with a score of 86.5.
  • Visual GroundingThe GLM-4.5V performed exceptionally well in visual localization tasks, scoring 91.3 in the RefCOCO+loc (val) test.
  • Spatial ReasoningIn terms of spatial reasoning ability, the GLM-4.5V achieved an excellent score of 87.3 in the CV-Bench test.
  • CodingIn programming tasks, the GLM-4.5V scored 82.2 on the Design2Code benchmark, demonstrating its capabilities in code generation and comprehension.
  • Video UnderstandingThe GLM-4.5V also performs well in video understanding, scoring 74.6 in the VideoMME (w/o sub) test.

Project address for GLM-4.5V

  • GitHub repositoryhttps://github.com/zai-org/GLM-V/
  • HuggingFace model library: https://huggingface.co/collections/zai-org/glm-45v-68999032ddf8ecf7dcdbc102
  • Technical Papers: https://github.com/zai-org/GLM-V/tree/main/resources/GLM-4.5V_technical_report.pdf
  • Desktop assistant applicationhttps://huggingface.co/spaces/zai-org/GLM-4.5V-Demo-App

How to use GLM-4.5V

  • Registration and LoginVisit the Z.ai official website and register an account using your email address. After registration, log in to your account.
  • Select ModelAfter logging in, select GLM-4.5V from the model selection drop-down menu.
  • Experience features:
    • Web front-end replicationUpload your web design template, and the model will automatically generate front-end code.
    • Visual reasoningUpload images or videos, and the model will perform tasks such as visual understanding, object recognition, and scene analysis.
    • Image Search GameUpload the target image, and the model will find a matching image in the complex scene.
    • Video UnderstandingUpload a video file, and the model will extract key information and generate a video summary or event detection results.

API call pricing for GLM-4.5V

  • enter2 yuan/M tokens
  • Output6 yuan/M tokens
  • Response speedReaching 60-80 tokens/s

Application scenarios of GLM-4.5V

  • Web front-end replicationUpload your web design and quickly generate front-end code to help developers efficiently develop web pages.
  • Visual Q&AUsers upload images and ask questions, and the model generates accurate answers based on the image content. This can be used in fields such as education and intelligent customer service.
  • Image Search GameIt can quickly locate target images in complex scenes, and is suitable for security monitoring, smart retail and entertainment game development.
  • Video UnderstandingAnalyze video content, extract key information to generate summaries or detect events, and optimize video recommendation, editing, and monitoring.
  • Image description generationGenerate accurate descriptive text for uploaded images to help visually impaired people understand the images and improve their social media sharing experience.