AB
AiBoss
Tutorials

Real-world testing of the Zhipu GLM-4.6V reveals it to be the most powerful domestically produced multimodal agent base model.

Zhipu officially launched its new visual reasoning model series, GLM-4.6V, and made it fully open source! This release includes two versions: GLM-4.6V: total parameters 106B, activation parameters per inference approximately 12B, visual understanding accuracy reaching...

Zhipu officially launches its new visual reasoning model GLM-4.6V series modelsand comprehensivelyopen source!

This release includes two versions:

  • GLM-4.6VThe total number of parameters is 106 bytes, with approximately 12 bytes of activation parameters per inference, achieving visual understanding accuracy comparable to other parameters.SOTASuitable for cloud and high-performance scenarios;
  • GLM-4.6V-FlashTotal parameters: 9B, lighter and faster, suitable for local deployment;

GLM-4.6V is the first to integrate Function Call (tool call) capabilities into a visual model.,letLarge ModelIt has both eyes and hands, supports native processing of complex visual tasks, and can proactively invoke tools to complete subsequent operations based on visual understanding.

For example, GLM-4.6V can directly "understand" papers with complex structures and a large number of charts and diagrams, and reorganize them into a well-illustrated article that everyone can understand.

With just a screenshot, the page structure can be dissected and a nearly identical front-end page can be replicated.

Open z.ai and select the model GLM-4.6V in the upper left corner of the page.

Official website:https://chat.z.ai

GitHub:https://github.com/zai-org/GLM-V

Hugging Face:https://huggingface.co/collections/zai-org/glm-46v

The GLM-4.6V can access four tools: image recognition, image processing, image search, and shopping search.

Below the input box, the official documentation provides a set of typical function examples, including universal search, image and text scanning, intelligent document reading, and video comprehension.intelligentPrice comparison and mathematical problem-solving, etc.Selecting any function will cause the GLM-4.6V to...automaticCall the matching tool.

Case 1 Universal Search

Prompt wordsWhere is this place, and what month is best to visit?

The GLM-4.6V has native visual understanding capabilities, directly calling image recognition tools to identify the content in images and then searching for relevant knowledge to provide a response.

Case 2: Image and Text Scan

Prompt wordsExtract information from the image and convert it into an Excel spreadsheet.

The GLM-4.6V has a very accurate understanding of content and layout.

Let's try something a little more complex:

Prompt wordsPlease scan out the ingredients, ingredient list, and other information for this cat food, and analyze whether it is suitable for a 2-year-old kitten to eat long-term.

The GLM-4.6V also accurately identifies the raw material composition and product components, and performs analysis based on this information.

Case 3: Intelligent Document Reading

Last week, Professor Pan Jianwei's team from the University of Science and Technology of China published their work in a top international journal.PRLPublished in Physical Review Lettersup to dateThe research findings represent a breakthrough in the field of quantum physics, ending the century-long debate between Einstein and Bohr.

I found the original paper and asked GLM-4.6V to analyze it for us.

Prompt wordsIn layman's terms, explain what this paper is about, why it is said to have ended the century-long debate between Einstein and Bohr, and what this achievement means for the real world and ordinary people, beyond its academic value.

The GLM-4.6V can not only understand complex charts and graphs, but also reorganize key information and explain it clearly in a visually appealing way.

Case 4 Video Comprehension

Prompt wordsThis is a classic scene from The Secret Life. What specific camera techniques were used, and what are the highlights of its shot composition?

The interpretation provided by GLM-4.6V was very professional. It explained the content of the entire video, which shots were used, and what emotions these shots conveyed... It was much more profound than my understanding.

Case 5: Mathematical Problem Solving

Prompt wordsAnswer the questions in the diagram.

The GLM-4.6V can combine visual information with external knowledge for combined reasoning, and the problem-solving approach is very clear.

case6 intelligentparity

Prompt wordsPlease help me search for affordable earrings similar to those worn by Zhao Lusi in the picture.

The GLM-4.6V app found several affordable alternatives for me, and the identification was quite accurate, with options available on different platforms.

Case 7: Creation of graphic and textual content

Prompt wordsSearch for the development process of visual models and generate a report with illustrations.

Case 8: Replicating Front-End Web Pages

Prompt wordsTo replicate the webpage in the screenshot, all image materials used on the page must be real images and videos; do not use placeholders or placeholder elements.

Visual understanding, structural reasoning, and code generation are all done in one step, producing web pages that are almost identical to the original images. Even the floating window structure in the screenshots is recognized and restored! The navigation bar also provides space for navigation.

In actual testing, the GLM-4.6V can not only recognize details in the picture, but also connect the meaning of the image with the meaning of natural language, understand what the picture is expressing, and the relationship between these information. The whole process is quite smooth.

When using this tool, it's recommended to keep Deep Thinking enabled for higher model response quality. It's advisable to disable the tool when replicating the front-end; otherwise, customize the settings according to the task or keep the default settings in the official options.

thispowerfulThe visual capabilities will also be integrated into Zhipu's Coding Plan, which costs as little as 20 yuan per month and can be used directly.up to dateIts modeling capabilities are excellent for everyday use.

As these capabilities mature, visual information will become deeply involved in decision-making, planning, and action itself, and images from the real world will become primary sources of information that the system can directly understand and access.

The improvement in visual model capabilities is not just for... AI A pair of eyes and a pair of hands, but for the next generation.intelligentbodyEngage with the real world to open doors.

In the future, robots will no longer need to be precisely programmed to perform a certain action, but will be able to understand natural commands such as "go get the red sweater on the far right of the wardrobe".