Domestic open-source alternative to Nano Banana, with 10 case studies testing GLM-Image.
Zhipu has released and open-sourced its latest multimodal model, GLM-Image, achieving a text accuracy of 0.9116, making it the most accurate open-source model for text rendering to date. Its NED metric is also top-notch, demonstrating high adherence to prompts and minimizing misspellings...
Nano Banana Pro has become a viral sensation online, with various infographics, posters, and knowledge cards flooding social media. However, when people actually use it in their daily work, they often encounter this problem: when there's a lot of text, the generated images often contain crooked and messy Chinese characters. The layout might be fine in some parts, but the text information is so incorrect that it's unusable.
This long-standing pain point that has plagued creators and ordinary workers has finally been solved by domestically produced models.
Zhipu Release andopen sourceAlreadyup to dateofMultimodalThe GLM-Image model achieves a text accuracy of 0.9116, making it the most accurate text rendering model currently available.open sourceModel.NEDThe metrics are also top-notch, forPrompt wordsThe compliance rate is very high, and there are fewer typos and omissions.
It really caught my eye. Without further ado, let's test it out together.
Currently, GLM-Image offers three ways to experience it: either by using it online in BigModel or by calling the API. All three costs 0.1 yuan per image and natively supports images of any size from 1024*1024 to 2048*2048.
The online generator does not support resizing or scaling; it will always output a fixed 1280*1280 image.
Official website:
https://bigmodel.cn/trialcenter/modeltrial/image
Zhipu Qingyan APP or web versionAIDrawingintelligentbody, alsofreeExperience GLM-Image.
Official website:
https://chatglm.cn/main/gdetail/65a232c082ff90a2ad2f15e2
I preferrecommendEveryone Claude Using the API via code will yield better results and also supports custom sizes.
Claude For code configuration steps, please refer to this tutorial:Step-by-step guide to connecting a GLM-4.5 power supply. Claude Code:open sourceThe Ultimate Model Configuration Guide
Request example (note the need to replace the API KEY).Prompt words(and image size):
curl -X POST "https://open.bigmodel.cn/api/paas/v4/images/generations" \ -H "Authorization: Bearer YOU-API-KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "glm-image", "prompt": "图像提示词", "size": "1056 × 1408" }'
Case 1: Science Popularization Illustration
Prompt wordsA watercolor-style hand-painted model structure illustration for educational purposes, with a left-right layout. The beige, antique watercolor paper textured background features natural blurring and coffee stain effects, creating a soft and translucent hand-painted feel.
Left side – Diffusion architecture (overall gradually becoming clearer): The title "Diffusion architecture (overall gradually becoming clearer)" is located at the top.
Showing three vertically arranged, progressively layered watercolor paintings:
Abstract, colorful watercolor noise clumps with disordered color mixing, labeled "Random Noise Input".
The watercolor blocks in the image began to converge towards several main areas, significantly reducing overall noise and eliminating the scattered specks across the screen. A general compositional direction vaguely emerged, such as the distinction between the ground and the sky, and low-saturation, blurry tree shadows and house blocks in the distance. However, all elements existed only as blurry color gamuts, lacking clear edges, sharp outlines, and identifiable details.
Marked "Outline begins to appear". A clear and complete watercolor pastoral landscape, including trees, houses and sky, with rich details, marked "Image clearly formed".
Right side – Autoregressive Architecture (block-by-block stitching): The title "Autoregressive Architecture (block-by-block stitching)" is located at the top.
Demonstrates the assembly process of three vertically arranged modular units:
A small watercolor flower square appears in the upper left corner, with a complete and clear shape, labeled "Partially formed first".
While keeping the position, size, and shape of the flower unchanged from the first step, add watercolor blocks of leaves and stems next to the flower. The new content should blend naturally with the original flower. Do not move, scale, or redraw existing flowers. Mark "Continue generating on existing results".
With all elements remaining in their original positions from the first two steps, continue adding new flowers, green leaves, and butterflies to eventually create a complete bouquet image, labeled "gradually pieced together to form the complete result".
Overall style requirements: Watercolor hand-painted textbook illustration style, with soft and natural colors. All text labels should use handwritten Chinese fonts. Watercolor edges should be naturally blurred, and lines should not be completely closed to avoid an industrial or digital feel, maintaining a warm and approachable science popularization atmosphere.
Case 2: Xiaohongshu Cover
Prompt wordsGenerate a Xiaohongshu (Little Red Book) note cover featuring a commuting OOTD (Outfit of the Day) theme, with the content: "OOTD, a week's worth of outfits without repeating."
Prompt wordsGenerate a Xiaohongshu (Little Red Book) note cover with the theme of a rental apartment renovation vlog, in 3:4 aspect ratio, with the content: "500 yuan to transform a rental apartment 🏠 from an old and dilapidated place to a cute and whimsical home, even the landlord asked me if I had moved!"
Case 3 comics
Prompt wordsGenerate a humorous folk-themed comic illustration, consisting of 4 panels, featuring a human girl and an anthropomorphic weasel. The overall style is lighthearted and funny, with slight elements of Chinese folklore. The art style leans towards cartoonish, but the characters' expressions are exaggerated and clear.
Frame 1
A nighttime scene unfolds on a narrow path. A human girl is walking with a bag on her back when a weasel, walking upright, suddenly blocks the way, its expression serious and mysterious.
(Dialogue bubble) The weasel said, "Do I look like a human to you?"
2nd square
The girl stopped, tilted her head in thought, and looked casual.
(Dialogue bubble) The girl says, "You look like the Jade Emperor to me."
3rd square
The weasel's pupils dilated instantly, its expression one of extreme terror, and cold sweat poured down its back. There was no dialogue; the emotions were exaggerated.
4th square
The weasel suddenly rushed forward and covered the girl's mouth, growling nervously in a low voice.
(Dialogue bubble) The weasel says, "Sis, don't get a new account!"
Overall style requirements:
Combining Chinese folk customs with modern jokes
Exaggerated facial expressions and obvious body movements
The dialogue is entirely in Chinese.
The visuals are clean, and the comic panel layout is clear.
The atmosphere leans towards humor, plot twists, and internet memes.
No horror elements, leaning towards light comedy
Case 4: Images from Xiaohongshu
Prompt wordsInfographic in the style of Xiaohongshu (Little Red Book), cartoon style, hand-drawn style text, off-white background.
The image features a cartoon-style combination of a brain and a paintbrush, symbolizing cognition and generation, with the words "GLM-Image" hand-drawn on a sticky note next to it.
A hand-drawn headline at the top: Domestic Chips Outperform...open sourceImage SOTA
The subtitle below: Zhipu × Huawei | GLM-Image
The addition of cartoon chips, servers, and glitter stickers gives the overall visual a futuristic yet cute feel.
Overall style: hand-drawn, fresh, and futuristic cartoonish, with ample white space and a clear focus. All images and text are hand-drawn, with no realistic elements.
Watermark in the bottom right corner: "K-Sister Research Society"
Prompt wordsThe infographic is in the style of Xiaohongshu (Little Red Book), featuring a cartoon style, hand-drawn text, and a light mint green background.
Left-right structure diagram of the screen:
On the left is a cartoon "speech bubble + brain," labeled "Autoregressive Model | Understanding Instructions | Overall Composition."
On the right is a cartoon "drawing brush + text strokes" image, labeled "Diffusion Decoder | Detail Depiction | Text Strokes".
A hand-drawn arrow connects the two parts, and the title is written above: Understand the instructions and write the correct words.
Highlighting with fluorescent pen underlines emphasizes the keywords "cognitive generation" and "knowledge + reasoning".
Overall style: hand-drawn, cute, clear and easy to understand, with concise information and plenty of white space.
Watermark in the bottom right corner: "K-Sister Research Society"
Case 5 Information Labeling
Prompt wordsGenerate a bedroom decor showcase image in the style of Xiaohongshu (a Chinese social media platform), with realistic shooting quality and a vertical composition. The image depicts a small bedroom with a warm and soothing overall feel, leaning towards a Japanese and cute style. On one side is a single bed with green sheets and a white bedspread, and teddy bears and cushions on the bed. Next to the bed is a cartoon-style rug with a cute cow pattern and soft colors.
In the center of the room stands a white desk with an open laptop displaying brightly colored illustrations. Stationery, a water glass, and small ornaments are also on the desk. Behind the desk is a window with white sheer curtains hanging down, offering glimpses of city lights at night. Warm yellow desk lamps and nightlights are placed by the window and on the desk, creating a soft evening atmosphere. To the left are multi-tiered storage shelves and bookshelves, filled with books, storage boxes, toys, nightlights, and everyday items. The items are plentiful but the overall space is tidy, giving it a slightly authentic, lived-in feel.
The image overlays multiple product tags, similar in style to those used in product recommendation images on Xiaohongshu (a Chinese social media platform). The tag text is yellow or light yellow with a slight outline and shadow, ensuring clear readability. Each tag must be anchored near its corresponding item, arranged irregularly, and includes the product name and price information, such as "Decorative painting 💰25", "Cow rug 💰45", "Night light 💰29.9", "Curtains 💰52", "Teddy bear 💰107", etc., creating a complete list of room furnishing expenses without obscuring the main visual element.
Case 6 Recipe
Prompt wordsGenerate an infographic for "Stir-fried Green Peppers with Pork" that includes step-by-step recipe information. Requirements:
Top-down view, minimalist style, white background
The Chinese name of the dish is displayed at the top center.
Label all ingredients with their Chinese names, quantities, and calorie content.
Use dashed lines and icons to illustrate cooking steps.
The finished product's presentation is displayed at the bottom.
According to the traditional method of making this dish,automaticSuitable match:
1. Ingredient list (including precise quantities and calorie content)
2. Cooking step icons (e.g., chopping vegetables, stir-frying, seasoning, etc.)
3. Final presentation style
Case 7 Infographic
Prompt wordsGenerate a high-end close-up poster of a Chinese cabbage in a product photography style. The image uses a centered composition, with a fresh and plump Chinese cabbage as the main subject. The leaves are layered, with light green outer leaves and pale yellow inner core. The veins are clearly visible, and the surface has a natural, refreshing, and moist texture. Overall, the cabbage appears clean, tender, and substantial.
The background is a clean, light gray color, free of any clutter or texture, highlighting the shape and texture of the ingredients themselves.
The lighting uses soft, natural light, shining from the side and front, so that the leaves are evenly lit and have distinct layers; the bottom and back form slight, soft shadows, which enhance the three-dimensionality without excessive contrast, resulting in a clean and sophisticated image.
Top text layout
Centered text at the top of the screen:
The main title, "Chinese Cabbage," is written in a natural and casual handwritten style. The font color is taken from the fresh light green of the outer leaves of a Chinese cabbage, with soft strokes that are friendly and natural.
Subtitle / Copy: "Clears heat and relieves irritability, moistens dryness and promotes bowel movement, strengthens the spleen and nourishes the stomach" copy is generated based on the concept of traditional Chinese medicine diet therapy, highlighting the moisturizing and conditioning properties of Chinese cabbage.
Subtitle font: handwritten, with a warm brown color, positioned below the main title, creating a soft and restrained visual effect.
Bottom content area (cooking method demonstration)
At the bottom of the screen, three cooking methods suitable for Chinese cabbage are arranged horizontally. Each method includes a small picture of the finished product and a Chinese text description, with a consistent and concise style.
The accompanying picture shows stir-fried Chinese cabbage after being quickly stir-fried over high heat. The stems are crisp and tender, and the leaves are glossy, presenting the most basic home-style flavor.
The picture shows cabbage and tofu stewed together. The soup is light in color, reflecting its mild and nourishing nature, and is suitable for consumption in all seasons.
The dish of stir-fried cabbage with vinegar features sliced cabbage that has been quickly stir-fried until slightly charred, with a refreshing color and a prominent sour and savory flavor that whets the appetite.
Overall style requirements:
The overall style is minimalist product advertising, with ample white space and restrained information, emphasizing the natural form and health attributes of ingredients. The aesthetic is high-end, clean, and understated, suitable for healthy eating, fresh food brands, or wellness-related visual content. It features realistic photographic quality, without any illustrative or cartoonish feel, and high resolution, making it suitable for posters, e-commerce main images, or health education displays.
Case 8 Poster
Prompt wordsThis modern, creative poster features a close-up of a donkey with its mouth outlined in pink to form a smiling face against a clear blue sky. Above the image is the pink artistic font that reads "Hello, working people!" To the left are white text depicting various aspects of working life (e.g., "Work pressure is heavier than a mountain, wallet lighter than paper," "A flurry of activity, but only a meager salary," "Going to work is like going to a funeral, leaving work is like going to a disco," "Coffee sustains me every day, but my salary is gone every year," "Working overtime until I'm bald, and my savings are down to double digits"). The color scheme is predominantly blue and brown, with bright and contrasting colors. The close-up composition emphasizes the donkey's face and the creative smile, creating a humorous and lighthearted atmosphere that satirizes the daily lives of working people.
Case 9 Movie Poster
Prompt wordsThis is a retro-style movie poster titled "Sunset at the Seaside." The image features a backdrop of an orange-red sunset over the sea, with the silhouettes of a man and woman sitting on a beach bench, creating a tranquil atmosphere. The poster uses a contrasting color scheme of fluorescent green and orange-red, and the handwritten "Sunset at the Seaside" lettering is visually striking. Coupled with the tagline, "Tired of the city's 4 PM sunsets? Sometimes I want to see a sunset by the sea," it conveys a longing for nature and romance. Retro elements at the bottom (the year "2025," the brand logo "cc ORIGINAL POSTER," etc.) enhance the nostalgic feel, resulting in an overall style that is both artistic and full of retro-chic vibes.
Case 10 Packaging Design
Prompt wordsThis is a 3D realistic tomato packaging box. The main design is a creative tomato shape, made of red paper material with a surface simulating the yellow spots and textures of a real tomato skin. A green simulated stem is placed on top. The box features a perforated window that displays several plump, round cherry tomatoes inside, arranged in a layered manner. A white label on the front clearly displays the word "Tomato." The box also features the "K-Sister Research Society" brand logo and product information such as "Natural Good Fruit," "Sweet and Refreshing," "Exploding Juice," and "Naturally Ripened." A recycling symbol is clearly marked on the packaging.practicalThe design embodies the concepts of sexiness and environmental protection. Against a pure white background, the packaging uses a level viewing angle to highlight its three-dimensional shape and rich material details, such as the glossy texture of the tomato, the paper's grain, and the exquisite printed logo. The design style is simple and fresh, with red, white, and green as the main colors, conveying a natural and healthy brand tone and successfully creating a creative and eco-friendly food packaging design. The image is high-definition and rich in detail.
This time, Zhipu did not continue to use the mainstream Diffusion architecture, but instead adopted a self-developed autoregressive + diffusion encoder hybrid architecture.
The Diffusion architecture is essentially a process of gradually clarifying highly chaotic noise, much like the process of opening our eyes from squinting to seeing the image clearly.
The models generated by the Diffusion architecture have a strong sense of unity and a consistent style, making them very suitable for posters and illustrations. Nano Banana Pro,MidjourneySeedream 4.5 and others are typical examples of Diffusion architecture models.
During the Diffusion generation process, text is also treated as a complex shape to be restored. With high text density, characters are easily stretched and distorted during the diffusion process, which can easily result in the "scratched characters" we often see.
Autoregressive (AR) generation is a step-by-step, sequential process where each step references previously generated content as context. The model first generates a character, then uses that character to determine the next character, ensuring strong correlation between the preceding and following content.
In the hybrid architecture of GLM-Image, the autoregressive mechanism intervenes first, responsible for...Prompt wordsThe text content is written in the correct order, and then the diffusion encoder completes the image details and overall visual presentation. The accuracy of the text is higher, and the overall image still has texture and style.
Overall, GLM-Image demonstrates excellent Chinese command understanding and high text generation accuracy. It also exhibits minimal garbled text even with high information density. Adding quotation marks to text further enhances accuracy. For content creators, it's a true productivity booster.
GLM-Image, based on the Huawei A2 chip and the MindSpeed training framework, has successfully completed the entire process from data preprocessing to...Large ModelThe complete training process illustrates the cutting-edge capabilities of domestically developed full-stack computing platforms.MultimodalThe model also has a real-world path to be fully trained and continuously iterated, so it is no longer dependent on foreign computing power.
GLM-Image is not only usable and easy to use, but also very cost-effective. Currently GLM-Image generates an image for only 0.1 yuan per image.The Nano Banana Pro costs about one yuan per sheet, which is 10 times more expensive.
The value of GLM-Image lies not only in its capabilities, but also in its validation of a replicable technological path, paving the way for future advancements.MultimodalThe model provides a reference engineering paradigm.
Within the domestic technology system, the path of cutting-edge models is not only feasible, but has also begun to be steadily developed.
Original link:Domestic Nano Banana open source10 Real-World Test Cases to Help You Quickly Understand GLM-Image