Google's Nano Banana Pro Professional Generation Guide: 10 Tips (Chinese & English Version)
Google's Nano Banana Pro guide introduces the core features and application techniques of the professional image generation model, Nano Banana Pro. The article emphasizes the model's breakthroughs in generating professional assets, covering text rendering...
Google's Nano Banana Pro guide introduces the core functions and application techniques of the professional image generation model, Nano Banana Pro. The article emphasizes the model's breakthroughs in generating professional assets, covering ten core capabilities including text rendering, character consistency, visual compositing, Google search integration, advanced editing, 2D/3D conversion, and high-resolution output. It provides detailed best practice examples for each function, guiding users on how to interact with the model like a creative director using natural language commands to efficiently produce high-quality, commercial-grade image content.
Nano-Banana ProThis represents a significant leap forward compared to previous generations, upgrading from "entertainment-oriented" image generation to "functional" professional asset creation.It excels in text rendering, character consistency, visual composition, world knowledge (search), and high-resolution (4K) output.
This article includes the following content::
0. The Golden Rule of Prompt Keywords
1. Text rendering, infographics, and visual compositing
2. A thumbnail of role consistency and viral spread
3.Using Google search for authenticity verification
4. Advanced editing, repair, and coloring
5. Dimensional Transformation (2D) ↔ 3D)
6. High resolution and texture enhancement
7. Thinking and reasoning ability
8. One-time storyboards and concept designs
9. Structural control and layout guidance
Tips for the Golden Rule
Nano-Banana Pro is a thinking model that can match keywords and supports understanding intent, physics, and composition. For best results, abandon traditional label instructions (e.g., dog, park, 4K, realistic) and treat it as a design partner for creation.
Refinement is better than regeneration
The model excels at understanding conversational modification commands. If an image already meets 80% of expectations, don't generate a new image from scratch. Simply state your specific adjustment requirements.
For example: "The effect is good. Please adjust the lighting to a sunset atmosphere and change the text to neon blue."
Use natural language and complete sentence structures
Engage with the model as if you were explaining your needs to a human designer. Use standard grammar and vivid adjectives.
❌ Negative example: "Cool cars, neon lights, city, night, 8K".
✅ Positive example: "A cinematic wide-angle shot: a futuristic sports car speeds through the rainy streets of Tokyo at night, the neon lights reflecting off the wet pavement and the metal chassis of the car in a dazzling display."
The description needs to be specific and concrete.
The more vague the prompts, the worse the generated effect. Please clearly define the subject, environment, lighting, and atmosphere.
main bodyIt should not be simply said that "a lady"; it should be described as "an elegant older woman wearing a vintage Chanel-style suit".
MaterialTexture needs to be described. For example, "matte texture", "brushed stainless steel", "soft velvet" or "creased paper".
Provide background information ("Why" and "For whom")
Because the model possesses the ability to "think," providing context helps in making logical artistic decisions.
Example"Create an image of a sandwich for a high-end Brazilian cookbook." (The model will infer from this that professional plating, shallow depth of field, and perfect lighting are required.)
Text rendering, infographics and visual composition
Nano-Banana Pro boasts advanced generation capabilities, rendering clear, readable, and stylized text, and synthesizing complex information into visual formats.
Best Practice Recommendations:
- Information compressionThe model can be asked to "compress" dense text or PDF content into visual charts.
- Style specificationClearly specify the style you want, such as "elegant magazine layout", "technical schematics", or "hand-drawn whiteboard style".
- Text citationUse quotation marks to explicitly indicate the specific text that needs to be displayed.
Prompt word examples:
Financial Statement Infographic (Data Input):
[Enter the latest Google search results]Financial Report[PDF file]
"Please generate a concise, modern infographic summarizing the key financial highlights of this financial report. It should include charts for 'Revenue Growth' and 'Net Profit,' and highlight the CEO's key quotes in a stylized quotation box format."
Retro infographic:
“Create a retro infographic in the style of the 1950s to introduce the history of American restaurants. The infographic should include separate sections such as ‘Food,’ ‘Jukebox,’ and ‘Decor.’ Ensure that all text is clear, legible, and in line with the style of the time.”
Technical drawings:
“Draw an orthographic blueprint to describe the building in the form of a plan, elevation, and section. Clearly label the ‘North Elevation’ and ‘Main Entrance’ using a professional architectural font. The format should be 16:9.”
Whiteboard summary (for teaching purposes):
The concept of "Transformer Neural Network Architecture" is summarized using hand-drawn whiteboard diagrams, suitable for university lectures. The encoder and decoder modules are marked with different colored markers, and "self-attention" and "feedforward" are clearly labeled.
Role consistency and viral thumbnails
Nano-Banana Pro supports up to 14 reference images (6 of which are high-fidelity). This enables the "identity lock" function—seamlessly placing a specific person or character into a new scene while maintaining their facial features.
Best Practice Recommendations:
- Identity LockThe explicit instruction is: "Ensure that the facial features of the person are completely consistent with Figure 1."
- Facial expressions and actionsWhile locking onto an identity, users can freely describe changes in emotions or posture.
- Viral compositionIt combines the main subject with eye-catching graphics and text in one go to create a composition with strong communicative power.
Example prompt:
"Viral thumbnails" (logo + text + graphic):
Design a viral video thumbnail using the person in Figure 1. Facial Consistency: Keep the person's facial features identical to those in Figure 1, but change their expression to make them look excited and surprised. Action: Place the person on the left side of the frame, pointing their finger to the right. Subject: Place a high-resolution image of delicious avocado toast on the right side. Graphic: Add a striking yellow arrow connecting the person's finger and the toast. Text: Overlay the eye-catching, trendy text "3minuteFudede!" in the center. Use thick white lines and shadows. Background: A blurred, bright kitchen background. High saturation and high contrast.
The "Furry Friend" Scenario (Group Consistency):
[Enter 3 pictures of different plush toys]
Please create a fun ten-page story about three furry friends going on a tropical vacation. The plot should be exciting and suspenseful, ending with a heartwarming conclusion. The three characters should maintain the same clothing and appearance, but their expressions and angles should vary across the ten illustrations. Each character should appear only once in each illustration.
Brand equity creation:
[Enter a product image]
"Please create nine stunning fashion editorials, with a style reminiscent of award-winning fashion magazine spreads. Use these as a reference for your brand's style, making subtle adjustments and variations to showcase professional design. Please create one image at a time, for a total of nine images."
Using Google search for authenticity verification
Nano-Banana Pro can access Google search and generate images based on real-time data, current events, or fact-checking, reducing information errors on time-sensitive topics.
Best Practices Recommendations:
- It can be required to visualize dynamic data (such as weather, stock market, news).
- Before generating an image, the model will "think" (reason) about the search results and then begin creating.
Example prompt:
Event visualization:
"Based on current travel trends, generate an infographic showing the best time to visit U.S. national parks in 2025."
Advanced editing, repair, and coloring
The model excels at performing complex editing tasks through conversational commands, including "partial redrawing" (removing/adding objects), "image restoration" (restoring old photos), "intelligent colorization" (comics/black and white photos), and "style transfer".
Best Practices:
- Semantic instructionsNo need to manually apply masks; simply describe your modification needs naturally.
- Physical logic understandingIt can generate complex commands such as "fill the glass with liquid" to test its ability to generate physical simulations.
Example prompts:
Object removal and completion:
“Remove the tourists from the background of this photo and fill the space with reasonable textures (pebbles and storefronts) that match the surrounding environment.”
Comics/Comic Coloring:
[Enter black and white comic book storyboard]
"Colorize this comic panel. Use a vibrant anime-style color scheme. Ensure the energy beam lighting effect is neon blue, and that the character's clothing colors match their official color scheme."
Localization (text translation + cultural adaptation):
[Enter an image of a London bus stop advertisement]
“Localize this concept to the Tokyo setting, including translating the slogans into Japanese. Change the background to the bustling streets of Shibuya at night.”
Lighting/Seasonal Control:
[Enter a picture of a house in summer]
“Transform this scene into winter. Keep the building structure exactly the same, but add snow to the roof and yard, and change the lighting to a cold, gloomy afternoon.”
Dimensional transformation (2D ↔ 3D)
A powerful new feature allows you to convert 2D floor plans into 3D visualizations and vice versa. This is ideal for interior designers, architects, and even meme creators.
Example prompt:
2D floor plan to 3D interior design rendering:
Generate a professional interior design rendering based on the uploaded 2D floor plan.layoutThe design uses a collage format, with a main image at the top (wide-angle view of the living room) and three smaller images below (master bedroom, home office, and 3D top view).styleAll images feature a modern minimalist style, paired with warm oak floors and off-white walls.qualityPhotorealistic rendering with soft, natural light.
2D to 3D emoji conversion:
"Transform the 'All's Well' dog emoji into a realistic 3D rendering. Keep the composition the same, but make the dog look like a plush toy and the flames look like real flames."
Nano-Banana Pro natively supports image generation from 1K to 4K. This feature is especially useful for rendering fine textures or producing large-format prints.
Best Practices Recommendations:
- High-resolution requestIf your API or interface allows, please explicitly request the generation of 2K or 4K high-resolution images.
- High-fidelity detailed descriptionThe prompts describe high-fidelity details such as minor imperfections and complex surface textures.
4K texture generation:
“Utilizing native high-fidelity output, we have created a breathtaking and atmospheric mossy forest ground environment. We have mastered complex lighting effects and delicate textures, ensuring that every moss and every ray of light is rendered at pixel-level resolution to meet the needs of 4K wallpapers.”
Complex logic (thinking patterns):
“Create a hyper-realistic infographic of a premium cheeseburger, breaking it down to show the texture of the toasted burrito bun, the caramelized crust of the patty, and the glistening, melted cheese. Label each layer with its flavor profile.”
Nano-Banana Pro defaults to "Think" mode, where the model generates intermediate thinking images (at no extra cost) to optimize the composition before rendering the final output. This is helpful for data analysis and solving visual problems.
Example prompts:
Solve the equation:
Solve the equation log_{x^2+1}(x^4-1)=2 using C language on a whiteboard. Please clearly write out the solution steps.
Visual reasoning:
“Analyze this room image to generate a ‘before’ image, showing what the room might have looked like during construction, including the frame and unfinished drywall.”
One-time storyboards and concept designs
You can generate sequence images or storyboards directly without using panel templates, ensuring a coherent narrative flow within a single session. This feature is widely used for creating "film concept art" (e.g., releasing fake leaked images of upcoming films).
Example prompt:
Create a compelling nine-part story, comprising nine images, featuring a woman and a man shooting an award-winning luxury luggage advertisement. The story should have emotional ups and downs, ending with an elegant photograph of the woman holding the brand logo. The male and female leads must maintain the same identity and attire, but can be photographed from different angles and distances. Generate the images one by one. Ensure each image is in 16:9 landscape format.
Reference images aren't limited to characters or objects to be edited. You can use them to maintain strict control over the composition and layout of the final image. This is undoubtedly a revolutionary feature for designers who need to transform sketches, wireframes, or specific grids into beautiful assets.
Best Practices Recommendations:
- Draft and sketchesUpload a hand-drawn sketch and precisely specify the positions of text and objects.
- WireframeGenerate high-fidelity UI mockups using screenshots of existing layouts or wireframes.
- Grid: Using mesh images, drive the model to generate image assets specifically designed for tile-style games or LED displays.
Example prompt:
From sketch to final advertisement:
"Create an advertisement for [the product] based on this sketch."
Create a UI model based on the wireframe.:
"Create a model for [the product] according to the following guidelines."
Pixel Art and LED Displays:
"Please create a pixel art unicorn that perfectly fits this 64x64 grid image. Use high-contrast colors."
(Hint: Developers can then programmatically extract the center color of each cell to drive the connected 64×64 LED dot matrix display.)
Sprite Image Example:
"A sprite of a woman performing a backflip on a drone, 3x3 grid, frame-by-frame animation sequence, square aspect ratio. Please draw exactly according to the structure of the attached reference image."
(Tip: You can extract each cell and create a GIF animation.)
Now that we've grasped the basics of cue words, we can begin constructing them:
- Experiment in the interfaceGoogle AI Studio is the fastest way to test prompts and parameters.
- View Featured Apps:existApp GalleryExperience cool apps powered by Nano-Banana.
- Turning ideas into applicationsIn AI Studio Build, easily turn your most successful suggestions into apps that you can share with your friends.
- Building applicationsReady to write code? Please refer to...Developer Guideor Gemini API Sample LibraryGet the guide and code snippets.
- In-depth exploration of technologyRead the full article Gemini API See the documentation for details on rate limits, pricing, and integration.
Original English text
Nano-Banana Pro is a significant leap forward from previous generation models, moving from "fun" image generation to "functional" professional asset production. It excels in text rendering, character consistency, visual synthesis, world knowledge (Search), and high-resolution (4K) output.
Following the developer guide on how to get started with AI Studio and the API, this guide covers the core capabilities and how to prompt them effectively.
By Guillaume Vernade, Gemini Developer Advocate, Google DeepMind
Here's what you'll find in this article:
- The Golden Rules of Prompting
- Text Rendering, Infographics & Visual Synthesis
- Character Consistency & Viral Thumbnails
- Grounding with Google Search
- Advanced Editing, Restoration & Colorization
- Dimensional Translation (2D ↔ 3D)
- High-Resolution & Textures
- Thinking & Reasoning
- One-Shot Storyboarding & Concept Art
- Structural Control & Layout Guidance
- What's next?
Section 0: The Golden Rules of Prompting
Nano-Banana Pro is a "Thinking" model. It doesn't just match keywords; it understands intent, physics, and composition. To get the best results, stop using "tag soups" (e.g., dog, park, 4k, realistic) and start acting like a Creative Director.
Edit, Don't Re-roll
The model is exceptionally good at understanding conversational edits. If an image is 80% correct, do not generate a new one from scratch. Instead, simply ask for the specific change you need.
Example: "That's great, but change the lighting to sunset and make the text neon blue."
Use Natural Language & Full Sentences
Talk to the model as if you were briefing a human artist. Use proper grammar and descriptive adjectives.
❌ Bad: "Cool car, neon, city, night, 8k."
✅ Good: "A cinematic wide shot of a futuristic sports car speeding through a rainy Tokyo street at night. The neon signs reflect off the wet pavement and the car's metallic chassis."
Be Specific and Descriptive
Vague prompts yield generic results. Define the subject, the setting, the lighting, and the mood.
Subject: Instead of "a woman," say "a sophisticated elderly woman wearing a vintage chanel-style suit."
Materiality: Describe textures. "Matte finish," "brushed steel," "soft velvet," "crumpled paper."
Provide Context (The "Why" or "For whom")
Because the model "thinks," giving it context helps it make logical artistic decisions.
Example: "Create an image of a sandwich for a Brazilian high-end gourmet cookbook." (The model will infer professional plating, shallow depth of field, and perfect lighting).
Text Rendering, Infographics & Visual Synthesis
Nano-Banana Pro has SOTA capabilities for rendering legible, stylized text and synthesizing complex information into visual formats.
Best Practices:
- Compression: Ask the model to "compress" dense text or PDFs into visual aids.
- Style: Specify if you want a "polished editorial," a "technical diagram," or a "hand-drawn whiteboard" look.
- Quotes: Clearly specify the text you want in quotes.
Example Prompts:
Earnings Report Infographic (Data Ingestion):
[Input PDF of Google's latest earnings report]
"Generate a clean, modern infographic summarizing the key financial highlights from this earnings report. Include charts for 'Revenue Growth' and 'Net Income', and highlight the CEO's key quote in a stylized pull-quote box."
Retro Infographic:
"Make a retro, 1950s-style infographic about the history of the American diner. Include distinct sections for 'The Food,' 'The Jukebox,' and 'The Decor.' Ensure all text is legible and stylized to match the period."
Technical Diagram:
"Create an orthographic blueprint that describes this building in plan, elevation, and section. Label the 'North Elevation' and 'Main Entrance' clearly in technical architectural font. Format 16:9."
Whiteboard Summary (Educational):
"Summarize the concept of 'Transformer Neural Network Architecture' as a hand-drawn whiteboard diagram suitable for a university lecture. Use different colored markers for the Encoder and Decoder blocks, and include legible labels for 'Self-Attention' and 'Feed Forward'."
Character Consistency & Viral Thumbnails
Nano-Banana Pro supports up to 14 reference images (6 with high fidelity). This allows for "Identity Locking"—placing a specific person or character into new scenarios without facial distortion.
Best Practices:
- Identity Locking: Explicitly state: "Keep the person's facial features exactly the same as Image 1."
- Expression/Action: Describe the change in emotion or pose while maintaining the identity.
- Viral Composition: Combine subjects with bold graphics and text in a single pass.
Example Prompts:
The "Viral Thumbnail" (Identity + Text + Graphics):
"Design a viral video thumbnail using the person from Image 1. Face Consistency: Keep the person's facial features exactly the same as Image 1, but change their expression to look excited and surprised. Action: Pose the person on the left side, pointing their finger towards the right side of the frame. Subject: On the right side, place a high-quality image of a delicious avocado toast. Graphics: Add a bold yellow arrow connecting the person's finger to the toast. Text: Overlay massive, pop-style text in the middle: 'Done in 3 minutes!' (Done in 3 mins!). Use a thick white outline and drop shadow. Background: A blurred, bright kitchen background. High saturation and contrast."
The "Fluffy Friends" Scenario (Group Consistency):
[Input 3 images of different plush creatures]
"Create a funny 10-part story with these 3 fluffy friends going on a tropical vacation. The story is thrilling throughout with emotional highs and lows and ends in a happy moment. Keep the attire and identity consistent for all 3 characters, but their expressions and angles should vary throughout all 10 images. Make sure to only have one of each character in each image."
Brand Asset Generation:
[Input 1 image of a product]
"Create 9 stunning fashion shots as if they’re from an award-winning fashion editorial. Use this reference as the brand style but add nuance and variety to the range so they convey a professional design touch. Please generate nine images, one at a time."
Grounding with Google Search
Nano-Banana Pro uses Google Search to generate imagery based on real-time data, current events, or factual verification, reducing hallucinations on timely topics.
Best Practices:
- Ask for visualizations of dynamic data (weather, stocks, news).
- The model will "Think" (reason) about the search results before generating the image.
Example Prompts:
Event Visualization:
"Generate an infographic of the best times to visit the U.S. National Parks in 2025 based on current travel trends."
Advanced Editing, Restoration & Colorization
The model excels at complex edits via conversational prompting. This includes "In-painting" (removing/adding objects), "Restoration" (fixing old photos), "Colorization" (Manga/B&W photos), and "Style Swapping."
Best Practices:
- Semantic Instructions: You do not need to manually mask; simply tell the model what to change naturally.
- Physics Understanding: You can ask for complex changes like "fill this glass with liquid" to test physics generation.
Example Prompts:
Object Removal & In-painting:
"Remove the tourists from the background of this photo and fill the space with logical textures (cobblestones and storefronts) that match the surrounding environment."
Manga/Comic Colorization:
[Input black and white manga panel]
"Colorize this manga panel. Use a vibrant anime style palette. Ensure the lighting effects on the energy beams are glowing neon blue and the character's outfit is consistent with their official colors."
Localization (Text Translation + Cultural Adaptation):
[Input image of a London bus stop ad]
"Take this concept and localize it to a Tokyo setting, including translating the tagline into Japanese. Change the background to a bustling Shibuya street at night."
Lighting/Seasonal Control:
[Input image of a house in summer]
"Turn this scene into winter time. Keep the house architecture exactly the same, but add snow to the roof and yard, and change the lighting to a cold, overcast afternoon."
Dimensional Translation (2D ↔ 3D)
A powerful new capability is translating 2D schematics into 3D visualizations, or vice versa. This is ideal for interior designers, architects, and meme creators.
Example Prompts:
2D Floor Plan to 3D Interior Design Board:
"Based on the uploaded 2D floor plan, generate a professional interior design presentation board in a single image. Layout: A collage with one large main image at the top (wide-angle perspective of the living area), and three smaller images below (Master Bedroom, Home Office, and a 3D top-down floor plan). Style: Apply a Modern Minimalist style with warm oak wood flooring and off-white walls across ALL images. Quality: Photorealistic rendering, soft natural lighting."
2D to 3D Meme Conversion:
"Turn the 'This is Fine' dog meme into a photorealistic 3D render. Keep the composition identical but make the dog look like a plush toy and the fire look like realistic flames."
High-Resolution & Textures
Nano-Banana Pro supports native 1K to 4K image generation. This is particularly useful for detailed textures or large-format prints.
Best Practices:
- Explicitly request high resolutions (2K or 4K) if your API/Interface allows.
- Describe high-fidelity details (imperfections, surface textures).
4K Texture Generation:
"Harness native high-fidelity output to craft a breathtaking, atmospheric environment of a mossy forest floor. Command complex lighting effects and delicate textures, ensuring every strand of moss and beam of light is rendered in pixel-perfect resolution suitable for a 4K wallpaper."
Complex Logic (Thinking Mode):
"Create a hyper-realistic infographic of a gourmet cheeseburger, deconstructed to show the texture of the toasted brioche bun, the seared crust of the patty, and the glistening melt of the cheese. Label each layer with its flavor profile."
Thinking & Reasoning
Nano-Banana Pro defaults to a "Thinking" process where it generates interim thought images (not charged) to refine composition before rendering the final output. This allows for data analysis and solving visual problems.
Example Prompts:
Solve Equations:
"Solve log_{x^2+1}(x^4-1)=2 in C on a white board. Show the steps clearly."
Visual Reasoning:
"Analyze this image of a room and generate a 'before' image that shows what the room might have looked like during construction, showing the framing and unfinished drywall."
One-Shot Storyboarding & Concept Art
You can generate sequential art or storyboards without a grid, ensuring a cohesive narrative flow in a single session. This is also popular for "Movie Concept Art" (e.g., fake leaks of upcoming films).
Example Prompt:
"Create an addictively intriguing 9-part story with 9 images featuring a woman and man in an award-winning luxury luggage commercial. The story should have emotional highs and lows, ending on an elegant shot of the woman with the logo. The identity of the woman and man and their attire must stay consistent throughout but they can and should be seen from different angles and distances. Please generate images one at a time. Make sure every image is in a 16:9 landscape format."
Structural Control & Layout Guidance
Input images aren't limited to character references or subjects to edit. You can use them to strictly control the composition and layout of the final output. This is a game-changer for designers who need to turn a napkin sketch, a wireframe, or a specific grid layout into a polished asset.
Best Practices:
- Drafts & Sketches: Upload a hand-drawn sketch to define exactly where the text and object should sit.
- Wireframes: Use screenshots of existing layouts or wireframes to generate high-fidelity UI mockups.
- Grids: Use grid images to force the model to generate assets for tile-based games or LED displays.
Sketch to Final Ad:
"Create a ad for a [product] following this sketch."
UI Mockup from Wireframe:
"Create a mock-up for a [product] following these guidelines."
Pixel Art & LED Displays:
"Generate a pixel art sprite of a unicorn that fits perfectly into this 64×64 grid image. Use high contrast colors."
(Tip: Developers can then programmatically extract the center color of each cell to drive a connected 64×64 LED matrix display).
Sprites:
"Sprite sheet of a woman doing a backflip on a drone, 3×3 grid, sequence, frame by frame animation, square aspect ratio. Follow the structure of the attached reference image exactly.."
(Tip: You can then extract each cell and make a gif)
What's next?
Now that you have mastered the basics of prompting, here is how you can start building:
- Experiment in the UI: Google AI Studio is the fastest way to test prompts and parameters.
- Check, it's really cool. Nano-banana powered app in the App Gallery.
- Vibe-code you dream app: Transform your best prompt into an app that you can easily share with your friends in AI Studio Build.
- Build Applications: Ready to code? Check out the developer guide or the Gemini API Cookbook For guides and code snippets.
- Technical Deep DiveRead the full Gemini API Documentation for details on rate limits, pricing, and integration.