The most comprehensive Nano Banana tutorial online, including 4 free usage methods.
Google's newly released AI image generation model, Google Gemini 2.5 Flash Image (Nano Banana), continues to dominate various platforms thanks to its consistency and lightning-fast image generation speed. This article will discuss Nano...
Googleup to dateReleasedAIImage generation model Google Gemini 2.5 Flash Image (Nano Banana) continues to dominate various platforms with its consistency and lightning-fast image generation speed.
I also receive daily requests from group members urging me to update: "K-sister, please quickly introduce how to use Nano Banana, it's amazing!"
This article will discuss Nano Banana. The content mainly includes three parts: a technical explanation of the Nano Banana core team, various ways to use Nano Banana, and...freeFive ways to use Nano Banana.
Feel free to add your thoughts and share your opinions in the comments section!
In a recent Release Notes interview, we invited four core members of the Nano Banana team to help us better understand the technology behind Nano Banana's key features:
Native image generation
Nano Banana's core breakthrough is native image generation, which continuously references the context during the generation process to complete complex tasks step by step.
Unlike traditional text-based image tools, Imagen is more like a single-point expert, while Nano Banana is an all-rounder that can cross modalities and support complex interactions.
This will also be the direction of the team's development: future models should not only be drawn beautifully, but also have stronger understanding and reasoning abilities.
Consistency between character and scene
Consistency is another highlight of Nano Banana. Past models often crashed when editing images. For example, if I just wanted to change the curtains, the bed and sofa would also change, or if I changed the angle of a person's image, the face would change.
The team demonstrated a case study on-site, using a close-up of the host's face to generate a full-body image of him wearing a giant banana costume.
prompt: zoom out and show him wearing a giant banana costume. keep his face visble.
Prompt wordsZoom out to show him wearing a giant banana costume, ensuring his face is visible.
Nano Banana maintains excellent consistency; the face remains the same, but the scene and clothing have been completely changed, yet the image still looks very natural. The entire generation process takes only about ten seconds.
Another detail is,Gemini The collaboration between our team and the Imagen team has resulted in more natural-looking images. Previously, the results sometimes looked "pasted on," but now we can achieve a coherent overall effect.
Text rendering
Many people might think that "writing a few words in an image" is no big deal, but Nano Banana considers text rendering a core long-term metric. They believe that text is structured content, and if a model can learn how to process text, it can also grasp more complex structures such as textures in images.
Currently, Nano Banana is suitable for some...SimpleThe text rendering effect is very good, but there are also some shortcomings.
prompt: now write"Gemini nano on the image.
Prompt wordsWrite " on the picture"Gemini nano
A few years ago, almost no models could handle text well, even very short ones.Prompt wordsIt frequently crashed. Therefore, the Nano Banana team decided to track this metric long-term; regardless of the experiments conducted, continuous observation would prevent performance degradation. They discovered that many seemingly unrelated changes could also improve text rendering.
The team actually started by "finding the shortcomings of the model" and gradually explored a path that could drive overall quality improvement.
Comprehension and Creativity
Nano Banana possesses global knowledge, can understand ambiguous instructions, and can even unleash some creativity.
The team mentioned an interesting concept during the interview—reporting bias. For example, when you visit a friend's house, you almost never mention the ordinary sofa in their home when you chat with others afterwards. But if you show someone a photo, the sofa will be there.
Therefore, if we want to truly understand the world, we may need more description through text, while visual signals are like a shortcut to understanding the world, which can directly show the environment, objects, and relationships without the need for additional explanation.
Understanding and generation are thus complementary. As the model interprets images and language, it accumulates a more solid understanding of the world, resulting in more stable and natural creations. Sometimes, it can even generate content that exceeds user expectations.
fromPopularIt offers various gameplay options, including single-image, multi-image, and video generation. Feel free to add more suggestions!
PopularGameplay
1. Turn any image into a figurine.
prompt: turn this photo into a character figure. Behind it, place a box with the character’s image printed on it, and a computer showing the Blender modeling process on its screen. In front of the box, add a round plastic base with the character figure standing on.
Transform this photo into a character. Place a box behind the character, with the character's image printed on it, and display it on a computer screen above the box.BlenderModeling process. Add a circular plastic base to the front of the box, and the character stands on it.
2. Draw the real scene based on the map.
prompt: draw what the red arrow sees.
Prompt wordsDraw what the red arrow points to.
Draw a DEM with contour lines.
draw the real world view from the red circle in the direction of the arrow.
Draw with contour linesDigital Elevation Model.
Draw a real-world view from the red circle in the direction of the arrow.
The above case study comes from blogger X @Simon
3. Cartoons become reality
prompt: Depict as a live big budget costume test on set, shot on film.
Variant Prompt: For easier additional editing. Depict as a live big budget costume test on set, shot on film against green screen.
It depicts a high-budget costume fitting taking place on set, shot using film.
Variant tip: To make additional editing easier, depict it as a big-budget costume test on set, shot on set, using film against a green screen.
The above case study comes from blogger @Brent Lynch.
4.360-degree product display
prompt: This exact car in this exact environment.
Change Perspective: Perfect side angle view.
This car and its exact surroundings.
Changing the perspective: The perfect side view
After generating images from different perspectives, use Keling 2.1 to generate a video from the first and last frames.
The above case study comes from blogger X @Rory Flynn
5. Restoring old photos
prompt: Restore and colorize the picture without altering, removing, or adding any detail or element.
Prompt words: Restore the image color, but do not change, delete, or add any details or elements.
The above case study comes from blogger @Rodrigo Bressane.
6. Isometric 3D View
Transform design drawings into 3D views.
The above case study comes from blogger X @levelsio
The transformation from 2D drawings to 3D models looks quite impressive. However, the current feedback indicates that the generated images are not accurate enough; for example, they are somewhat blurry, and the window positions are inaccurate.
7. Switch perspectives
prompt: aerial perspective of a camera behind the blurry ceiling fan looking down at the girl sitting in an hospital waiting room.
An aerial perspective view of a girl sitting in a hospital waiting room, taken from behind a blurry ceiling fan and behind a camera.
Single image editing
Prompt wordsChange the text in the image to: Why don't you ask God Gemini?
Prompt wordsChange the character's clothes to a down jacket.
3. Refer to the character generation scenario
Prompt wordsThe image depicts the characters having dinner with SpongeBob SquarePants.
Prompt wordsRemove the person on the left side of the image.
5. Change elements
Prompt wordsReplace the old Twitter logo in the background of the image with the current X-shaped logo.
6. Understanding Fuzzy Instructions
Prompt wordsTo make the people in the picture look like Native Americans.
The first step is to replace the background with a green screen; subsequent background replacements will produce better results.
prompt:Replace the background with a solid color green screen
Prompt wordsReplace the background with a solid green screen.
prompt: replace the background with the attached image. Make sure [subject] is lit to match the image;
replace the background with [describe your scene]. Make sure [subject] is lit to match the scene.
Prompt wordsReplace the background with an additional image. Ensure the lighting of the [subject] matches the image.
Replace the background with [Describe your scene]. Ensure the lighting of the [subject] matches the scene.
8. Using Nano Banana for interior design
Nano Banana after "renovation":
9. Real-world scenes become game assets
prompt: Concisely name the key entity in this image (e.g. person, object, building). Create 3d pixel art of the isolated key entity in isometric perspective, 8-bit sprite on a white background. No drop shadow.
Prompt wordsConcisely name key entities in the image (e.g., people, objects, buildings). Create independent 3D pixel elements from a uniform perspective, 8-bit transparent images, and no shadow effects.
10. Transform city buildings into 3D
The content in 【】 can be modified according to the actual city.
prompt: Turn this photograph of a 【Parisian building】 into a isometric tile, in the style of the five other 3D.
Prompt wordsConvert the Parisian buildings in this image into 3D isometric models.
The above case study comes from blogger X @Emm | scenario.com
Multiple Image Editing
1. Attitude Reference
Prompt: take the anime man and woman in the first image and put them in the poses of the stick man in blue and stick woman in red. erase the stick figures.
Prompt wordsPlace the anime male and female figures in the first image into the poses of a blue little man and a red little woman, and then erase the little figures.
The above case study comes from blogger X @Justine Moore
prompt: Model pose like the sketch.
Prompt wordsThe model's pose became like a sketch.
2. Image position reference
Nano Banana can generate accurate images simply by marking locations on the graph.
Add to the diagramPrompt wordsGenerate comic:
The above case comes from X blogger @けいすけ/ AIMalinka & Kairi
prompt: A model is posing and leaning against a pink bmw. She is wearing the following items, the scene is against a light gray background. The green alien is a keychain and it's attached to the pink handbag. The model also has a pink parrot on her shoulder. There is a pug sitting next to her wearing a pink collar and gold headphones.
Prompt wordsA model leans against a pink BMW against a light gray background. She is wearing the following items: a green alien keychain hanging from a pink handbag, and a pink parrot perched on her shoulder. Next to her is a poodle wearing a pink collar and gold headphones.
This case study comes from blogger @Travis Davids.
Series IP/Anime Characters
For example, we have this character image.
prompt: First, please set up the basic color palette and the shadows and saturation.
Prompt wordsFirst, set the basic color palette and shadows and saturation.
prompt:Next, please do the character model sheet.
Next, please create the character model table.
prompt:Next, please provide the [basic action set].
Next, please provide the basic motion set.
prompt: Please give me the costume design set.
Prompt wordsPlease give me a clothing design kit.
prompt: Please make an expression sheet.
Prompt wordsPlease create an emoji set.
Convert the image to line art, then color it using the brand's colors.
step:
– Prepare the original image
– Convert to line art
– Color the line art with a color palette
– Change the character to brand colors
– Prepare the original image
– Convert to line art
– Use a color palette to color the line art
– Change the character's color to the brand color
Video (storyboard) creation
1. A first-person perspective of riding a horse through the 20th century.
First, generate various scenes using Nano Banana:
prompt: dashcam google street view shot | Hobbiton streets | hobbits carrying out daily tasks like gardening and smoking pipes | sunny day.
Prompt wordsDashcam footage from Google Street View | Streets of Hobbiton | Hobbits performing daily tasks, such as gardening and smoking | Sunny day
prompt: dashcam google street view shot | Seat of Seeing on Amon Hen | Ruined pavilion atop the hill, a hobbit-like figure from behind climbing the steps, the winding path down visible overlooking the river and lands beyond | panoramic view under emerging stars at dusk.
Prompt words: Dashcam footage from Google Street View | Amon Hen's Viewpoint: A dilapidated pavilion atop a mountain, a hobbit-like figure climbs the steps from behind, revealing a winding path leading down, overlooking the river and distant lands, a panoramic view under the first stars at dusk.
usePrompt wordsCreate a first-person perspective image of horseback riding:
prompt: dashcam google street view shot
Prompt wordsFirst-person perspective: Riding a horse across a meadow. Twelfth century.
Use the first and last frame animation of Keling 2.1 to generate video clips.
prompt: "scene_description": "The rider gallops out from the ruins of the ivy-covered statues, leaving the storm-lit plains behind. The path winds through rugged terrain as the pace remains fast. Ahead, a towering dark castle glows with eerie green light atop jagged cliffs, its spires piercing the stormy sky. Cloaked figures march steadily toward the fortress across a massive stone bridge.", "visual_style": "dark epic fantasy, cinematic, continuous POV", "camera_movement": "smooth forward gallop, first-person view without cuts, transitioning naturally from the ruined statues across the plains to the castle bridge", "main_subject": "the white horse’s head and rider’s gloved hands, centered as they race toward the looming fortress", "background_setting": "storm-darkened mountains and cliffs, a vast stone bridge spanning a deep chasm, leading to the glowing green-lit castle", "lighting_mood": "ominous twilight with green highlights from the fortress and flashes of distant lightning"
prompt: "scene_description": "The rider gallops out from the ruins of the ivy-covered statues, leaving the storm-lit plains behind. The path winds through rugged terrain as the pace remains fast. Ahead, a towering dark castle glows with eerie green light atop jagged cliffs, its spires piercing the stormy sky. Cloaked figures march steadily toward the fortress across a massive stone bridge.", "visual_style": "dark epic fantasy, cinematic, continuous POV", "camera_movement": "smooth forward gallop, first-person view without cuts, transitioning naturally from the ruined statues across the plains to the castle bridge", "main_subject": "the white horse’s head and rider’s gloved hands, centered as they race toward the looming fortress", "background_setting": "storm-darkened mountains and cliffs, a vast stone bridge spanning a deep chasm, leading to the glowing green-lit castle", "lighting_mood": "ominous twilight with green highlights from the fortress and flashes of distant lightning
By editing these videos together, this long, time-traveling video was created.
This case study comes from blogger @TechHalla.
2. Graffiti - 3D Images - Videos
This case comes from blogger @Alex Patrascu.
3. Let the figures in famous paintings meet in the real world.
This case comes from blogger @Alex Patrascu.
4.AIcartoon
This case comes from blogger @Framer.
GoogleGemini(Pro membership required)
exist Gemini On the official website homepage, select Gemini For the 2.5 Pro model, select "Create images" under "Tool" in the dialog box.
The Nano Banana model is used by default at this point.
Upload image, enterPrompt wordsIt can be used immediately.
Google AI Studio (free)
Open Geogle AI On the Studio website, click on settings in the upper right corner.
Select Nano Banana in the settings.
Upload an image and enter...Prompt wordsThat's fine.
LMArena (free)
Select Direct Chat mode at the top of the LMArena homepage.
Continue by selecting the gemini-2.5-flash-image-preview (nano-banana) model, and you can use it directly.
Lovart (Limited Time Offer)free)
Click on the Nano Banana model entry on the Lovart homepage.
You can then directly use the Nano Banana model, however...freeThere is a deadline; it ends on September 2nd.
Freepik
Official website: Freepik
Select the Nano Banana model.freeUse it daily.freeGenerate 10 images.
In the past, creating a set of high-quality, consistent drawings required professional skills and a significant amount of time.
Now with Nano Banana, you can use just a short command to recall and reproduce what you say in a few seconds.
The content industry has been hit the hardest. Those who dare to use it now... AI There aren't many companies doing advertising creative work yet, but things might change drastically after this.
With just a prompt from the brand, dozens of ad creatives can be generated within hours.
The same applies to animation and short dramas. Details that used to require an entire team to work on for months can now be completed by a single person while making changes.
In the future, storylines may be generated "on demand"—as soon as viewers leave comments, the model immediately continues to create the story.
Content creation will enter a completely new production mode.