Nano Banana 2 Real-world Test - Google Gemini 3.1 Flash Image Tops Arena Raw Image Ranking
Nano Banana 2 is here. Yesterday, Google launched its latest image model, Nano Banana 2 (Gemini 3.1 Flash Image), which immediately topped the Arena raw image rankings! Lovart was among the first to experience it...
Nano Banana 2 is here.
Yesterday, Google launchedup to dateImage model Nano Banana 2 (Gemini 3.1 Flash Image), immediately topped the Arena raw image rankings upon release!
Not only is the understanding and adherence to instructions stronger, but the rendering of multilingual text is also more accurate, and it can directly output images with different aspect ratios and resolutions.
In short, it's higher quality, faster, and cheaper!
Lovart is available to users immediately, and this Pro member can also...freePlaying with the Nano Banana 2 is really great.
This article will share some creative ways to use the Nano Banana 2.
We open the Lovart website and select from the image templates. Nano Banana 2.
Official websitehttps://www.lovart.ai/
World Knowledge
Nano Banana 2 access Gemini Knowledge base and real-time search. For example, if we want Nano Banana 2 to generate a real-world image, it will first search for visual reference information and then render the image according to our requirements.
When taking raw photos, we choose Thinking mode, enable online search..
Prompt wordsHigh-definition photography, looking out the window of a cozy café, directly facing the Yellow Crane Tower, with real-time weather outside, 4K resolution, 16:9 aspect ratio.
Nano Banana 2 searches for and references the architectural features of the Yellow Crane Tower, and queries real-time weather information in Wuhan as a reference to generate images.
The generated image quality is excellent; if I didn't know there were no coffee shops opposite the Yellow Crane Tower, it would be really hard to tell that this is... AI The generated image.
Replacing the content within the brackets [ ] will generate images in other regions' photographic styles. This feature is especially popular on X.
The Nano Banana 2 can accurately recognize and understand the markings on the map.
Prompt wordsGenerates a panoramic view of the locations marked in red in the reference image, in an anime style.
Lovart accurately identified the location marked in red as the Oriental Pearl TV Tower, and based on my screenshot, it also fully displayed the view of the Huangpu River and the Lujiazui skyline.
Editing images in Lovart is very convenient; parameters such as lighting, exposure, contrast, and saturation can be adjusted directly in the raw image interface.
Lovart's infinite canvas style is also very easy to use.SimpleWe can click on the image in the canvas to add it as a reference image to the dialog. Let's try generating another night scene version.
Generate a panoramic night view of the Oriental Pearl Tower with an anime style and neon lighting effects.
The stylization is done quite well.
Text rendering
Nano Banana Pro's most popular feature is generating infographics, but Chinese text is only relatively accurate at 4K resolution.
For example, we can generate an infographic using Google's official description of Nano Banana 2.
Transform this article into a cartoon-style infographic, using hand-drawn illustrations to visually explain core concepts. Include a few simple cartoon elements, icons, and arrows to highlight keywords and core concepts, aiding user understanding. Chinese is used by default. The aspect ratio is 16:9.
When generating infographics with Nano Banana Pro, 4K resolution is the preferred option; otherwise, some text will inevitably appear blurry or misaligned. However, placing 4K images within articles takes a long time to load, which may negatively impact the reading experience.
My usual procedure is to first generate a 4K infographic, download it to my local machine, resize it, and then use it. Processing one or two images is fine, but if there are many images, it takes quite a bit of time.
Nano Banana Pro generates
Nano Banana 2 has upgraded its text rendering, and now even without selecting 4K resolution, text rendering is much more accurate.
The text generation is indeed free of garbled characters, but I found two typos. By pressing CTRL and left-clicking to select the location of the typos, we can make very accurate local corrections.
This one is fine:
Stronger consistency among subjects
Nano Banana 2's subject consistency is also stronger than Nano Banana Pro, maintaining similarity for up to 5 characters and fidelity for up to 14 objects in a single instance.
Analyze the overall composition of the reference image. Identify all key subjects (whether individuals, groups/couples, vehicles, or specific objects) and their spatial relationships/interactions. Generate a coherent 3x3 "relationship table" grid showing nine shots taken in the same environment that are identical to the aforementioned subjects.
Adjust standard film shot types according to content (e.g., keep the group intact if it's a group; shoot the entire object if it's an object):
First line (establishing the environment): Extreme long shot (ELS): The subject appears small in a vast environment. Long shot (LS): The subject or group is fully visible from top to bottom (from head to toe/from wheels to roof). Medium long shot (American lens/three-quarter view): Shot from above the knee (people) or from a three-quarter view (objects).
Second row (core coverage): 4. Medium shot (MS): Shot from waist up (or the center of the object). Focus on interaction/action. 5. Medium close-up (MCU): Shot from chest up. Intimate composition of the main subject. 6. Close-up (CU): Focus on the "front" of a face or object.
Third line (details and angles): 7. Extreme close-up (ECU): Highly focused on key features (eyes, hands, logos, textures), presenting macro-like details. 8. Low-angle shot (looking up): Looking up at the subject from the ground (creating an epic/heroic feel). 9. High-angle shot (bird's-eye view): Looking down at the subject from above.
Ensure strict consistency: the same characters/objects, clothing, and lighting must appear in all 9 shots. Depth of field should vary realistically (background blurring in close-ups).
Create a professional 3x3 cinematic storyboard grid containing 9 shots. The grid should represent the effect of a specific subject/scene in the input image at all focal lengths.
First row: Wide-angle environmental shots, full-body shots, three-quarter profile shots (above the knees). Second row: Waist-high and above perspective shots, chest-high and above perspective shots, close-ups of the face/front. Third row: Macro details, low angles, high angles.
All frames must have realistic textures, consistent cinematic color grading, and be correctly composed based on the number and type of subjects or objects being analyzed.
Reference image
Lovart immediately recognized this as a classic scene of a couple in the movie "Titanic" at dusk, and generated nine other shots of the same scene, maintaining excellent consistency in the characters.
The characters' facial features, clothing, and poses are all very well preserved. However, while scenes 1 and 9 look good individually, when compared side-by-side, the difference within the same scene is quite significant.
However, it is used for AI The panel layout for the comic book doesn't have much impact; it works very well.
Lovart integrates top-tier models from across the board, enabling seamless multi-model interaction. Once you've created a storyboard, you can directly generate video in one step – it's simply amazing! AI Short dramas, comicsHigh efficiencyproductive forces.
Prompt wordsA video was generated based on a reference image. As the rain fell, the entire bamboo forest seemed shrouded in mist. A young man in a blue robe, holding an oil-paper umbrella, slowly walked along the slippery stone steps, the water shimmering beneath his feet, the distant sound of a waterfall clearly audible. The forest was eerily quiet, save for the raindrops tapping on his umbrella, yet he seemed to be listening for other sounds. Reaching the depths of the forest, he stopped, looking up at the swaying bamboo shadows, as if someone was waiting for him in the shadows, or perhaps someone was chasing him. His fingers tightened on the umbrella handle, rainwater trickling down the wood grain. Finally, he took a deep breath, suppressing his hesitation, and stepped into the thicker mist to keep an old promise.
The overcast weather, bamboo forest, and thick fog, combined with the character's tense expression as he looks around, create a strong atmosphere. The camera cuts are also very smooth, and the information progresses naturally.
Batch generation
Lovart can generate 100 images in a single batch, and combined with Nano Banana 2's powerful and fast text-to-image generation capabilities, the image output efficiency skyrockets.
Please create an engaging comic strip story, referencing the girl's world travels depicted in the image. The story should be dramatic, filled with emotional highs and lows. Maintain consistency in the characters' clothing and appearance. Create 10 images, one at a time, in a 3:4 aspect ratio.
In the borderless canvas, we can see Lovart's image generation process. The dialog box on the right is working in the library, and the generated images will be laid out one by one on the canvas, making it very easy to see all the materials and find them.
If you want to change the style or modify the details, simply click on the image to add it to the dialog box for further editing.
Let's take a look at the final result:
The consistency is maintained very well, and the design of the visuals tells a story.
Prompt wordsThis project generates 81 outfit photos for the character in the reference image, grouped into sets of nine. Each set maintains consistency in the character's clothing and setting, employing a combination of wide shots, close-ups, medium shots, and extreme close-ups. Different shooting angles and character poses are randomly generated, ensuring relaxed and natural postures, realistic human anatomy, normal proportions, and actions consistent with everyday shooting logic. Lighting and shadow are natural and harmonious, matching the environment, resulting in an overall realistic style.
The generated daily outfits can be published directly~ Lovart's batch image generation is so reliable!
Business and product posters
Nano Banana 2 can generate images ranging from 512px to 4K, with different aspect ratios and resolutions.Designers should understand the importance and convenience of this feature.
In e-commerce, the same set of materials often needs to be prepared in different formats for different platforms: banners are more suitable for horizontal layouts, store main images need to be square, product detail pages require vertical images, and WeChat Moments posters also tend to be vertical. If you use a single large image to crop all of them, the composition will be very messy, and making each one separately is particularly time-consuming.
We can directly generate multiple variants from the Lovart reference image at once:
Prompt wordsGenerate 10 variations of different sizes for use on banner pages, store main images, WeChat Moments posters, and product detail pages.
Although the aspect ratios of different images have changed, the style and subject matter remain consistent.
Once we've created a great main visual, we can use it like a template.One-clickIt can be extended to different scenarios, saving a lot of time.
The development of raw image models in the past two years has been remarkable.AI The generated images are becoming clearer and more realistic. Nano Banana 2 has made a more crucial improvement, as it can directly capture real-time information and generate images based on real-time data.
Nano Banana 2 is almost three to four times faster than Nano Banana Pro, with more stable text comprehension and rendering, and it's also cheaper. In business, it's often said that speed, quality, and cost are mutually exclusive; you can't have it all. Fast comes at a price, cheap comes at a price, and good comes at a price. The advent of Nano Banana 2 directly breaks this impossible triangle.
Of course, improved design efficiency is not just about producing drawings faster, but about making the entire process smoother, clearer, and simpler.
Lovart's Nano Banana 2 is even more user-friendly, with its conversational generation and infinite canvas, allowing all images to be displayed on the same page. Reference images and any previous version can be easily added to the context, making image editing incredibly convenient and accurate. Combined with the batch generation function, a single inspiration can generate countless usable materials.
Individual design capabilities will gradually be... AI By scaling up, one person can complete the entire process from idea to delivery. For teams, collaboration becomes lighter and communication costs decrease.
What will truly differentiate us in the future may no longer be who can draw more realistically, but who can turn ideas into reality faster, more accurately, and more controllably.