How to Use Nano Banana: Complete Guide to Google’s AI Image Generator

Google’s AI image generator lets you create and edit images by typing a description in plain English. Nano Banana runs inside the Gemini app — open it, describe what you want, and the model produces an image in seconds. Since launching in August 2025, over 500 million images have been edited through the Gemini app alone, with hundreds of millions more across other Google surfaces.

This guide walks through every way to use the tool: generating images from scratch, editing existing photos, rendering text, maintaining character consistency across scenes, and choosing the right model for your specific task.

Where to Access Nano Banana

Google built Nano Banana into nearly every product it operates. The fastest way to start is the Gemini app, but you have at least seven other entry points depending on whether you need consumer features, creative tools, or developer APIs.

Gemini App (Easiest Way)

Visit gemini.google.com or open the Gemini mobile app. Select “Create images” from the tools menu, or simply describe what you want in the chat. Free-tier users get a limited number of image generations per day. After hitting the limit, the app falls back to the original Nano Banana model. Google AI Plus, Pro, and Ultra subscribers receive higher quotas and keep access to Nano Banana Pro through the regeneration menu.

You can choose between Fast, Thinking, and Pro model modes from the dropdown. Nano Banana 2 is the default image generator across all three modes. Thinking mode adds extra reasoning before generation, which helps with complex scenes but takes longer.

Other Platforms

Google’s image generation model extends well beyond the main Gemini app. Each platform serves a different workflow.

  • AI Mode in Search — go to google.com/ai, select “Thinking with 3 Pro” from the dropdown, click the plus sign icon, and choose “Create Images Pro”
  • NotebookLM — turn research sources into visual Slide Decks and Infographics powered by Nano Banana Pro
  • Google Slides and Vids — select “Help me visualize” in the Gemini sidebar to generate infographics, charts, and slide designs
  • Flow — Google’s AI filmmaking tool uses Nano Banana 2 as the default image model, available for zero credits
  • AI Studio — free for experimentation at aistudio.google.com, requires a paid API key for production use
  • Vertex AI — enterprise deployment through Google Cloud
  • Google Ads — powers AI-generated image suggestions for ad campaigns

How to Generate Images: Step by Step

Every image starts with a prompt. The quality of your output depends almost entirely on how well you describe what you want. Google’s AI image generator uses deep reasoning to fully understand your prompt before producing any pixels, which means detailed instructions consistently outperform vague ones.

Here is the process from start to finish:

  1. Open the Gemini app and select your preferred model mode (Fast for speed, Thinking for complex scenes, Pro for maximum quality)
  2. Type your prompt using the formula: Subject + Action + Location/Context + Composition + Style
  3. Wait a few seconds for the image to generate
  4. Review the result and refine with follow-up prompts — the model maintains conversation context, so you can say “make the background darker” or “change the dress to blue”
  5. Download the final image or continue iterating until satisfied

The Prompt Formula

Google Cloud’s official prompting guide breaks effective prompts into five components. The Subject defines who or what appears in the image. The Action describes what the subject is doing. Location and Context set the environment, lighting, and time of day. Composition controls camera angle and framing. Style determines the visual aesthetic.

A strong prompt looks like this: “A striking fashion model wearing a tailored brown dress, posing with a confident stance, in a deep cherry red studio, medium-full shot center-framed, fashion magazine editorial style, shot on medium-format analog film with pronounced grain.”

Compare that to a weak prompt: “a woman in a dress.” Both produce results, but the detailed version gives you control over the output instead of leaving everything to the model’s defaults.

Best Practices for Prompts

Start with a strong verb that tells the model the primary operation you want to perform. “Generate,” “Create,” “Transform,” “Edit,” or “Draw” each set a different expectation for the output.

Use positive framing. Describe what you want to see, not what you want to avoid. Write “empty street at dawn” instead of “a street with no cars and no people.” The model responds better to constructive descriptions because negative constraints are harder to enforce reliably.

Control the camera with photographic terminology. Terms like “low angle,” “aerial view,” “close-up,” “wide shot,” “shallow depth of field,” and “golden hour lighting” give the model specific visual parameters to work with.

Iterate conversationally. After the initial generation, refine the image through follow-up prompts. You do not need to rewrite the entire description each time. Say “move the lamp to the left” or “make the colors warmer” and the model adjusts the existing image.

Prompt ComponentWhat It ControlsExample
SubjectWho/what appears“a golden retriever puppy”
ActionWhat they’re doing“jumping to catch a frisbee”
Location/ContextWhere, when, lighting“in a sunlit park at golden hour”
CompositionCamera angle, framing“low-angle shot, rule of thirds”
StyleVisual aesthetic“photorealistic, shallow depth of field”

How to Edit Images with Nano Banana

Image editing is where this Gemini image model truly separates itself from competitors. Instead of manually painting masks or selecting areas with cursor tools, you describe the change you want in natural language. The model uses semantic masking to identify the target element from your description and applies the edit while keeping the rest of the image intact.

Upload any image to the Gemini app and start giving instructions. The editing capabilities fall into four main categories.

Object removal. Use the verb “remove” and name the specific element: “Using this image, remove the stone bust from the foreground. Keep everything else exactly the same, preserving the original style, lighting, and composition.” The explicit instruction to preserve everything else prevents the model from making unwanted changes to surrounding areas.

Object addition. Define both the new element and its placement: “Using this image, add a regal Doberman Pinscher sitting on the gravel path to the far left. Ensure the new object matches the lighting and perspective of the original image.” Describing how the added element should interact with existing lighting prevents it from looking pasted in.

Object replacement. Combine removal and addition in a single step: “Replace the stone bench with a sculptural, liquid-mercury-like flowing metal wave. Keep everything else exactly the same.” This produces cleaner results than doing removal and addition separately because the model handles the spatial transition in one pass.

Style transfer. Repaint an entire image in a different artistic style while preserving layout and subject matter: “Change this image to a sleek, minimalist futurism style with all organic textures replaced by smooth white polymer and chrome. Ensure the composition and position of all objects remain exactly the same.” This works for watercolor, oil painting, pixel art, 3D render, anime, and dozens of other aesthetics.

Text Rendering and Localization

Generating accurate, legible text inside images was one of the hardest problems in AI image generation until Nano Banana Pro solved it. The current models use a character-by-character verification loop that plans the text layout, evaluates each character after rendering, and corrects errors before finalizing the output.

To get the best typographic results, enclose your desired words in quotes within the prompt. Write “Happy Birthday” rather than just Happy Birthday. Then describe the typography style: “bold sans-serif font,” “neon cursive signage,” “hand-lettered chalk on blackboard,” or name a specific font family.

Nano Banana 2 also supports in-image localization. You can write a prompt in one language and specify the target language for rendered text. Provide the exact translated text in quote marks to guide the rendering, or let the model translate automatically based on cultural cues in your prompt. This makes it possible to produce localized marketing materials, greeting cards, and signage across multiple languages from a single workflow.

Text TaskPrompt ApproachBest Model
Short text (1-5 words)Enclose in quotes, specify font styleAny model
Paragraph textEnclose in quotes, specify layout and fontNano Banana Pro
Multilingual textProvide exact translated text in quotesNB2 or Pro
Infographic labelsDescribe data + layout, model generates textNB2 or Pro
Calligraphy / artisticDescribe letterform style in detailNano Banana Pro

Character Consistency

Maintaining a consistent character across different images remains one of the most difficult challenges in AI generation. If you simply describe a character with the same adjectives in multiple prompts, the model produces a slightly different person each time because it interprets the description from scratch.

The 360-Degree Character Sheet Technique

The most reliable method uses reference images instead of text descriptions. Generate a character sheet showing your character from multiple angles — front view, side profile, three-quarter view, and back view — all in a single image. Use a prompt like: “A character sheet showing a young woman with short red hair and a green jacket from four angles: front, left profile, three-quarter, and back view. White background, consistent lighting.”

Once you have this sheet, upload it as a reference image in every subsequent prompt. The model reads the visual reference and maintains facial features, clothing, proportions, and pose characteristics across completely different scenes and compositions.

Nano Banana 2 supports up to 5 characters and the fidelity of up to 14 objects in a single workflow. Assign a distinct name to each character in your prompt so the model can track them independently. For example: “Using the reference sheet, show Character A (the woman in green) sitting across from Character B (the man in blue) at a coffee shop table.”

Nano Banana vs. Nano Banana Pro vs. Nano Banana 2

Google currently offers three distinct image generation models, each designed for different workflows and budgets. Choosing the right one saves both time and money.

The original Nano Banana runs on Gemini 2.5 Flash. It is the fastest and cheapest option, designed for high-speed editing, style transfers, and rapid iteration. The model uses pattern-matching rather than deep reasoning, which means it excels at tasks where you want to modify an existing image while keeping its overall structure.

Nano Banana Pro runs on Gemini 3 Pro. It delivers the highest visual quality and uses a reasoning engine that plans scene logic before generating pixels. Pro connects to Google Search for real-time data, making it the best choice for infographics, data visualizations, and images that require factual accuracy. Text rendering is most reliable with Pro.

Nano Banana 2 runs on Gemini 3.1 Flash. It offers approximately 95% of Pro’s capabilities at roughly half the cost. The model supports 4K resolution (Pro maxes at 2K), Image Grounding through web search, and 14 aspect ratios including extreme formats like 1:8 and 8:1. For most new projects, Nano Banana 2 should be the default choice.

SpecNano BananaNano Banana ProNano Banana 2
EngineGemini 2.5 FlashGemini 3 ProGemini 3.1 Flash
Max Resolution1K2K4K
Input TokensN/A65,536131,072
Aspect RatiosStandard (10)Standard (10)Extended (14)
Reference ImagesLimitedUp to 14Up to 14
Text RenderingBasicExcellentVery Good
Image GroundingNoText onlyVisual + Text
Real-time DataNoYes (Google Search)Yes (Google Search)
Thinking ModeNoNoYes (toggleable)
Best ForQuick edits, style transferComplex compositions, infographicsGeneral purpose, production

Maximum Resolution by Model (pixels)

Real-World Use Cases

The Gemini image model has moved beyond creative experiments into production workflows. Companies are integrating the API into applications that serve millions of users, and the results show how versatile image generation and editing have become.

By integrating Nano Banana 2’s advanced image generation and editing capabilities, Whering has been able to transform low-quality user photos into professional, studio-grade assets while preserving authentic textures.

Bianca Rangecroft, CEO, Whering

Interior design and product visualization. Upload a photo of any room and drag in images of furniture, artwork, or decor. The model fuses the product into the scene with correct lighting, perspective, and shadows. Google built a demo app called Home Canvas in AI Studio that shows this workflow in action.

Virtual try-on for clothing and accessories. Take a photo of yourself and an image of a clothing item, then let the model combine them. The spatial understanding capabilities produce results where the clothing wraps naturally around the body with realistic folds and draping.

Video generation with consistent characters. One of the biggest challenges in AI video is maintaining character consistency across 8-second clips. Use Nano Banana to generate a reference frame with the exact character pose and expression you need, then pass that frame to Veo 3 for video generation. This eliminates the subtle face-shifting that breaks longer-form content.

Localized marketing at scale. Generate an ad creative in one language, then use in-image localization to produce variants for every target market. The model translates text and adapts visual elements to match cultural context, turning a single prompt into dozens of production-ready assets.

Performance gains in production. HubX achieved a 74-76% reduction in latency by switching to Nano Banana 2, effectively making their face editing workflows 4x faster without losing Pro-level quality. KLIPY uses the precision text rendering to create accurate copy directly in meme-style assets, stickers, and emojis at high speed.

FAQ

keyboard_arrow_up