Nano Banana Prompt Generator: How to Write Perfect Prompts for AI Image Generation (2026)

A Nano Banana prompt generator takes your rough idea and transforms it into a detailed, structured prompt that Google’s Nano Banana AI image models can interpret with precision. The Nano Banana prompt generator automates the process of adding style keywords, lighting instructions, composition details, and technical parameters so you get professional-grade visuals without memorizing prompt engineering syntax. Community repositories like the GitHub awesome-nano-banana-pro-prompts collection already contain over 10,000 curated prompts across 16 languages, proving how central prompt structure has become to the AI image generation workflow.

Below you will find the complete prompting formula, 50+ ready-to-use templates for every use case, and a breakdown of how Nano Banana, Nano Banana 2, and Nano Banana Pro handle prompts differently.

What Is a Nano Banana Prompt Generator?

A prompt generator for Nano Banana sits between your imagination and the AI model. You type something simple like “dog in fantasy armor” and the tool expands it into a multi-layered instruction set that specifies lighting, camera angle, material texture, color palette, and composition. The model then applies deep reasoning capabilities to parse every element before rendering a single pixel.

How Nano Banana Prompt Generators Work

Nano Banana is part of Google’s Gemini model family, originally discovered through blind testing on LMArena where users rated its outputs without knowing which model produced them. A prompt generator works by mapping your input against a database of tested style keywords, composition rules, and technical parameters. It fills in the gaps you leave — if you say “sunset portrait,” the generator adds “golden hour backlighting, shallow depth of field f/1.8, warm amber tones, medium shot, center-framed” based on what has historically produced the best results for that category.

The output is not a rigid formula. Generators adapt based on the model you target. A prompt optimized for Nano Banana (Flash) emphasizes conversational editing instructions, while a Nano Banana Pro prompt frontloads structured scene logic because the Pro model uses a reasoning engine before generating.

Why Prompts Matter for Nano Banana

Google’s own Cloud Blog confirms that Nano Banana models apply deep reasoning to fully understand a prompt before generating an image. A vague prompt like “a dog” triggers the model’s default assumptions about breed, pose, lighting, and background. A structured prompt removes ambiguity. The difference between “a dog” and “a golden retriever puppy playing in a sunlit garden, photorealistic, cinematic lighting, shallow depth of field” is the difference between a stock photo and a portfolio piece. The prompt is the single biggest factor controlling output quality, more than model version or resolution settings.

Nano Banana vs Nano Banana 2 vs Nano Banana Pro: Which Model to Prompt

Each Nano Banana model processes prompts through a fundamentally different pipeline. Choosing the wrong model for your task wastes tokens and produces suboptimal results. The three models cover a spectrum from rapid iteration to deep compositional reasoning.

Nano Banana (Gemini 2.5 Flash) — Speed and Editing

Nano Banana, built on Gemini 2.5 Flash, uses rapid pattern-matching to analyze existing pixels and predict how they should change. It excels at conversational editing where you refine an image through multiple rounds of short instructions: “make the sky more dramatic,” “remove the person on the left,” “shift the color palette to warm tones.” The prompting paradigm is iterative. You start with a base image and sculpt it through follow-up prompts rather than front-loading every detail into a single instruction.

Best use cases include inpainting (removing or replacing objects within an image), style transfer (converting a photo to watercolor, synthwave, or oil painting), and quick concept exploration where speed matters more than pixel-level precision.

Nano Banana 2 (Gemini 3.1 Flash Image) — The New Default

Nano Banana 2 closes the gap between Nano Banana and Nano Banana Pro. It delivers richer textures, sharper detail, stronger lighting, and better prompt adherence without the latency tradeoff of Pro. Supporting 131,072 input tokens, Nano Banana 2 can process complex multi-image prompts where you feed up to 14 reference images alongside your text instructions.

Character consistency is where Nano Banana 2 stands out. It can maintain the resemblance of up to 5 distinct characters and the fidelity of up to 14 objects simultaneously across generations. Configurable thinking levels let you choose between Minimal (fast exploration), High (complex layouts), and Dynamic (model decides based on prompt complexity).

Nano Banana Pro (Gemini 3 Pro Image) — Reasoning and Precision

Nano Banana Pro uses the “Deep Think” reasoning engine, which plans scene logic before generating any pixels. Where Flash pattern-matches, Pro reasons. It breaks down your prompt into spatial relationships, text placement rules, lighting physics, and compositional hierarchy before rendering. This makes it the only Nano Banana model suitable for complex text rendering, infographics, and typography-heavy designs.

Pro supports 65,536 input tokens, native 2K and 4K output, and integrates with live Google Search for real-time data grounding. It generates multilingual text in 10+ languages with high fidelity — a capability the Flash models cannot match.

Quick Comparison Table

FeatureNano Banana (Flash)Nano Banana 2Nano Banana Pro
Primary roleEditing and iterationBalanced generationReasoning and precision
Prompting styleConversational, iterativeDetailed, single-shotStructured, front-loaded
SpeedFastestFastModerate
Text renderingBasicImprovedBest (Deep Think)
Max resolution2K2K4K
Reference imagesUp to 14Up to 14Up to 14
Search groundingNoNoYes
Input tokens131,072131,07265,536
Best forInpainting, style transferGeneral image generationInfographics, typography

The Nano Banana Prompt Formula: Best Practices

Google’s official prompting documentation defines a five-part formula that works across all Nano Banana models. Mastering this formula eliminates the trial-and-error cycle most users fall into.

The Core Formula: Subject + Action + Context + Style

Every effective Nano Banana prompt contains five elements in this order:

  1. Subject — the main object or person (“a striking fashion model wearing a tailored brown dress”)
  2. Action — what the subject is doing (“posing with a confident stance”)
  3. Location/Context — the environment (“a deep cherry red studio backdrop”)
  4. Composition — camera framing (“medium-full shot, center-framed”)
  5. Style — artistic direction (“fashion magazine editorial, analog film grain, cinematic lighting”)

The formula works because it mirrors how the model’s internal reasoning decomposes a scene. Subject anchors the generation, action creates dynamism, context provides spatial grounding, composition controls the virtual camera, and style sets the rendering pipeline.

Four Golden Rules of Nano Banana Prompting

Use positive framing exclusively. Describe what you want to see, never what you want to avoid. Write “empty street at dawn” instead of “street with no cars.” Nano Banana models parse positive instructions more reliably than negations, which can produce the opposite of your intent.

Control the camera with photographic terminology. Terms like “low angle,” “aerial view,” “shallow depth of field (f/1.8),” “macro lens,” and “wide-angle” map directly to the model’s training data from professional photography. Generic instructions like “close up” produce less consistent results than “tight headshot, 85mm portrait lens, f/2.0.”

Be specific about materials and textures. “Brushed steel with fingerprint smudges” outperforms “metal surface.” “Cracked leather with visible grain” beats “leather texture.” The model generates better micro-detail when you name exact material properties.

Iterate conversationally. No prompt is perfect on the first try. Generate a base image, then refine with follow-up instructions: “increase the contrast in the shadows,” “shift the background from blue to teal,” “add rim lighting from the upper right.”

Start With a Strong Verb

The first word of your prompt tells the model its primary operation. Each verb activates a different processing mode:

  • Generate/Create — new image from scratch
  • Transform/Convert — style transfer on existing image
  • Replace — swap one element for another
  • Remove — delete an object and fill the gap
  • Add — insert a new element into an existing scene
  • Change — modify a specific property (color, size, position)
  • Translate — convert text in an image to another language

Starting with a verb instead of a noun (“Generate a sunset over mountains” vs “sunset over mountains”) produces more predictable results because it removes ambiguity about whether you want generation, editing, or analysis.

Ready-to-Use Nano Banana Prompt Templates

These templates follow the five-part formula and have been tested across Nano Banana, Nano Banana 2, and Nano Banana Pro. Copy them directly or modify the bracketed elements to match your subject.

Text-to-Image Generation Prompts

Photorealistic portrait. “Generate a photorealistic portrait of a [age] [ethnicity] [gender] with [hair description], wearing [clothing]. [Expression]. Shot on a 85mm lens at f/1.4, natural window light from the left, shallow depth of field, [background]. Color grading: [warm/cool/neutral] tones.”

Cinematic landscape. “Create a sweeping landscape of [location/terrain] during [time of day]. Dramatic [cloud type] clouds, [lighting condition]. Wide-angle lens, rule of thirds composition, [foreground element] in the lower third. Style: National Geographic photography, high dynamic range.”

Product shot. “Generate a commercial product photograph of [product] on [surface material]. Three-point softbox studio lighting, clean white background gradient, slight reflection on surface. Shot with a macro lens, f/8 for deep focus. Style: Apple product photography, minimalist.”

Concept art. “Create concept art of [subject/scene]. [Art style: Ghibli, cyberpunk, solarpunk, dark fantasy]. Dynamic composition with [foreground/midground/background elements]. Atmospheric perspective, volumetric lighting, [color palette]. Style: AAA game concept art, matte painting quality.”

Illustration. “Generate a [style: flat vector / watercolor / ink wash / children’s book] illustration of [subject]. [Color palette]. Clean lines, [composition]. Suitable for [use case: editorial, book cover, poster].”

3D Figurine and Collectible Prompts

The 3D figurine trend is one of Nano Banana’s signature use cases. These prompts produce collectible-quality renders.

“Turn this portrait photo into a 1/7 scale commercial figure in realistic style, placed on a computer desk. Transparent acrylic base with engraved name plate. Studio lighting from above, slight ambient occlusion. Material: painted PVC with matte finish. Show subtle sculpting details on clothing folds.”

“Create a gashapon capsule diorama of [character]. Chibi proportions (3-head ratio). Sitting inside a transparent capsule with [themed background]. Pastel color palette. Material: glossy painted PVC. Cute, collectible aesthetic.”

“Generate a 1/6 scale action figure of [character] in [pose]. Metallic chrome armor with weathering effects. Ball-joint articulation visible at shoulders and hips. Display base with nameplate. Studio product photography lighting.”

Style Transfer Prompts

Style transfer works best on Nano Banana (Flash) with its conversational editing pipeline. The key instruction is preserving composition while changing visual style.

“Change this image to synthwave style. Add neon grid ground plane, retro sunset gradient sky (pink to purple to dark blue), chrome reflections on all surfaces. Keep the composition and position of all objects exactly the same.”

“Transform this photograph into a watercolor painting. Visible brushstrokes, paper texture bleeding through, slight color bleeding at edges. Preserve the original composition. Style: Turner-inspired atmospheric watercolor.”

“Convert this image to Van Gogh’s Starry Night style. Thick impasto brushstrokes, swirling sky patterns, vibrant complementary colors (blue-orange, yellow-purple). Maintain the exact composition and subject positioning.”

Photo Editing Prompts

Object removal. “Remove the [element] from this image. Fill the area with natural continuation of the [surrounding context: background, floor, sky]. Keep everything else exactly the same. Match lighting and texture seamlessly.”

Object addition. “Add a [element] to the [location in image]. Match the existing lighting direction (coming from [direction]), perspective, and color temperature. The added element should cast an appropriate shadow.”

Background replacement. “Replace the background with [new background description]. Keep the subject exactly as-is, including hair edges and semi-transparent areas. Match the new background’s lighting to the subject’s existing lighting direction.”

How to Prompt for Character Consistency

Maintaining a consistent character across multiple images is one of the hardest challenges in AI image generation. Nano Banana 2 handles this better than most competitors thanks to its multi-reference architecture.

The 360-Degree Character Sheet Method

The most reliable technique for character consistency uses a two-step process.

Step 1: Generate a character reference sheet showing your character from multiple angles in a single image. Prompt: “Create a character reference sheet for [character description]. Show the character in the same outfit from four angles: front view, three-quarter left, three-quarter right, and back view. White background, consistent lighting, neutral pose. Label each view.”

Step 2: Upload that reference sheet and use it to place the character in new scenes. Prompt: “Using the attached character reference sheet, generate [character name] in [new scene/pose/environment]. Maintain exact facial features, hair style, body proportions, and clothing details from the reference. [New scene lighting and composition instructions].”

This method works because the model sees all angles simultaneously, building an internal 3D understanding of proportions, clothing details, and facial features before rendering the new scene.

Multi-Reference Image Logic

Nano Banana 2 supports up to 14 reference images per prompt, enabling sophisticated composite references. You can combine one image for facial features, another for pose, and a third for garment texture. The formula is: [Reference images] + [Relationship instruction] + [New scenario].

“Using Image 1 as the face reference, Image 2 as the body pose reference, and Image 3 as the fabric texture, generate a fashion photograph of this person in this pose wearing a garment made from this fabric. Set in a sun-drenched minimalist studio, editorial lighting.”

“Using the napkin sketch as structure and the fabric sample as texture, transform into a high-fidelity 3D armchair render in a sun-drenched minimalist living room. Photorealistic materials, ray-traced lighting.”

Prompting for Text Rendering and Typography

Text rendering separates Nano Banana Pro from every other model in the family. The Deep Think reasoning engine plans character placement, kerning, and readability before committing pixels, producing text that is actually legible at output resolution.

Rules for Sharp, Legible Text

Getting clean text output requires specific prompting techniques that differ from standard image generation.

Put exact text in double quotation marks. The model treats quoted strings as literal text to render: “Happy Birthday,” “URBAN EXPLORER,” “Sale 50% Off.” Without quotes, the model may interpret text as stylistic guidance rather than literal characters.

Keep phrases short. Single words and 2-4 word phrases render most reliably. Full sentences increase the chance of character errors. If you need longer text, break it into multiple text elements with explicit positioning: “Title text ‘EXPLORE’ centered at the top, subtitle ‘The World Awaits’ below in smaller type.”

Guide the font style descriptively rather than naming specific fonts. “Clean, bold sans-serif” or “elegant, traditional serif with thin strokes” gives the model creative latitude while maintaining your intent. Overly specific requests like “Helvetica Neue Light 14px” may not map to the model’s training data.

Define text hierarchy explicitly with three levels. Level 1: headline (large, bold). Level 2: subheaders (medium, semi-bold). Level 3: body copy (small, regular weight). This mirrors how the Deep Think engine structures typographic layouts internally.

Multilingual Text and Translation

Nano Banana Pro supports state-of-the-art multilingual text generation in 10+ languages, including Latin, Cyrillic, CJK, Arabic, and Devanagari scripts. The translation workflow preserves the original image composition while swapping text content.

“Using this image, translate the text on the storefront sign into Japanese. Keep the font size, color, and positioning exactly the same. Maintain all other visual elements unchanged.”

This works for marketing materials, restaurant menus, street signage, product labels, and any image where text needs localization without redesigning the entire visual.

Prompting for Infographics and Data Visualization

Nano Banana Pro’s Deep Think engine handles structured data and hierarchical information better than any other image model. The key is giving the model a clear data structure and layout pattern.

S-curve layout for sequential processes. “Create a professional process infographic showing ‘How to Brew the Perfect Espresso.’ Use an S-curve zigzag pattern flowing top to bottom. Include five steps, each with an icon and a short label. Warm neutral palette (cream, brown, copper). Clean sans-serif typography. 30% white space minimum.”

Bento grid for topic overviews. “Generate a Bento grid infographic comparing four AI image generators. Each cell contains: model logo placeholder, model name, 3 key stats, and a color-coded rating bar. Cool blue palette with white text. Modern, minimalist design.”

Data encoding through color. Sequential palettes (light to dark shades of one color) work for magnitude and numerical data. Qualitative palettes (distinct colors) work for categorical groups. Specify this in your prompt: “Use a sequential blue palette (light to dark) for the bar chart values, indicating increasing complexity.”

The 3-level text hierarchy (headline, subheader, body copy) applies here too. Infographic prompts that define typography levels produce cleaner, more readable output than those that leave font decisions to the model.

Advanced Prompting: Search Grounding, Lighting, and Camera Control

These techniques push beyond basic prompting into professional-grade control over the generation pipeline. They work best with Nano Banana Pro but many lighting and camera tricks apply across all models.

Real-Time Data with Google Search Grounding

Nano Banana Pro can actively search the web to generate images based on real-time information through its integration with Google Search grounding on Vertex AI. This means your prompts can reference current events, live data, or trending topics without manually providing that context.

“Search for the current weather conditions in Tokyo, then visualize them as a miniature diorama inside a snow globe. Photorealistic rendering, studio macro photography lighting.”

“Look up today’s top 3 trending topics on social media, then create an editorial magazine cover that incorporates all three topics as visual elements. Bold typography, Vogue-style layout.”

The formula for search-grounded prompts is: [Source/Search request] + [Analytical task] + [Visual translation]. The model retrieves data, processes it, and then generates an image informed by that data.

Lighting Design

Lighting control uses the same terminology professional photographers and cinematographers use.

  • Studio setups: “three-point softbox setup” produces even, commercial lighting. Add “beauty dish at 45 degrees” for portrait glamour lighting
  • Dramatic: “Chiaroscuro lighting with harsh, high contrast” creates Renaissance painting mood. “Single hard light from upper left, deep shadows” for noir
  • Natural: “Golden hour backlighting creating long shadows” for warm outdoor scenes. “Overcast diffused daylight” for even, shadowless illumination
  • Rembrandt lighting: creates a triangle of light on the shadow-side cheek, ideal for moody, editorial portraits

Camera, Lens, and Film Stock

Hardware references in prompts trigger specific visual characteristics from the model’s training data.

  • GoPro — immersive wide-angle distortion, action camera perspective
  • Fujifilm — authentic, warm color science with subtle film grain
  • Disposable camera — raw flash, red-eye artifacts, nostalgic grain, slight color shift
  • Hasselblad — medium format, extreme detail, smooth tonal gradation

Lens specifications control framing and depth. “Wide-angle 24mm” captures vast environments. “Macro lens with 1:1 reproduction ratio” reveals microscopic detail. “85mm portrait lens at f/1.8” creates creamy background bokeh.

Film stock references add temporal character. “1980s Kodak Portra 400, warm grain” for nostalgic warmth. “Fuji Velvia 50 slide film, saturated colors” for vivid landscapes. “Ilford HP5 pushed to 1600, heavy grain” for gritty black and white.

Technical Specs: Resolutions, Aspect Ratios, and Limits

Understanding the technical boundaries prevents wasted generation attempts and helps you optimize prompts for your target output format.

SpecificationNano Banana 2 (Flash)Nano Banana Pro
Max resolution2K4K
Available resolutions512px, 1K, 2K512px, 1K, 2K, 4K
Input tokens131,07265,536
Output tokens32,76832,768
Reference imagesUp to 14Up to 14
Max file size (API)50 MB50 MB
Max file size (Console)7 MB7 MB
Supported formatsPNG, JPEG, WebP, HEIC, HEIFPNG, JPEG, WebP, HEIC, HEIF
Content credentialsC2PA + SynthIDC2PA + SynthID

Resolution and Aspect Ratio Reference

Both models support standard aspect ratios: 1:1, 3:2, 2:3, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, and 21:9. Nano Banana 2 (Flash) additionally supports extreme ratios: 1:4, 4:1, 1:8, and 8:1, useful for banner ads, social media stories, and panoramic compositions.

Use 4K resolution for shots where micro-textures matter — brushed steel, fabric weave, skin pores, architectural detail. For concept exploration and rapid iteration, 1K resolution generates faster and lets you finalize composition before upscaling.

Input and Output Limits

Every image you upload as a reference consumes tokens from your input budget. A typical high-resolution reference image uses 1,000-3,000 tokens depending on complexity. With 14 reference images at maximum resolution, you can exhaust a significant portion of Flash’s 131,072 token limit. Plan your reference strategy accordingly — crop reference images to show only the relevant detail rather than uploading full-resolution photos.

All outputs from both models include C2PA Content Credentials and SynthID watermarks. These are invisible to viewers but can be detected by verification tools, ensuring transparency about AI-generated content.

The key is to start a prompt with a strong verb that tells the model the primary operation: Generate, Create, Transform, Replace, Remove, Add, Change, Convert.

Google Cloud — Ultimate Prompting Guide for Nano Banana

Max Input Tokens by Model

FAQ

keyboard_arrow_up