How to Use Nano Banana Pro: Complete Guide to Google’s AI Image Model
Nano Banana Pro (Gemini 3 Pro Image) is Google DeepMind’s most capable image generation and editing model. Released on November 20, 2025, it builds on the original Nano Banana with 4K resolution output, accurate text rendering in multiple languages, multi-image composition with up to 14 reference inputs, and Google Search grounding for data-driven infographics.
This guide covers exactly how to access Nano Banana Pro, create your first image, write effective prompts, use advanced features like reference images and style transfer, and choose the right access method for your workflow.
Where to Access Nano Banana Pro
Google has integrated Nano Banana Pro across six platforms. The fastest path for most users is the Gemini App — no code, no setup, just open and generate. Developers and enterprise users have API access through Google AI Studio and Vertex AI.
| Platform | Access Method | Best For |
|---|---|---|
| Gemini App | Create Images, choose Thinking model | Casual and professional users |
| AI Mode in Search | google.com/ai, select “Thinking with 3 Pro” | Quick generation from Search |
| NotebookLM | Slide Decks, Infographics | Research visualization |
| Google Workspace | Slides (Help me visualize), Vids | Presentations, video content |
| Flow | AI filmmaking tool, paid plans | Creatives, filmmakers |
| AI Studio / Vertex AI | Gemini API with billing | Developers, enterprise |
Free-tier users can generate images through the Gemini App with a limit of approximately 2-3 images per day at 1K resolution. When the daily quota runs out, the app falls back to standard Nano Banana. Google AI Plus, Pro, and Ultra subscribers receive higher quotas — up to 1,000 images per day at native 4K on the Ultra plan.
How to Create Your First Image: Step by Step
The entire process takes under a minute once you know the workflow. Here are the five steps from opening the app to downloading a finished image.
- Open the Gemini App at gemini.google.com and select “Create Images” from the menu, then choose the “Thinking” model — this activates Nano Banana Pro
- Write a descriptive prompt that covers subject, action, environment, composition, and style — for example: “A cinematic poster featuring a surfer dog under neon lights, sci-fi color grading, and fog at sunset”
- Optionally upload reference images (up to 14) to guide style, character appearance, or object placement — 6 for objects, 5 for faces, 3 for style references
- Set your output preferences: resolution (1K, 2K, or 4K) and aspect ratio (1:1 for Instagram, 16:9 for thumbnails, 9:16 for Reels, or any of the 10+ supported ratios)
- Click Generate, review the results, and refine with follow-up prompts — adjust lighting, replace specific areas, apply style transfer, or change composition through conversational editing
Nano Banana Pro processes your prompt through an internal reasoning step — it forms “intermediate mental images” to plan layout and lighting before rendering the final output. This thinking process is why the model produces more coherent results than standard image generators, especially for complex scenes with multiple elements.
Prompting Guide: How to Write Effective Prompts
The quality gap between a vague prompt and a specific one is dramatic. Google’s official prompting formula for Nano Banana Pro follows five components, and mastering this structure produces consistently better results.
The Prompt Formula
Every effective Nano Banana Pro prompt combines five elements: Subject + Action + Location/Context + Composition + Style. Instead of keywords, describe the scene narratively — the model responds to specific photographic language.
A strong prompt reads like a creative brief, not a search query. Compare these two approaches:
| Weak Prompt | Strong Prompt |
|---|---|
| “Fashion model in studio” | “A striking fashion model wearing a tailored brown dress, sleek boots, and holding a structured handbag. Posing with a confident stance, slightly turned. Deep cherry red studio backdrop. Medium-full shot, center-framed. Fashion magazine editorial, shot on medium-format analog film, pronounced grain, cinematic lighting.” |
The strong prompt specifies clothing details, pose, background color, framing, and photographic style — giving the model enough information to produce a precise result rather than a generic interpretation.
Prompting Best Practices
Be specific and use positive framing. Describe what you want, not what you do not want. Say “empty street” instead of “street with no cars.” Nano Banana Pro interprets positive instructions more reliably than negations.
Control the camera with photographic terms. Terms like “low angle,” “aerial view,” “shallow depth of field (f/1.8),” “wide-angle lens,” and “golden hour lighting” give you precise control over perspective and mood. The model understands these terms the same way a photographer would.
Iterate conversationally. After generating an initial image, refine it through follow-up prompts: “Make the background warmer,” “Add fog in the distance,” “Change the jacket to leather.” This conversational editing is one of Nano Banana Pro’s strongest capabilities — each refinement builds on the previous result rather than starting from scratch.
Start with a strong verb. The first word of your prompt should tell the model the primary operation: “Create,” “Generate,” “Design,” “Transform,” “Edit,” “Combine.” This anchors the model’s reasoning process from the start.
Advanced Features: Reference Images, Text Rendering, and 4K
Beyond basic text-to-image generation, Nano Banana Pro includes several professional-grade capabilities that separate it from standard AI image generators.
Multi-Image Reference (Up to 14 Inputs)
You can upload up to 14 reference images in a single prompt to guide the output. The model maintains character consistency across references — the same person appears with identical features in different scenes, lighting conditions, and outfits. Supported formats include PNG, JPEG, WebP, HEIC, and HEIF, with a maximum of 6 object references, 5 face references, and 3 style references per generation.
This feature enables workflows that were previously impossible without manual compositing: placing a specific product into a completely new environment, maintaining a brand mascot’s appearance across a campaign, or generating consistent storyboard frames where characters look the same in every panel.
Text Rendering in Multiple Languages
Nano Banana Pro renders legible text directly in generated images — from short taglines to full paragraphs, in multiple languages. This capability makes it practical for creating posters, mockups, social media graphics, infographics, and product packaging concepts without a separate text overlay step. The model handles calligraphy, block letters, and stylized fonts, and can translate text between languages while preserving the visual design.
4K Resolution and Aspect Ratios
Output resolution scales from 1K through 2K to 4K (available on Ultra subscription and API). The model supports over 10 aspect ratios natively: 1:1, 3:2, 2:3, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9, and Nano Banana 2 adds 1:4, 4:1, 1:8, and 8:1. No cropping or resizing needed — choose your target format before generation.
Google Search Grounding
Nano Banana Pro connects to Google Search to pull real-time data into generated visuals. Request an infographic about current weather, a recipe visualization with real ingredients, or a data chart based on live information — the model incorporates factual data from the web into the image.
Nano Banana Pro uses Gemini’s state-of-the-art reasoning and real-world knowledge to visualize information better than ever before.
Google — Introducing Nano Banana Pro
Nano Banana Pro vs Nano Banana 2: Which to Use
Both models serve different use cases, and choosing the right one depends on whether you prioritize visual fidelity or cost efficiency.
API Cost Per Image: Nano Banana Pro vs NB2 ($)
| Feature | Nano Banana Pro | Nano Banana 2 |
|---|---|---|
| Model base | Gemini 3 Pro | Gemini 3.1 Flash |
| Image quality | Highest fidelity | Near-Pro quality |
| API cost (2K) | $0.134/image | $0.101/image |
| API cost (4K) | $0.24/image | $0.151/image |
| Context window | 65,536 input tokens | 131,072 input tokens |
| Speed | Slower (more reasoning) | Faster |
| Best for | Premium creative work, complex compositions | High-volume production, cost-sensitive workflows |
For professional creative work requiring maximum fidelity — product photography, brand campaigns, complex multi-element compositions — Nano Banana Pro remains the superior choice. For high-volume generation where cost matters and near-Pro quality suffices, Nano Banana 2 delivers 25-50% savings at comparable visual quality.
Common Use Cases
- Infographics and data visualizations with real-world data via Search grounding
- Product mockups and packaging concepts with accurate text rendering
- Social media content across all major platform formats
- Brand campaigns with consistent characters across multiple scenes
- Storyboards and visual narratives with multi-image reference
- Educational materials and diagrams from notes or documents
- Marketing assets with localized text in multiple languages
- Professional headshots and portrait editing
All generated images include C2PA Content Credentials and SynthID watermarks for provenance tracking and AI content identification.
