ComfyUI Nano Banana: Complete Setup and Workflow Guide (2026)
Google’s Gemini image generation has landed inside ComfyUI, and most creators know it by its code name: Nano Banana. The integration through official Partner Nodes means you can generate and edit images directly on the canvas without downloading model weights, running a local GPU, or leaving the node-based workflow you already use. This guide covers every step from a fresh install through advanced multi-reference generation, with real costs and settings drawn from the latest 2026 builds.
Whether you pick the standard model at $0.039 per image or Nano Banana Pro with native 4K output, the entire process runs through the Gemini API on Google’s servers. That single architectural choice changes what hardware you need, how much you pay, and which creative workflows become practical.
What Is Nano Banana and Why Use It in ComfyUI?
The name started as an internal code name at Google DeepMind and stuck after the model went viral on social media. Technically, Nano Banana maps to Gemini 2.5 Flash Image, while Nano Banana Pro maps to Gemini 3 Pro Image (technical API name: gemini-3-pro-image-preview). Google never replaced the nickname, and every major ComfyUI integration now uses it as the primary label.
ComfyUI adds a layer that the Gemini web interface cannot match. As a node-based AI workflow editor with over 1,000 community custom nodes, it lets you chain Google Gemini Image generation with local Stable Diffusion or SDXL checkpoints, add ControlNet for spatial control, batch-process dozens of variations, and save every workflow as a reusable JSON template. 72% of new ComfyUI users now choose the Desktop version, which further simplifies the initial setup.
The most significant practical advantage is hardware independence. Stable Diffusion and Flux require an NVIDIA GPU with 6 GB or more VRAM to run locally. Nano Banana runs entirely through the API, so the computational load falls on Google’s infrastructure. Any machine capable of running ComfyUI itself, even a laptop without a discrete GPU, can produce high-quality images in 3 to 15 seconds per generation.
Nano Banana vs Nano Banana 2 vs Nano Banana Pro
Three model tiers exist under the Nano Banana umbrella, and choosing the wrong one wastes either money or quality. The comparison below uses data current as of early 2026.
| Feature | Nano Banana | Nano Banana 2 | Nano Banana Pro |
|---|---|---|---|
| Underlying model | Gemini 2.5 Flash Image | Gemini 2.5 Flash Image (upgraded) | Gemini 3 Pro Image |
| Max resolution | ~1024 px | ~1024 px | Native 4K (4096 x 4096) |
| Reference images | Up to 5 | Up to 5 characters + 14 objects | Up to 14 |
| Text rendering | Limited | Improved | 94% accuracy across 10 languages |
| Character consistency | Basic | Up to 5 characters | 5-face identity preservation |
| Cost per image | $0.039 | $0.039 | $0.13 — $0.24 |
| Generation speed | 3 — 5 sec | 3 — 5 sec | 5 — 15 sec |
| Context window | Standard | Standard | 64K input / 32K output |
Nano Banana 2 closes the gap between the original model and Pro. It delivers richer textures, sharper detail, stronger lighting, and noticeably better prompt adherence without the latency or cost increase of Pro. For most creators, it replaces the original as the default choice.
Nano Banana Pro remains the right pick for three specific scenarios: production assets that need native 4K resolution, designs that require accurate text in images (94% rendering accuracy versus roughly 71% on comparable models), and complex compositions referencing up to 14 source images at once with 5-face identity preservation.
Nano Banana 2 also introduces two features absent from the other tiers. Configurable thinking levels let you choose Minimal for fast exploration or High/Dynamic for complex multi-object scenes. Search-grounded world knowledge triggers Google Image Search when the prompt references lesser-known landmarks or niche objects, reducing hallucination in educational or architectural visuals.
Cost per Image by Model Tier
How to Install Nano Banana in ComfyUI
Three installation paths exist, each suited to a different type of user. All three require a paid Gemini API key, which is covered in the next section.
- Update ComfyUI to the latest version (nightly build or stable release from October 2025 onward).
- Double-click anywhere on the canvas and search for “Google Gemini Image.”
- If the node appears, drag it onto the canvas and connect it to a Save Image node.
- If the node does not appear, switch to the nightly update channel in ComfyUI settings and restart.
- Optionally, load a pre-built template by navigating to Template and searching for “Nano Banana Pro.”
That five-step sequence covers the official Partner Nodes method, which was developed collaboratively between the ComfyUI team and Google. It receives priority support and updates alongside each ComfyUI release, making it the safest choice for most workflows.
ComfyUI-NanoBanano custom node. Created by ShmuelRonen, this alternative adds batch processing up to 4 images per request, built-in cost tracking at roughly $0.039 per generation, temperature control from 0.0 to 1.0, and multi-modal operations including generate, edit, style transfer, and object insertion with up to 5 reference images. Install it by cloning the repository into your custom_nodes folder and running pip install -r requirements.txt, then restart ComfyUI.
ComfyUI_Nano_Banana for Vertex AI. Built by ru4ls, this node supports both the standard Generative AI SDK and Vertex AI connections. Enterprise teams that need Vertex AI compliance, audit logging, or VPC-level access control should consider this option over the other two.
| Method | Best for | Batch support | Cost tracking | Vertex AI |
|---|---|---|---|---|
| Partner Nodes | Most users | No | No | No |
| ComfyUI-NanoBanano | Power users | Up to 4 | Yes | No |
| ComfyUI_Nano_Banana | Enterprise | Yes | No | Yes |
Setting Up Your Gemini API Key
Every Nano Banana node in ComfyUI requires an API key from Google AI Studio. The free tier of the Gemini API explicitly excludes image generation, so a paid account with billing enabled is mandatory before any of the workflows in this guide will function.
Getting the Key from Google AI Studio
Navigate to Google AI Studio and sign in with a Google account that has billing permissions. Open the API Keys section through the left navigation panel, click Create API Key, and select or create a project. The generated key starts with “AIza” followed by roughly 35 alphanumeric characters. Copy it immediately because you cannot view the full key again after leaving the page.
Billing and Free Credits
Add a payment method under the billing section of Google Cloud Platform. New users receive $300 in free credits, which covers approximately 2,240 standard-resolution Nano Banana Pro generations at $0.134 per image, or roughly 7,700 standard Nano Banana generations at $0.039 each. Even while free credits remain active, a payment method must be on file.
Secure Storage
Set the key as an environment variable instead of pasting it directly into every node. On macOS or Linux, add export GEMINI_API_KEY="your_key_here" to your shell profile. On Windows, use System Properties to add it as a system environment variable. Several ComfyUI nodes auto-detect this variable, which prevents accidental key exposure when sharing workflows or committing JSON templates to version control.
Your First Nano Banana Workflow: Text to Image
The fastest way to confirm your setup works is to load a pre-built template. Open the Template menu in ComfyUI, search for “Nano Banana Pro,” and click to populate the canvas. The template places a Google Gemini Image node with connected output and prompt fields, so you only need to enter your API key and a text description.
Nano Banana handles natural language prompts well. You do not need the weighted token syntax common in Stable Diffusion workflows. A prompt like “a cozy coffee shop interior with morning light streaming through large windows, watercolor painting style” produces useful results without brackets, parentheses, or CLIP weighting. Set your preferred aspect ratio (1:1, 16:9, 9:16, 4:3, or 3:4), click Queue Prompt, and wait 3 to 15 seconds for the image to appear.
Image editing uses the same node with one addition. Connect a Load Image node to the reference input of the Gemini Image node, then write an editing prompt describing the change you want: “change the wall color to deep blue” or “add falling snow and a winter atmosphere.” The model interprets instructions contextually, deciding what to preserve and what to modify without requiring a manual mask. This makes iterative creative work faster than traditional inpainting workflows, especially when you want to explore variations without rebuilding from scratch.
Nano Banana 2 will likely become the new default: fast enough to explore, smart enough to refine, and polished enough to ship.
Jo Zhang, ComfyUI Blog
Recommended Defaults and Settings
Good defaults prevent most re-renders. The values below come from testing across current 2026 ComfyUI builds and cover both pure Nano Banana workflows and hybrid pipelines that chain API-generated images with local model refinement.
Nano Banana 2 thinking levels. Set to Minimal when you are brainstorming and generating rapid variations. Switch to High or Dynamic when the scene involves complex layouts, multi-object compositions, or detailed textual instructions. The difference in generation time is small, but High mode produces more coherent multi-element scenes.
Seed discipline. Lock the seed as soon as you find a direction worth exploring. Then iterate on the prompt with the seed fixed so you can see exactly what each word change does. ComfyUI embeds the full workflow graph into PNG metadata by default, so every image you save is fully reproducible.
Hybrid workflow sampler settings. When passing Nano Banana output into local Stable Diffusion or SDXL refinement nodes, use DPM++ 2M Karras as the sampler. Set steps to 18-24 for SD1.5 or 22-28 for SDXL. Keep CFG between 4.5 and 6.5. A short, stable negative prompt works best: “blurry, extra fingers, overlapping limbs, watermark, low-res, jpeg artifacts.”
Batch size. Set batch to 2-4 images when exploring concepts. If VRAM is tight on local model nodes, use batch count instead of batch size to avoid memory spikes on the GPU.
Pricing and Cost Optimization
The per-image cost depends entirely on which model tier you select and what resolution you need. There are no hidden fees for batch generation or aspect ratio changes.
- Nano Banana and Nano Banana 2: $0.039 per image at up to ~1024 px
- Nano Banana Pro at standard resolution: $0.13 per image
- Nano Banana Pro at native 4K: $0.24 per image
- Third-party platforms using credit systems: 4 credits per image on average
Google Cloud’s $300 free credit for new accounts stretches further than you might expect. At the standard Nano Banana rate, 7,700 images cost nothing. At the Pro rate, you still get 2,240 generations before a single charge hits your card.
The simplest cost strategy: use Nano Banana 2 for all exploration, drafting, and concept work. Switch to Pro only for the final render when you need 4K resolution or precise text. If you use the ComfyUI-NanoBanano custom node, its built-in cost tracker shows cumulative spending directly on the canvas so you never lose visibility.
Upscaling adds another option. Generate at standard resolution through Nano Banana 2, then upscale to 2K or 4K using a latent upscale node (1.5x or 2x before VAE Decode) or a post-decode image resize with Lanczos filtering. This hybrid approach produces near-Pro quality at the standard-tier price.
Troubleshooting Common Errors
Most issues fall into four categories, and all of them resolve in under a minute once you know where to look.
API Key and Authentication Errors
The single most common failure is attempting to use a free-tier API key. Image generation requires a paid Gemini API account with billing enabled. If you see authentication errors, open Google AI Studio, confirm your key has image generation permissions, and test a simple generation directly in the Studio interface before returning to ComfyUI. Also check for invisible leading or trailing spaces in the key field by selecting all text in the input and re-pasting.
Node Not Found in Search
If searching the canvas for “Google Gemini Image” returns no results, your ComfyUI version predates Partner Node support. Update to the nightly channel or any stable release from October 2025 onward. For custom nodes like ComfyUI-NanoBanano, confirm that the git clone completed without errors, that pip install -r requirements.txt ran successfully, and that you restarted ComfyUI after installation.
Generation Timeouts or Failures
Nano Banana is entirely API-based, so a stable internet connection matters more than local hardware. If generations fail intermittently, check your network. CUDA out-of-memory errors only apply when you combine Nano Banana output with local model processing in a hybrid workflow. In that case, lower batch size first, then reduce resolution. Local model nodes also require image dimensions divisible by 64.
Model and CLIP Mismatch in Hybrid Workflows
When feeding Nano Banana output into local Stable Diffusion refinement, make sure the checkpoint, CLIP encoder, and VAE all belong to the same model family. The easiest guard is to use a single Checkpoint Loader node that loads all three components together, eliminating cross-wiring between mismatched model components.
Advanced Features
Once your basic workflow runs smoothly, these capabilities push Nano Banana into production-grade territory.
Subject and character consistency lets you maintain recognizable faces and objects across multiple generations. Nano Banana 2 preserves up to 5 character identities and 14 object appearances in a single workflow, making it practical for storyboarding and narrative sequences. Nano Banana Pro extends this to 5-face identity preservation with even higher fidelity, suitable for commercial character design where likeness must remain stable across dozens of outputs.
Text rendering separates Nano Banana Pro from nearly every competing image model. At 94% accuracy in benchmark tests compared to roughly 71% on comparable systems, it handles product mockups with readable labels, social media graphics with clean typography, technical diagrams with legible annotations, and greeting cards in any of 10 supported languages including Latin scripts and CJK character sets. If your workflow needs text in images, Pro is the only tier worth considering.
Search-grounded world knowledge is exclusive to Nano Banana 2. When a prompt references a specific architectural style, a lesser-known landmark, or a niche biological specimen, the model triggers Google Image Search grounding to inform its output. This reduces the hallucinated details that plague isolated models and produces representations that reflect actual visual characteristics, particularly useful for educational infographics, travel content, and architectural visualization.
Hybrid workflows with local models unlock the full potential of ComfyUI’s node graph. Use Nano Banana for initial composition and concept generation, then pass the output through local SDXL or Flux nodes for style refinement. Add ControlNet nodes for precise spatial control over poses, edges, or depth. For the upscale step, a 1.5x to 2x latent upscale before VAE Decode preserves coherence better than raw pixel scaling, though a post-decode Lanczos resize works well for hitting exact output dimensions.
