Nano Banana 2.1 Advanced (Google)
Nano Banana 2.1 Advanced by Google generates and edits images with up to 14 references, optional Web or Image Search grounding, and 1K to 4K output.
Overview
Nano Banana 2.1 Advanced is Google’s multimodal image generation and image editing model in the Gemini family. You describe a scene or an edit in natural language, and the model outputs a finished still image at 1K, 2K, or 4K. You can also feed it reference images to keep the same character or product consistent across iterations. It’s useful when you need clean text inside images, strong prompt adherence, and repeatable creative direction across a set of assets.
What you can build
- Product photos that keep packaging text and label layout intact while changing the setting
- Multi-image composites that merge several reference items into one new scene
- Character-consistent scenes for storyboards, campaigns, or social content variations
- Infographics, posters, menus, and slides with legible typography and clear layout
- Photo edits like relighting, background swaps, outfit changes, and style changes
- Mask-style edits by marking a region in the input image and describing the change
- Sketch-to-render concept images for architecture, interiors, and industrial design
Inputs
- A written prompt that describes the image you want, or the edit you want applied to a reference image. Very long prompts are supported, but dense instructions work best when you keep them structured.
- Optional reference images to edit or to use as visual anchors. You can supply up to 14 images in one request, mixing object references and character references.
- A target output resolution choice: 1K, 2K, or 4K.
- An aspect ratio choice. When you provide a reference image, the model can match that input ratio automatically. Without a reference image, it typically defaults to a square (1:1) image unless you set a ratio.
- Optional Web Search grounding. Turn it on when the image must reflect current real-world facts.
- Optional Image Search grounding. Turn it on when you want the model to look at web images for visual context. This mode also uses Web Search.
- Optional system-level instructions that steer style and behavior for the whole request, separate from the creative prompt.
- Optional video context as an MP4 file (kept under 15 MB). Use this when motion context or a sequence helps guide the image.
- Optional document context as a PDF file (kept under 15 MB). Use this when you want the image to follow a reference spec, brief, or slide.
- Optional random seed for more repeatable outputs when you rerun the same request.
- An output file format choice for the generated image, such as JPEG, PNG, or WebP.
- A safety filter level that controls how aggressively the system blocks sensitive content.
Outputs
- One or more generated image files in the selected format (JPEG, PNG, or WebP), at the requested aspect ratio and resolution.
- Each returned image includes file metadata such as a downloadable location, and may include width and height in pixels.
- A short text description of what the model produced, which can help with cataloging or selecting between variants.
Common pixel sizes you can expect (examples): - 1:1 at 1K is 1024×1024, at 2K is 2048×2048, and at 4K is 4096×4096
- 16:9 at 1K is 1376×768, at 2K is 2752×1536, and at 4K is 5504×3072
- 9:16 at 1K is 768×1376, at 2K is 1536×2752, and at 4K is 3072×5504
- 8:1 at 1K is 3072×384, at 2K is 6144×768, and at 4K is 12288×1536
Recommended settings
- Infographics and dense layouts: use 2K or 4K, and a higher thinking level if available.
- Small text (labels, captions, packaging): prefer 2K over 1K. 1K can blur small typography.
- Character consistency across multiple images: provide a clean, front-facing reference portrait for each character, and keep character count to 4 or fewer.
- Product consistency: provide well-lit product shots as object references, and keep object references to 10 or fewer.
- Real-world facts (weather, event results, current products): enable Web Search grounding. Add Image Search when the visual look must match what’s online.
Limitations
- The model can still miss fine typography in very small text, long paragraphs, or page-length documents rendered into a single image.
- Character consistency across edits is strong but not perfect, especially with large pose changes or occlusions.
- During edits, it may keep more of the original pose or structure than you asked for.
- It can confuse spatial details like left versus right in complex scenes.
- Like other foundation models, it can hallucinate details. Don’t treat generated diagrams or data graphics as verified truth.
- Search grounding has constraints. Some deployments don’t allow using real-world photos of people pulled from web search.
- Non-structured or low-quality inputs reduce output quality. Blurry photos, heavy compression, screenshots of screenshots, and messy markups can cause unwanted artifacts.
Safety & compliance
- Safety filtering can block requests or outputs that violate policy. Use the safety level controls when your use case needs stricter or looser filtering.
- Generated images may include provenance features such as SynthID watermarking and content credentials, depending on the deployment.
- You’re responsible for rights and consent for any reference images you upload, including likeness rights, trademarks, and copyrighted material.
Example prompts
Great starting points for nano banana 2.1 advanced.
API quick start
Run nano banana 2.1 advanced with a single API call.
{
"prompt": "The same surfer standing on the beach at …",
"resolution": "1K",
"inputImage": "https://your-cdn.com/input.png",
"thinkingLevel": "medium"
}curl -X POST "https://api.wiro.ai/v1/Run/google/nano-banana-2-1-advanced" \
-H "Content-Type: application/json" \
-H "x-api-key: YOUR_WIRO_API_KEY" \
--data-binary @- <<'JSON'
{
"prompt": "The same surfer standing on the beach at …",
"resolution": "1K",
"inputImage": "https://your-cdn.com/input.png",
"thinkingLevel": "medium"
}
JSON