Try MiniMax FastH3 V2 from Fastvideo →
Models
Agents
Workflows
Studio
PricingBlogDocs
ExploreDiscover models by categoryBrowse All ModelsBrowse the complete catalogSee FavoritesSign in to view saved models
Generative Media AgentCreate and edit media by chattingWorkflow AgentBuild visual workflows with Agent
OverviewThe platform at a glanceLearnSkills, knowledge, guardrailsAnatomyWhat makes agents reasonBuild Your AgentPick skills, set tier, deploy
Pre-built AgentsBrowse the catalog
Agent Usecases
Ad Campaign ManagerApp Event ManagerApp Review RepliesBarber BookingCustomer Win-BackEcommerce ListingsRestaurant Reviews
Sign InStart Building

Task History

Click to see output list

No tasks yet

Go to Models
Explore models/
Image GenerationActive

Nano Banana 2.1 Advanced (Google)

bygoogle

Nano Banana 2.1 Advanced by Google generates and edits images with up to 14 references, optional Web or Image Search grounding, and 1K to 4K output.

Text to ImageImage to ImageImage EditingFast Inference
Model ID
nano banana 2.1 advanced
Provider
google
Updated
1791549518
nano banana 2.1 advanced
63
Comments
Average rating : 4.92 (65 users)
Providergoogle
Modelnano banana 2.1 advanced
Text to ImageImage to ImageImage EditingFast Inference
Support Wiro on G2
wiro playground—google/nano banana 2.1 advanced
Reset to defaults
0 / 14
Maximum 14 image allowed
Delete All

Optional. Default is empty (text-to-image). Nano Banana 2.1 lets you mix up to 14 reference images: up to 10 images of objects with high-fidelity to include in the final image, and up to 4 images of characters to maintain character consistency.

Required. The prompt to generate the image.

Required. Default is 1K. Resolution of the output image.

Optional. Default is medium. How much the model reasons before generating. Higher levels can improve complex compositions but use more thinking tokens.

Sample outputs
Sample 1
Sample 2
Sample 3
Sample 4
Updated 1791549518

Overview

Nano Banana 2.1 Advanced is Google’s multimodal image generation and image editing model in the Gemini family. You describe a scene or an edit in natural language, and the model outputs a finished still image at 1K, 2K, or 4K. You can also feed it reference images to keep the same character or product consistent across iterations. It’s useful when you need clean text inside images, strong prompt adherence, and repeatable creative direction across a set of assets.

What you can build

  • Product photos that keep packaging text and label layout intact while changing the setting
  • Multi-image composites that merge several reference items into one new scene
  • Character-consistent scenes for storyboards, campaigns, or social content variations
  • Infographics, posters, menus, and slides with legible typography and clear layout
  • Photo edits like relighting, background swaps, outfit changes, and style changes
  • Mask-style edits by marking a region in the input image and describing the change
  • Sketch-to-render concept images for architecture, interiors, and industrial design

Inputs

  • A written prompt that describes the image you want, or the edit you want applied to a reference image. Very long prompts are supported, but dense instructions work best when you keep them structured.
  • Optional reference images to edit or to use as visual anchors. You can supply up to 14 images in one request, mixing object references and character references.
  • A target output resolution choice: 1K, 2K, or 4K.
  • An aspect ratio choice. When you provide a reference image, the model can match that input ratio automatically. Without a reference image, it typically defaults to a square (1:1) image unless you set a ratio.
  • Optional Web Search grounding. Turn it on when the image must reflect current real-world facts.
  • Optional Image Search grounding. Turn it on when you want the model to look at web images for visual context. This mode also uses Web Search.
  • Optional system-level instructions that steer style and behavior for the whole request, separate from the creative prompt.
  • Optional video context as an MP4 file (kept under 15 MB). Use this when motion context or a sequence helps guide the image.
  • Optional document context as a PDF file (kept under 15 MB). Use this when you want the image to follow a reference spec, brief, or slide.
  • Optional random seed for more repeatable outputs when you rerun the same request.
  • An output file format choice for the generated image, such as JPEG, PNG, or WebP.
  • A safety filter level that controls how aggressively the system blocks sensitive content.

Outputs

  • One or more generated image files in the selected format (JPEG, PNG, or WebP), at the requested aspect ratio and resolution.
  • Each returned image includes file metadata such as a downloadable location, and may include width and height in pixels.
  • A short text description of what the model produced, which can help with cataloging or selecting between variants.

Common pixel sizes you can expect (examples): - 1:1 at 1K is 1024×1024, at 2K is 2048×2048, and at 4K is 4096×4096

  • 16:9 at 1K is 1376×768, at 2K is 2752×1536, and at 4K is 5504×3072
  • 9:16 at 1K is 768×1376, at 2K is 1536×2752, and at 4K is 3072×5504
  • 8:1 at 1K is 3072×384, at 2K is 6144×768, and at 4K is 12288×1536

Recommended settings

  • Infographics and dense layouts: use 2K or 4K, and a higher thinking level if available.
  • Small text (labels, captions, packaging): prefer 2K over 1K. 1K can blur small typography.
  • Character consistency across multiple images: provide a clean, front-facing reference portrait for each character, and keep character count to 4 or fewer.
  • Product consistency: provide well-lit product shots as object references, and keep object references to 10 or fewer.
  • Real-world facts (weather, event results, current products): enable Web Search grounding. Add Image Search when the visual look must match what’s online.

Limitations

  • The model can still miss fine typography in very small text, long paragraphs, or page-length documents rendered into a single image.
  • Character consistency across edits is strong but not perfect, especially with large pose changes or occlusions.
  • During edits, it may keep more of the original pose or structure than you asked for.
  • It can confuse spatial details like left versus right in complex scenes.
  • Like other foundation models, it can hallucinate details. Don’t treat generated diagrams or data graphics as verified truth.
  • Search grounding has constraints. Some deployments don’t allow using real-world photos of people pulled from web search.
  • Non-structured or low-quality inputs reduce output quality. Blurry photos, heavy compression, screenshots of screenshots, and messy markups can cause unwanted artifacts.

Safety & compliance

  • Safety filtering can block requests or outputs that violate policy. Use the safety level controls when your use case needs stricter or looser filtering.
  • Generated images may include provenance features such as SynthID watermarking and content credentials, depending on the deployment.
  • You’re responsible for rights and consent for any reference images you upload, including likeness rights, trademarks, and copyrighted material.

Example prompts

Great starting points for nano banana 2.1 advanced.

The same surfer standing on the beach at a professional surf competition at sunset, holding the same surfboard under his arm, now wearing a red competition jersey with the number 7 over his wetsuit. Colorful sponsor flags, a judges' tower and a cheering crowd along the shore, big golden waves breaking behind him. His face and pose are unchanged. Cinematic sports photography.Image Generation
The same woman on a red carpet at a glamorous evening gala, wearing a flowing emerald-green silk gown with delicate gold earrings, her curly hair styled in an elegant updo. Photographers' flashes and a softly blurred step-and-repeat wall behind her, warm spotlights. Her face is unchanged, high-end fashion editorial photography.Image Generation
The same rustic wooden table with a freshly baked Neapolitan pizza made from these ingredients, with a blistered charred crust, melted mozzarella, tomato sauce and fresh basil leaves, steam rising. A wood-fired oven glows softly in the background, warm cozy light, food photography.Image Generation
The same street scene fully restored and colorized: scratches, creases, dust and torn corners removed, with natural, historically accurate colors on the buildings, trams, clothing and sky. Sharp detail, soft daylight, the composition and every person exactly where they were.Image Generation

API quick start

Run nano banana 2.1 advanced with a single API call.

POST https://api.wiro.ai/v1/Run/google/nano-banana-2-1-advanced
{
  "prompt": "The same surfer standing on the beach at …",
  "resolution": "1K",
  "inputImage": "https://your-cdn.com/input.png",
  "thinkingLevel": "medium"
}
curl
curl -X POST "https://api.wiro.ai/v1/Run/google/nano-banana-2-1-advanced" \
  -H "Content-Type: application/json" \
  -H "x-api-key: YOUR_WIRO_API_KEY" \
  --data-binary @- <<'JSON'
{
  "prompt": "The same surfer standing on the beach at …",
  "resolution": "1K",
  "inputImage": "https://your-cdn.com/input.png",
  "thinkingLevel": "medium"
}
JSON
View full API docs

Discover, test, and run AI models, build workflows and agents with one unified API.

All systems operational
WiroAboutBlogCareersContact
ProductModelsAgentsPricingPartnerChangelogStatusFAQ
ModelsNano Banana 2GPT Image 2.5Seedream V5 ProSeedance 2.5Veo 3.1Kling V3FLUX 3FLUX.2 ProWan 3.0 PrimeGrok Imagine 1.5
PartnersGoogleOpenAIByteDanceBlack Forest LabsKling AIQwenAlibabaxAIMiniMaxElevenLabs
Getting StartedIntroductionAuthenticationProjectsCode ExamplesWiro MCP ServerSelf-Hosted MCPn8n IntegrationLLMs.txt
API ReferenceModelsRun a ModelModel ParametersTasksLLM & Chat StreamingWebSocketRealtime VoiceFiles
© 2026 Wiro AI. All rights reserved.
PrivacyTermsData Deletion