Try Google Gemini Omni Flash Video Generator from Google →
Models
Agents
WorkflowsStudioPricingBlogDocs
ExploreDiscover models by categoryBrowse All ModelsBrowse the complete catalogSee FavoritesSign in to view saved models
OverviewThe platform at a glanceLearnSkills, knowledge, guardrailsAnatomyWhat makes agents reasonBuild Your AgentPick skills, set tier, deploy
Pre-built AgentsBrowse the catalog
Agent Usecases
Ad Campaign ManagerApp Event ManagerApp Review RepliesBarber BookingCustomer Win-BackEcommerce ListingsRestaurant Reviews
Sign InStart Building

Task History

Click to see output list

No tasks yet

Go to Models
Explore models/
Image GenerationActive

nvidia / LocateAnything-3B

LocateAnything-3B

bynvidia

LocateAnything 3B by Nvidia localizes objects, UI elements, and text in images. It returns normalized boxes or points from natural-language prompts.

Image to ImageFast Inference
Model ID
LocateAnything-3B
Provider
nvidia
Updated
1781601543
LocateAnything-3B
0
Comments
Average rating : 2 (1 users)
Providernvidia
ModelLocateAnything-3B
Image to ImageFast Inference
wiro playground—nvidia/LocateAnything-3B
Reset to defaults
0 / 1
Maximum 1 image allowed
Drop image to upload

OR

Click to browse your device

Supports: JPG, JPEG, PNG, GIF, WEBP, HEIC

Input image to run grounding on.

Selects the grounding mode.

Natural-language description of the target. Required for detect, ground_single, ground_multi, ground_text, gui, and point.

Sample outputs
Sample 1
Sample 2
Sample 3
Sample 4
Updated 1781601543
## Overview LocateAnything-3B is an image grounding model from Nvidia. It combines a vision encoder with a language model, then generates a text response that encodes locations as boxes or points. It uses a block-based decoding scheme that predicts full coordinates together, instead of emitting coordinates one token at a time. This output format makes it useful for pipelines that must turn natural language into pixel-accurate regions. ## What you can build - Open-vocabulary object detection where labels come from your category list - Referring expression grounding like “the person in a red jacket” in crowded scenes - GUI element grounding on screenshots, such as “the search button” or “settings icon” - Text region localization for OCR pre-processing and document workflows - Pointing actions for robotics or agents that need a single target coordinate - Dataset annotation helpers that produce boxes you can review and correct ## Inputs - One RGB image as a JPG or PNG. Use the original resolution when you need tight boxes. - A task choice that controls what the model should return, such as detecting categories, grounding a single best match, grounding all matches, finding text, or pointing. - A natural-language target description for tasks that need it. Write a short phrase or instruction that uniquely identifies the region. - For category detection, provide the category names you want found in the image. A comma-separated list works well. - For GUI localization, choose whether you want a bounding box or a single click point. ## Outputs - A text response that includes structured location tokens. - Bounding boxes appear as a labeled region plus 4 integer coordinates in a 0–1000 normalized space. Convert to pixels by scaling by image width and height. - Points appear as 2 integer coordinates in the same 0–1000 normalized space. - When nothing matches, the model can return an explicit “none” marker for the location. - The response may include multiple labeled regions when the task asks for many matches. ## Recommended settings - Use “Ground Single” when you want one best match for a phrase. - Use “Ground Multi” when you expect several instances of the same description. - Use “Detect” when you have a fixed set of categories and want all instances. - Use “Detect Text” to get text region boxes without supplying a text phrase. - Use the GUI “Point” output when you need a click target instead of a box. ## Limitations - Outputs are quantized to a 0–1000 normalized coordinate grid. Very small objects may need higher-resolution inputs to stay precise. - Some scenes produce ambiguous matches. Prompts like “the biggest one” can fail when size cues are unclear. - Dense scenes can cause missed instances or mixed labels when many objects overlap. - Fast structured decoding can occasionally produce malformed location tokens. Your parser should validate tags and coordinate counts. - Low-quality inputs reduce accuracy. Blurry photos, heavy compression, extreme downscaling, glare, and scanned documents with artifacts can all degrade box quality. - Text localization depends on legibility. Small fonts, curved text, and stylized fonts can reduce recall. ## Safety & compliance - This model is released under an Nvidia research license. It is intended for research and development use, not unrestricted commercial deployment. - Don’t upload images that include sensitive personal data unless you have permission and a lawful purpose. - If you use it for GUI automation or document processing, add guardrails. Verify targets before taking actions like clicks or deletions. - Follow applicable laws and site terms when using it for surveillance, monitoring, or automated interaction.

Example prompts

Great starting points for LocateAnything-3B.

largest boatImage Generation
tennis ballImage Generation
umbrellaImage Generation
System OverviewImage Generation

API quick start

Run LocateAnything-3B with a single API call.

POST https://api.wiro.ai/v1/Run/nvidia/LocateAnything-3B
{
  "prompt": "largest boat",
  "inputImage": "https://your-cdn.com/input.png",
  "task": "ground_single",
  "outputType": "box"
}
View full API docs

Discover, test, and run AI models, build workflows and agents with one unified API.

All systems operational
WiroAboutBlogCareersContact
ProductModelsAgentsPricingPartnerChangelogStatusFAQ
Getting StartedIntroductionAuthenticationProjectsCode ExamplesWiro MCP ServerSelf-Hosted MCPn8n IntegrationLLMs.txt
API ReferenceModelsRun a ModelModel ParametersTasksLLM & Chat StreamingWebSocketRealtime VoiceFiles
© 2026 Wiro AI. All rights reserved.
PrivacyTermsData Deletion