Try Google Gemini Omni Flash Video Generator from Google →
Models
Agents
WorkflowsStudioPricingBlogDocs
ExploreDiscover models by categoryBrowse All ModelsBrowse the complete catalogSee FavoritesSign in to view saved models
OverviewThe platform at a glanceLearnSkills, knowledge, guardrailsAnatomyWhat makes agents reasonBuild Your AgentPick skills, set tier, deploy
Pre-built AgentsBrowse the catalog
Agent Usecases
Ad Campaign ManagerApp Event ManagerApp Review RepliesBarber BookingCustomer Win-BackEcommerce ListingsRestaurant Reviews
Sign InStart Building

Task History

Click to see output list

No tasks yet

Go to Models
DISCOVER AI MODELS

Search the full model catalog.

Discover and compare production-ready models. Filter by modality, provider, pricing, and production fit.

5 models1 active filters

Active filters

Wiro AI

Video Generation

Image Generation

Audio & Speech

Realtime Stream

Music Generation

3D Generation

LLM & Chat

Workflow

Partners

1
Need help choosing?

Compare models by capability, cost, and provider fit.

Showing 5 of 5 models
LocateAnything-3B - AI model cover image

LocateAnything-3B

nvidia

LocateAnything 3B by Nvidia localizes objects, UI elements, and text in images. It returns normalized boxes or points from natural-language prompts.

Image to Imageenglish
Run

Cosmos3-Super

nvidia

Cosmos 3 Super turns a reference image and motion prompt into a physics-grounded MP4 clip, with controls for aspect ratio, FPS, audio, and safety.

Text to Video
Run
parakeet-tdt-0.6b-v3 - AI model cover image

parakeet-tdt-0.6b-v3

nvidia

Multilingual speech-to-text for 25 European languages with auto language detection, punctuation, capitalization, and optional timestamps.

Speech to Texttransformersspeech-to-text
Run
PersonaPlex-Realtime - AI model cover image

PersonaPlex-Realtime

nvidia

Convert speech to speech with customizable voices using PersonaPlex. Supports various audio formats and offers control over text temperature and audio top K settings.

Speech to Speechspeech-to-textaudio
Run
nemotron - AI model cover image

nemotron

nvidia

Nemotron-Speech-Streaming-En-0.6b is the first unified model in the Nemotron Speech family, engineered to deliver high-quality English transcription across both low-latency streaming and high-throughput batch workloads. The model natively supports punctuation and capitalization and offers runtime flexibility with configurable chunk sizes, including 80ms, 160ms, 560ms, and 1120ms.

Speech to Textspeech-to-textaudio
Run

Discover, test, and run AI models, build workflows and agents with one unified API.

All systems operational
WiroAboutBlogCareersContact
ProductModelsAgentsPricingChangelogStatusFAQ
Getting StartedIntroductionAuthenticationProjectsCode ExamplesWiro MCP ServerSelf-Hosted MCPn8n IntegrationLLMs.txt
API ReferenceModelsRun a ModelModel ParametersTasksLLM & Chat StreamingWebSocketRealtime VoiceFiles
© 2026 Wiro AI. All rights reserved.
PrivacyTermsData Deletion