Try MiniMax FastH3 from Fastvideo →
Models
Agents
Workflows
Studio
PricingBlogDocs
ExploreDiscover models by categoryBrowse All ModelsBrowse the complete catalogSee FavoritesSign in to view saved models
Generative Media AgentCreate and edit media by chattingWorkflow AgentBuild visual workflows with Agent
OverviewThe platform at a glanceLearnSkills, knowledge, guardrailsAnatomyWhat makes agents reasonBuild Your AgentPick skills, set tier, deploy
Pre-built AgentsBrowse the catalog
Agent Usecases
Ad Campaign ManagerApp Event ManagerApp Review RepliesBarber BookingCustomer Win-BackEcommerce ListingsRestaurant Reviews
Sign InStart Building

Task History

Click to see output list

No tasks yet

Go to Models
AI MODELS

Production-ready models,organized for real work.

Browse and compare image, video, audio, and language models from leading providers, all available through the Wiro API.

More filters

Recently Added

522 models in this collection.

View all
Hair Cut Effects
wiro

Transform a single portrait into a viral cut-your-own-hair video. Say the words, make the first snip, and reveal a perf…

Image to VideoSocial Media & ViralFast Inference
Airplane Wing Stunt Trend
wiro

Transform a single portrait into a high-altitude wing-walking stunt video. Step out onto the wing of a plane in flight,…

Image to VideoSocial Media & ViralFast Inference
openai/gpt-6-astra model cover
gpt-6-astra
openai

OpenAI’s GPT 6 Astra handles long-context reasoning with optional image input and tool calls. Build agents for coding,…

LLMPartner LLMReasoning
wiro/video-caption model cover
video-caption
wiro

Burn captions into a video with TikTok-style timing. Transcribe spoken audio word-by-word or overlay fixed text, with f…

Video to VideoUtility
google/gemini-3.1-tts model cover
gemini-3.1-tts
google

Generate natural speech from text with Gemini 3.1 Flash TTS. Use voice options and expressive tags to control tone and…

Text to Speech
glm/5.2 model cover
5.2
glm

GLM 5.2 is a text-only MoE model with a usable 1M-token context for long-horizon coding and agent workflows. It support…

Partner LLMGlmLLM Tool Call
deepseek/v4-flash model cover
v4-flash
deepseek

DeepSeek V4 Flash is a fast MoE text model with a 1M-token context window and optional thinking mode. It’s built for ch…

Partner LLMDeepseekLLM Tool Call
deepseek/v4-pro model cover
v4-pro
deepseek

DeepSeek V4 Pro (v4-pro) is a long-context text model with optional thinking mode for agentic coding, math reasoning, a…

Partner LLMDeepseekLLM Tool Call
google/gemini-3.8-flash model cover
gemini-3.8-flash
google

Google’s Gemini 3.8 Flash reasons over text, images, audio, video, and PDFs, then returns strong text answers for long-…

Partner LLMLLM Tool Call
xai/grok-imagine-image-v2 model cover
grok-imagine-image-v2
xai

Create new images or edit a source photo with xAI’s Grok Imagine Image V2. Choose 1K or 2K output, quality level, and a…

Text to ImageImage to Image
p-video-edit
pruna

Pruna P Video Edit edits a short MP4 clip from one text instruction, with 1-4 reference images to guide identity or sty…

Video to Video
elevenlabs/speech-to-speech-v2 model cover
speech-to-speech-v2
elevenlabs

Re-voice an existing recording with a chosen ElevenLabs voice while keeping the original words, timing, and delivery. E…

Speech to Speech
elevenlabs/text-to-music-v2 model cover
text-to-music-v2
elevenlabs

Create original music from a plain-English brief with ElevenLabs Music v2. Choose track length, MP3 quality, and option…

Text to Music
Fast-H3
fastvideo

FastH3 is a four-forward distilled MiniMax H3 model that generates synchronized video and stereo audio from a text prom…

Text to VideoFast InferenceH200
unsloth/Qwen3.8-Flash-Next-GGUF model cover
Qwen3.8-Flash-Next-GGUF
unsloth

Run Unsloth’s GGUF quantizations of Qwen 3.8 Flash Next for long-context chat and reasoning. Includes optional thinking…

ChatLLMReasoning
qwen/Qwen3.8-27B-Obliterated model cover
Qwen3.8-27B-Obliterated
qwen

Refusal-reduced variant of Qwen 3.8 27B for long-context chat and coding. It can emit or hide thinking traces and suppo…

ChatLLMReasoning
claude/fable-5 model cover
fable-5
claude

Claude Fable 5 is a frontier text model with vision and PDF support. It handles 1M-token context for deep reasoning, co…

Partner LLMLLM Tool Call
claude/opus-5 model cover
opus-5
claude

Claude Opus 5 is a premium text model with vision for agentic coding and document work. Upload images or PDFs and get l…

Partner LLMLLM Tool Call
claude/sonnet-5 model cover
sonnet-5
claude

Claude Sonnet 5 is an agent-ready text model for coding, reasoning, and long-context work. It can read images and PDFs…

Partner LLMLLM Tool Call
openai/gpt-realtime-whisper model cover
gpt-realtime-whisper
openai

OpenAI’s GPT Realtime Whisper turns live audio into streaming transcript updates. Adjust delay levels to trade latency…

Popular Models

522 models in this collection.

View all
bytedance/seedream-v4-5-uncensored model cover
seedream-v4-5-uncensored
bytedance

Generate high-resolution images using Seedream v4.5 Uncensored. Supports text-to-image and image-to-image transformatio…

Text to ImageImage to Image
kling-v2.6-motion-control
klingai

Generates videos from images and reference videos with motion control. Supports custom prompts and character orientatio…

Image to Video
kling-v3
klingai

Generate high-quality videos from text prompts using Kling V3. Supports custom frames, duration, and aspect ratios.

Text to VideoImage to Video
google/nano-banana-2 model cover
nano-banana-2
google

An image editing tool designed for quick transformations using reference images and prompts. Supports multi-image mixin…

Text to ImageImage to Image
seedance-pro-v1.5-uncensored
bytedance

Seedance Pro v1.5 Uncensored by ByteDance generates short videos from text with optional native audio, strong prompt fo…

Text to VideoImage to Video
gemini-omni-flash-r2v
google

Generate 3–10s videos from reference images, or edit an existing clip with one instruction. Built on Google Gemini Omni…

Image to VideoVideo to Video
gemini-omni-flash
google

Generate 5s,10s 720p MP4 videos from text, or animate a still image as the opening frame. Google Gemini Omni Flash supp…

Text to VideoImage to Video
google/nano-banana-2-lite model cover
nano-banana-2-lite
google

Google’s Nano Banana 2 Lite generates and edits 1K images from text, with optional multi-image references. It’s built f…

Text to ImageImage to Image
google/nano-banana-pro model cover
nano-banana-pro
google

Google's Gemini 3 Pro Image Preview, also known as Nano Banana, model for text-to-image and image-to-image generation.

Text to ImageImage to Image
qwen/Qwen3.8-27B-Uncensored model cover
Qwen3.8-27B-Uncensored
qwen

Chat model based on Qwen 3.8 27B with refusal behavior reduced through direction removal. Its 262k context fits long do…

ChatLLM
seedance-2.5-uncensored
bytedance

Create 4–30s videos from a description, or steer motion using first and last frame images. Choose 480p or 720p output w…

Text to VideoImage to Video
qwen/Qwen3.8-27B-Obliterated model cover
Qwen3.8-27B-Obliterated
qwen

Refusal-reduced variant of Qwen 3.8 27B for long-context chat and coding. It can emit or hide thinking traces and suppo…

ChatLLM
unsloth/Qwen3.8-Flash-Next-GGUF model cover
Qwen3.8-Flash-Next-GGUF
unsloth

Run Unsloth’s GGUF quantizations of Qwen 3.8 Flash Next for long-context chat and reasoning. Includes optional thinking…

ChatLLM
wiro/Video Converter model cover
Video Converter
wiro

Convert a video to MP4, MOV, WebM, MKV, AVI, MPEG, or M4V with adjustable compression. Built by Wiro for clean exports…

Video to VideoUtility
wiro/Image Converter model cover
Image Converter
wiro

Convert an image to JPEG, PNG, WebP, TIFF, or AVIF. Set output quality from 0 to 100 to balance file size and visual fi…

Image to ImageUtility
panini-card
wiro

Turn a selfie into a Panini-style player card video. Enter name, height, weight, and birth date, then choose a country…

Image to VideoSocial Media & Viral
openai/gpt-image-2-custom model cover
gpt-image-2-custom
openai

Create or edit images with OpenAI GPT Image 2 using custom pixel sizes, quality tiers, and format controls. Add optiona…

Text to ImageImage to Image
wiro/smart resize model cover
smart resize
wiro

Smart Resize by Wiro converts one image into multiple exact sizes, using AI recomposition to keep key subjects in frame…

Image to Image
seedance 2.5 reference
bytedance

Create 4–30 second videos guided by reference clips, images, and audio. Preserve motion and timing while changing subje…

Image to VideoVideo to Video
openai/gpt-image-2 model cover
gpt-image-2
openai

Generate or edit images with GPT Image 2 from OpenAI. It delivers strong instruction following, sharp text rendering, a…

Text to ImageImage to Image

Generate Videos

165 models in this collection.

View all
Hair Cut Effects
wiro

Transform a single portrait into a viral cut-your-own-hair video. Say the words, make the first snip, and reveal a perf…

Image to VideoSocial Media & Viral
Airplane Wing Stunt Trend
wiro

Transform a single portrait into a high-altitude wing-walking stunt video. Step out onto the wing of a plane in flight,…

Image to VideoSocial Media & Viral
wiro/video-caption model cover
video-caption
wiro

Burn captions into a video with TikTok-style timing. Transcribe spoken audio word-by-word or overlay fixed text, with f…

Video to VideoUtility
p-video-edit
pruna

Pruna P Video Edit edits a short MP4 clip from one text instruction, with 1-4 reference images to guide identity or sty…

Video to Video
Fast-H3
fastvideo

FastH3 is a four-forward distilled MiniMax H3 model that generates synchronized video and stereo audio from a text prom…

Text to VideoFast Inference
seedance-2.5-uncensored
bytedance

Create 4–30s videos from a description, or steer motion using first and last frame images. Choose 480p or 720p output w…

Text to VideoImage to Video
seedance 2.5 reference
bytedance

Create 4–30 second videos guided by reference clips, images, and audio. Preserve motion and timing while changing subje…

Image to VideoVideo to Video
seedance 2.5
bytedance

Create 4–30 second videos from a prompt and optional first and last frame images, with optional synced audio. Powered b…

Text to VideoImage to Video
h3-r2v
minimax

MiniMax H3 R2V generates 4–15s videos with synced stereo audio from reference images, clips, and optional audio. Use it…

Image to VideoVideo to Video
h3
minimax

MiniMax H3 I2V generates 4–15s videos from a prompt and optional first or last frame. It outputs MP4 video at 24 FPS wi…

Text to VideoImage to Video
flux-3-v2v
blackforestlabs

Extend a source video past its final frame. FLUX 3 V2V generates 5–20 seconds of 720p or 1080p footage, with optional s…

Video to Video
flux-3
blackforestlabs

FLUX 3 I2V turns 1–10 images and a prompt into a 5–20s MP4 clip at 720p or 1080p, with optional native audio and keyfra…

Text to VideoImage to Video
ugc creator v2
wiro

Create UGC-style product ad videos from a single product photo and a short script. Choose a preset scene, duration, asp…

Image to VideoSocial Media & Viral
gen 4.5
runway

Runway Gen-4.5 turns text or a first-frame image into 2 to 10s 720p video clips. It follows sequenced actions and camer…

Text to VideoImage to Video
gemini-omni-flash-r2v
google

Generate 3–10s videos from reference images, or edit an existing clip with one instruction. Built on Google Gemini Omni…

Image to VideoVideo to Video
gemini-omni-flash
google

Generate 5s,10s 720p MP4 videos from text, or animate a still image as the opening frame. Google Gemini Omni Flash supp…

Text to VideoImage to Video
seedance 2.0 mini v2v
bytedance

ByteDance’s Seedance 2.0 Mini V2V edits short clips from 1–3 reference videos and a prompt. It outputs 480p or 720p MP4…

Image to VideoVideo to Video
seedance 2.0 mini
bytedance

Seedance 2.0 Mini by ByteDance generates short MP4 videos from a prompt, with optional start and end frames or referenc…

Image to VideoVideo to Video
kling-v3-turbo
klingai

Kling V3 Turbo turns a prompt, or a still image plus motion instructions, into a short MP4 video with native audio and…

Text to VideoImage to Video
Hero Vox Effects
wiro

Create a short talking-portrait video from a human photo and a script. Choose an effect preset, aspect ratio, and 5, 10…

Image to VideoSocial Media & Viral

Generate Images

86 models in this collection.

View all
xai/grok-imagine-image-v2 model cover
grok-imagine-image-v2
xai

Create new images or edit a source photo with xAI’s Grok Imagine Image V2. Choose 1K or 2K output, quality level, and a…

Text to ImageImage to Image
bytedance/seedream-v5-pro-uncensored model cover
seedream-v5-pro-uncensored
bytedance

Generate 1K or 2K images with strong typography and reference-guided edits. Built on ByteDance Seedream 5.0 Pro for pos…

Text to ImageImage to Image
pruna/p-image-ideogram-custom model cover
p-image-ideogram-custom
pruna

Pruna’s P Image Ideogram Custom generates poster-ready images from a short prompt. It’s tuned for readable text, typogr…

Text to Image
pruna/p-image-ideogram model cover
p-image-ideogram
pruna

P-Image Ideogram by Pruna generates high-quality text-to-image outputs with strong typography, offering 1K or 2K resolu…

Text to Image
bytedance/seedream-v5-pro model cover
seedream-v5-pro
bytedance

ByteDance Seedream V5 Pro generates images from text and edits a base image using up to 10 references. It’s tuned for i…

Text to ImageImage to Image
google/nano-banana-2-lite model cover
nano-banana-2-lite
google

Google’s Nano Banana 2 Lite generates and edits 1K images from text, with optional multi-image references. It’s built f…

Text to ImageImage to Image
microsoft/lens model cover
lens
microsoft

A foundational text-to-image model designed for efficient, high-resolution image generation with strong prompt followin…

Text to ImageFast Inference
openai/gpt-image-2-custom model cover
gpt-image-2-custom
openai

Create or edit images with OpenAI GPT Image 2 using custom pixel sizes, quality tiers, and format controls. Add optiona…

Text to ImageImage to Image
sensenova/U1-8B-Text-to-Image model cover
U1-8B-Text-to-Image
sensenova

SenseNova U1-8B turns text into high-detail images with strong layout control and clearer in-image text. Use it for pos…

Text to ImageFast Inference
openai/gpt-image-2 model cover
gpt-image-2
openai

Generate or edit images with GPT Image 2 from OpenAI. It delivers strong instruction following, sharp text rendering, a…

Text to ImageImage to Image
xai/grok-imagine-image model cover
grok-imagine-image
xai

xAI’s Grok Imagine Image generates images from prompts and can edit a source image using instructions. Pick an aspect r…

Text to ImageImage to Image
openai/gpt-image-1-5 model cover
gpt-image-1-5
openai

Generate or edit images using text prompts or image edits with GPT Image 1.5. Supports multiple sizes, formats, and qua…

Text to ImageImage to Image
bytedance/seedream-v4-5-uncensored model cover
seedream-v4-5-uncensored
bytedance

Generate high-resolution images using Seedream v4.5 Uncensored. Supports text-to-image and image-to-image transformatio…

Text to ImageImage to Image
fireredteam/FireRed-Image-Edit-1.1 model cover
FireRed-Image-Edit-1.1
fireredteam

FireRed-Image-Edit-1.1 significantly enhances identity consistency, multi-image conditioning, and domain-specialized ed…

Text to ImageImage to Image
bytedance/seedream-v5-lite-uncensored model cover
seedream-v5-lite-uncensored
bytedance

Generate high-quality images from text prompts or image inputs using the Seedream v5 Lite Uncensored model. Supports mu…

Text to ImageImage to Image
google/nano-banana-2 model cover
nano-banana-2
google

An image editing tool designed for quick transformations using reference images and prompts. Supports multi-image mixin…

Text to ImageImage to Image
bytedance/seedream-v5-lite model cover
seedream-v5-lite
bytedance

Generate high-quality images using Seedream V5 Lite, supporting both image-to-image and text-to-image transformations w…

Text to ImageImage to Image
fireredteam/FireRed-Image-Edit model cover
FireRed-Image-Edit
fireredteam

FireRed-Image-Edit is a general-purpose image editing model that delivers high-fidelity and consistent editing across a…

Text to ImageImage to Image
zai-org/GLM-IMAGE model cover
GLM-IMAGE
zai-org

GLM-Image is an image generation model adopts a hybrid autoregressive + diffusion decoder architecture. In general imag…

Text to ImageImage to Image
black-forest-labs/FLUX.2-klein-base-9B model cover
FLUX.2-klein-base-9B
black-forest-labs

FLUX.2 [klein] 9B Base is a 9 billion parameter rectified flow transformer capable of generating images from text descr…

Text to ImageImage to Image

Edit Images

81 models in this collection.

View all
xai/grok-imagine-image-v2 model cover
grok-imagine-image-v2
xai

Create new images or edit a source photo with xAI’s Grok Imagine Image V2. Choose 1K or 2K output, quality level, and a…

Text to ImageImage to Image
bytedance/seedream-v5-pro-uncensored model cover
seedream-v5-pro-uncensored
bytedance

Generate 1K or 2K images with strong typography and reference-guided edits. Built on ByteDance Seedream 5.0 Pro for pos…

Text to ImageImage to Image
bytedance/seedream-v5-pro model cover
seedream-v5-pro
bytedance

ByteDance Seedream V5 Pro generates images from text and edits a base image using up to 10 references. It’s tuned for i…

Text to ImageImage to Image
briaai/RMBG-2.0 model cover
RMBG-2.0
briaai

RMBG-2.0 removes image backgrounds with an 8-bit alpha matte for smooth edges. Built by BRIA AI for production cutouts…

Image to ImageEcommerce
google/nano-banana-2-lite model cover
nano-banana-2-lite
google

Google’s Nano Banana 2 Lite generates and edits 1K images from text, with optional multi-image references. It’s built f…

Text to ImageImage to Image
ideogram/remove-background model cover
remove-background
ideogram

Remove image backgrounds and get a clean transparent PNG cutout. Ideogram’s generative matting keeps hair, glass, and f…

Image to ImageIdeogram
Birthday Mode Effects
wiro

Turn your photo into a fun birthday celebration while keeping your face natural and unchanged for instantly share-ready…

Image to VideoImage to Image
Birthday Mode Effects with Caption
wiro

Turn your photo into a fun birthday celebration with your age featured on the scene, while keeping your face natural an…

Image to VideoImage to Image
pruna/p-image-try-on model cover
p-image-try-on
pruna

Pruna’s p-image-try-on creates virtual try-on images by applying one or more garment photos to a person photo. Generate…

Image to ImageFast Inference
nvidia/LocateAnything-3B model cover
LocateAnything-3B
nvidia

LocateAnything 3B by Nvidia localizes objects, UI elements, and text in images. It returns normalized boxes or points f…

Image to ImageFast Inference
wiro/Image Converter model cover
Image Converter
wiro

Convert an image to JPEG, PNG, WebP, TIFF, or AVIF. Set output quality from 0 to 100 to balance file size and visual fi…

Image to ImageUtility
wiro/smart resize model cover
smart resize
wiro

Smart Resize by Wiro converts one image into multiple exact sizes, using AI recomposition to keep key subjects in frame…

Image to Image
sensenova/U1-8B-Interleave model cover
U1-8B-Interleave
sensenova

u1-8b-interleave by SenseNova generates step-by-step text with matching images, optionally conditioned on up to 5 refer…

Image to ImageImage to Text
openai/gpt-image-2-custom model cover
gpt-image-2-custom
openai

Create or edit images with OpenAI GPT Image 2 using custom pixel sizes, quality tiers, and format controls. Add optiona…

Text to ImageImage to Image
openai/gpt-image-2 model cover
gpt-image-2
openai

Generate or edit images with GPT Image 2 from OpenAI. It delivers strong instruction following, sharp text rendering, a…

Text to ImageImage to Image
xai/grok-imagine-image model cover
grok-imagine-image
xai

xAI’s Grok Imagine Image generates images from prompts and can edit a source image using instructions. Pick an aspect r…

Text to ImageImage to Image
openai/gpt-image-1-5 model cover
gpt-image-1-5
openai

Generate or edit images using text prompts or image edits with GPT Image 1.5. Supports multiple sizes, formats, and qua…

Text to ImageImage to Image
wiro/Instagram Pose Multi model cover
Instagram Pose Multi
wiro

Generate stylish Instagram-style pose images with trendy angles, natural expressions, and a modern aesthetic. Built by…

Image to ImageSocial Media & Viral
bytedance/seedream-v4-5-uncensored model cover
seedream-v4-5-uncensored
bytedance

Generate high-resolution images using Seedream v4.5 Uncensored. Supports text-to-image and image-to-image transformatio…

Text to ImageImage to Image
fireredteam/FireRed-Image-Edit-1.1 model cover
FireRed-Image-Edit-1.1
fireredteam

FireRed-Image-Edit-1.1 significantly enhances identity consistency, multi-image conditioning, and domain-specialized ed…

Text to ImageImage to Image

Generate Text

28 models in this collection.

View all
sensenova/U1-8B-Interleave model cover
U1-8B-Interleave
sensenova

u1-8b-interleave by SenseNova generates step-by-step text with matching images, optionally conditioned on up to 5 refer…

Image to ImageImage to Text
sensenova/U1-8B-Visual-Understanding model cover
U1-8B-Visual-Understanding
sensenova

SenseNova U1 8B Visual Understanding creates infographic-style images and prompt-based edits from text and an optional…

Image to TextFast Inference
nvidia/parakeet-tdt-0.6b-v3 model cover
parakeet-tdt-0.6b-v3
nvidia

Multilingual speech-to-text for 25 European languages with auto language detection, punctuation, capitalization, and op…

Speech to TextFast Inference
coherelabs/cohere-transcribe-03-2026 model cover
cohere-transcribe-03-2026
coherelabs

CohereLabs cohere-transcribe-03-2026 is a 2B Conformer speech-to-text model for 14 languages. It creates accurate trans…

Speech to TextFast Inference
qwen/Qwen3-ASR-1.7B model cover
Qwen3-ASR-1.7B
qwen

A lightweight speech-to-text model optimized for fast inference. Converts audio input into text with support for multip…

Speech to TextFast Inference
mistralai/Voxtral-Mini-4B-Realtime-2602 model cover
Voxtral-Mini-4B-Realtime-2602
mistralai

Voxtral Mini 4B Realtime 2602 is a multilingual, realtime speech-transcription model and among the first open-source so…

Speech to TextRealtime STT
nvidia/nemotron model cover
nemotron
nvidia

Nemotron-Speech-Streaming-En-0.6b is the first unified model in the Nemotron Speech family, engineered to deliver high-…

Speech to TextAudio
elevenlabs/speech-to-text model cover
speech-to-text
elevenlabs

Speech to text model from ElevenLabs

Speech to Text
openai/whisper-medium model cover
whisper-medium
openai

OpenAI Whisper Medium transcribes speech to text with timestamps and optional speaker diarization. Works across many la…

Speech to TextWhisper
moondream3-preview/detect model cover
detect
moondream3-preview

Moondream3 is a cutting-edge vision-language model that delivers advanced visual reasoning with built-in object detecti…

Image to TextBf16
moondream3-preview/point model cover
point
moondream3-preview

Moondream3 is a cutting-edge vision-language model that delivers advanced visual reasoning with built-in object detecti…

Image to TextBf16
moondream3-preview/caption model cover
caption
moondream3-preview

Moondream3 is a cutting-edge vision-language model that delivers advanced visual reasoning with built-in object detecti…

Image to TextBf16
moondream3-preview/query model cover
query
moondream3-preview

Moondream3 is a cutting-edge vision-language model that delivers advanced visual reasoning with built-in object detecti…

Image to TextBf16
openai/whisper-large-v3-turbo-turkish model cover
whisper-large-v3-turbo-turkish
openai

Whisper is a pre-trained model for automatic speech recognition (ASR) and speech translation. Trained on 680k hours of…

Speech to TextWhisper
wiro/video-nsfw-detection model cover
video-nsfw-detection
wiro

NSFW video detection automatically analyzes video content to identify inappropriate or explicit material, ensuring comp…

Video to Text
VideoLLaMA3-2B
damo-nlp-sg

VideoLLaMA3-2B is a model designed for video understanding.

Video to Text
VideoLLaMA3-2B-Image
damo-nlp-sg

VideoLLaMA3-2B-Image is a model designed for image understanding.

Image to Text
wiro/VideoLLaMA3-7B-Image model cover
VideoLLaMA3-7B-Image
wiro

VideoLLaMA3-7B-Image is a model designed for image understanding.

Image to Text
wiro/VideoLLaMA3-7B model cover
VideoLLaMA3-7B
wiro

VideoLLaMA3-7B is a model designed for video understanding.

Video to Text
salesforce/blip2-flan-t5-xl model cover
blip2-flan-t5-xl
salesforce

BLIP-2 creates captions or detailed descriptions for images. This is BLIP-2 model, leveraging Flan T5-xl.

Image to Text

Generate 3D

4 models in this collection.

View all
tencentarc/Pixal3D model cover
Pixal3D
tencentarc
3D Generation
HY-World-2.0-World-Reconstruction
tencent
3D Generation
microsoft/Trellis-2 model cover
Trellis-2
microsoft
3D Generation
tencent/Hunyuan3D-2.1 model cover
Hunyuan3D-2.1
tencent
3D Generation

Generate Audio

31 models in this collection.

View all
google/gemini-3.1-tts model cover
gemini-3.1-tts
google
Text to Speech
elevenlabs/speech-to-speech-v2 model cover
speech-to-speech-v2
elevenlabs
Speech to Speech
openai/gpt-realtime-translate model cover
gpt-realtime-translate
openai
Speech to Speech
openai/gpt-realtime-2.1-mini model cover
gpt-realtime-2.1-mini
openai
Speech to Speech
openai/gpt-realtime-2.1 model cover
gpt-realtime-2.1
openai
Speech to Speech
openmoss/MOSS-TTS-v1.5 model cover
MOSS-TTS-v1.5
openmoss
Text to Speech
openbmb/VoxCPM2 model cover
VoxCPM2
openbmb
Text to Speech
k2-fsa/OmniVoice model cover
OmniVoice
k2-fsa
Text to Speech
humeai/tada-3b-ml model cover
tada-3b-ml
humeai
Text to Speech
fishaudio/s2-pro model cover
s2-pro
fishaudio
Text to Speech
nineninesix/kani-tts-2-en model cover
kani-tts-2-en
nineninesix
Text to Speech
resemble-ai/chatterbox-turbo model cover
chatterbox-turbo
resemble-ai
Text to Speech
resemble-ai/chatterbox-multilingual  model cover
chatterbox-multilingual
resemble-ai
Text to Speech
openmoss/MOSS-TTSD model cover
MOSS-TTSD
openmoss
Text to Speech
openmoss/MOSS-TTS-Realtime model cover
MOSS-TTS-Realtime
openmoss
Text to Speech
elevenlabs/Realtime Conversational AI model cover
Realtime Conversational AI
elevenlabs
Speech to Speech
openai/gpt-realtime-mini model cover
gpt-realtime-mini
openai
Speech to Speech
openai/gpt-realtime model cover
gpt-realtime
openai
Speech to Speech
qwen/Qwen3-TTS-12Hz-1.7B model cover
Qwen3-TTS-12Hz-1.7B
qwen
Text to Speech
nvidia/PersonaPlex-Realtime model cover
PersonaPlex-Realtime
nvidia
Speech to Speech

Generate Music

18 models in this collection.

View all
elevenlabs/text-to-music-v2 model cover
text-to-music-v2
elevenlabs
Text to Music
stabilityai/stable-audio-3-small-sfx model cover
stable-audio-3-small-sfx
stabilityai
Text to Music
stabilityai/stable-audio-3-small-music model cover
stable-audio-3-small-music
stabilityai
Text to Music
stabilityai/stable-audio-3-medium model cover
stable-audio-3-medium
stabilityai
Text to Music
google/lyria 3 model cover
lyria 3
google
Text to Song
tencent-ailab/SongGeneration 2 model cover
SongGeneration 2
tencent-ailab
Music Generation
wiro/video-background-music-v2 model cover
video-background-music-v2
wiro
Video to Video
ace-step/text-to-song-ACE-Step1.5 model cover
text-to-song-ACE-Step1.5
ace-step
Music Generation
wiro/Song Frame model cover
Song Frame
wiro
Image to Video
Faceless-Video-Generator
wiro
Text to Video
video-background-music-gen
wiro
Video to Video
ace-step/image-to-song-ACE-Step-v1-3.5B model cover
image-to-song-ACE-Step-v1-3.5B
ace-step
Music Generation
ace-step/text-to-song-ACE-Step-v1-3.5B model cover
text-to-song-ACE-Step-v1-3.5B
ace-step
Music Generation
wiro/image-to-song-with-reference-YuE model cover
image-to-song-with-reference-YuE
wiro
Music Generation
wiro/image-to-song-YuE model cover
image-to-song-YuE
wiro
Music Generation
wiro/text-to-song-with-reference-YuE model cover
text-to-song-with-reference-YuE
wiro
Music Generation
wiro/text-to-song-YuE model cover
text-to-song-YuE
wiro
Music Generation
wiro/music_gen model cover
music_gen
wiro
Music Generation

Realtime Stream

10 models in this collection.

View all
openai/gpt-realtime-translate model cover
gpt-realtime-translate
openai
Speech to Speech
openai/gpt-realtime-2.1-mini model cover
gpt-realtime-2.1-mini
openai
Speech to Speech
openai/gpt-realtime-2.1 model cover
gpt-realtime-2.1
openai
Speech to Speech
openbmb/VoxCPM2 model cover
VoxCPM2
openbmb
Text to Speech
mistralai/Voxtral-Mini-4B-Realtime-2602 model cover
Voxtral-Mini-4B-Realtime-2602
mistralai
Speech to Text
openmoss/MOSS-TTS-Realtime model cover
MOSS-TTS-Realtime
openmoss
Text to Speech
elevenlabs/Realtime Conversational AI model cover
Realtime Conversational AI
elevenlabs
Speech to Speech
openai/gpt-realtime-mini model cover
gpt-realtime-mini
openai
Speech to Speech
openai/gpt-realtime model cover
gpt-realtime
openai
Speech to Speech
nvidia/PersonaPlex-Realtime model cover
PersonaPlex-Realtime
nvidia
Speech to Speech

LLM & Chat

99 models in this collection.

View all
openai/gpt-6-astra model cover
gpt-6-astra
openai
LLM
glm/5.2 model cover
5.2
glm
Partner LLM
deepseek/v4-flash model cover
v4-flash
deepseek
Partner LLM
deepseek/v4-pro model cover
v4-pro
deepseek
Partner LLM
google/gemini-3.8-flash model cover
gemini-3.8-flash
google
Partner LLM
unsloth/Qwen3.8-Flash-Next-GGUF model cover
Qwen3.8-Flash-Next-GGUF
unsloth
Chat
qwen/Qwen3.8-27B-Obliterated model cover
Qwen3.8-27B-Obliterated
qwen
Chat
claude/fable-5 model cover
fable-5
claude
Partner LLM
claude/opus-5 model cover
opus-5
claude
Partner LLM
claude/sonnet-5 model cover
sonnet-5
claude
Partner LLM
google/gemini-3.7-flash model cover
gemini-3.7-flash
google
Partner LLM
qwen/Qwen3.8-27B model cover
Qwen3.8-27B
qwen
Chat
qwen/Qwen3.8-27B-Uncensored model cover
Qwen3.8-27B-Uncensored
qwen
Chat
bytedance/seed-v2.1-turbo model cover
seed-v2.1-turbo
bytedance
Partner LLM
bytedance/seed-v2-pro model cover
seed-v2-pro
bytedance
Partner LLM
bytedance/seed-v2-pro-uncensored model cover
seed-v2-pro-uncensored
bytedance
Partner LLM
bytedance/seed-v2.1-turbo-uncensored model cover
seed-v2.1-turbo-uncensored
bytedance
Partner LLM
google/gemini-3.5-flash-lite model cover
gemini-3.5-flash-lite
google
Partner LLM
google/gemini-3.6-flash model cover
gemini-3.6-flash
google
Partner LLM
openai/gpt-5-6-luna model cover
gpt-5-6-luna
openai
LLM

AI Models for E-commerce

17 models in this collection.

View all
ugc creator v2
wiro
Image to Video
briaai/RMBG-2.0 model cover
RMBG-2.0
briaai
Image to Image
ugc creator
wiro
Image to Video
wiro/Shopify Template model cover
Shopify Template
wiro
Image to Image
Product Studio
wiro
Image to Video
Product with Model
wiro
Image to Video
wiro/Virtual Try-On-V2 model cover
Virtual Try-On-V2
wiro
Image to Video
Animated Logo
wiro
Image to Video
3D Text Animations
wiro
Text to Video
Product Ads with Caption
wiro
Image to Video
Product Ads with Logo
wiro
Image to Video
Product Ads
wiro
Image to Video
wiro/camera-angle-editor model cover
camera-angle-editor
wiro
Image to Image
wiro/Product Photoshoot model cover
Product Photoshoot
wiro
Image to Image
wiro/Virtual Try-On model cover
Virtual Try-On
wiro
Image to Image
wiro/text-removal model cover
text-removal
wiro
Image to Image
wiro/remove-background model cover
remove-background
wiro
Image to Image

AI Models for Social Media Creators

66 models in this collection.

View all
Hair Cut Effects
wiro
Image to Video
Airplane Wing Stunt Trend
wiro
Image to Video
ugc creator v2
wiro
Image to Video
Hero Vox Effects
wiro
Image to Video
Birthday Mode Effects
wiro
Image to Video
Birthday Mode Effects with Caption
wiro
Image to Video
Hair Style Effects with Caption
wiro
Image to Video
Euphoria Effects
wiro
Image to Video
World Cup 2026 Effects
wiro
Image to Video
World Cup 2026 Effects with Caption
wiro
Image to Video
Scream Effects
wiro
Image to Video
panini-card
wiro
Image to Video
Sport Trend Effects
wiro
Image to Video
Insta Hot Girl Effects
wiro
Image to Video
Queer Editorial Effects
wiro
Image to Video
Tiktok Trend Effects
wiro
Image to Video
Wildlife Documentary Effect
wiro
Image to Video
Transformation Effect
wiro
Image to Video
Supernatural Presence Effect
wiro
Image to Video
Superhero Powers Effect
wiro
Image to Video
Built for Production

Everything you need
to ship with confidence.

Wiro is the production layer between your product and every model.

One unified API

Same auth, same format, every model.

Transparent pricing

Know your costs before you run.

Run history & files

Inputs, outputs, and cost logged.

Webhooks & workflows

Automate completions and build at scale.

Explore the full catalog

Hundreds of models.
One place to find them.

Filter by modality, provider, pricing, context, and more.

Advanced filters

Pinpoint the right model.

Compare side by side

Evaluate quality, cost, and performance.

Save and organize

Bookmark favorites and build collections.

View all models

Discover, test, and run AI models, build workflows and agents with one unified API.

All systems operational
WiroAboutBlogCareersContact
ProductModelsAgentsPricingPartnerChangelogStatusFAQ
Getting StartedIntroductionAuthenticationProjectsCode ExamplesWiro MCP ServerSelf-Hosted MCPn8n IntegrationLLMs.txt
API ReferenceModelsRun a ModelModel ParametersTasksLLM & Chat StreamingWebSocketRealtime VoiceFiles
© 2026 Wiro AI. All rights reserved.
PrivacyTermsData Deletion