Production-ready models,organized for real work.
Browse and compare image, video, audio, and language models from leading providers, all available through the Wiro API.



Recently Added
522 models in this collection.
Transform a single portrait into a viral cut-your-own-hair video. Say the words, make the first snip, and reveal a perf…
Transform a single portrait into a high-altitude wing-walking stunt video. Step out onto the wing of a plane in flight,…

OpenAI’s GPT 6 Astra handles long-context reasoning with optional image input and tool calls. Build agents for coding,…

Burn captions into a video with TikTok-style timing. Transcribe spoken audio word-by-word or overlay fixed text, with f…

Generate natural speech from text with Gemini 3.1 Flash TTS. Use voice options and expressive tags to control tone and…

GLM 5.2 is a text-only MoE model with a usable 1M-token context for long-horizon coding and agent workflows. It support…

DeepSeek V4 Flash is a fast MoE text model with a 1M-token context window and optional thinking mode. It’s built for ch…

DeepSeek V4 Pro (v4-pro) is a long-context text model with optional thinking mode for agentic coding, math reasoning, a…

Google’s Gemini 3.8 Flash reasons over text, images, audio, video, and PDFs, then returns strong text answers for long-…

Create new images or edit a source photo with xAI’s Grok Imagine Image V2. Choose 1K or 2K output, quality level, and a…
Pruna P Video Edit edits a short MP4 clip from one text instruction, with 1-4 reference images to guide identity or sty…

Re-voice an existing recording with a chosen ElevenLabs voice while keeping the original words, timing, and delivery. E…

Create original music from a plain-English brief with ElevenLabs Music v2. Choose track length, MP3 quality, and option…
FastH3 is a four-forward distilled MiniMax H3 model that generates synchronized video and stereo audio from a text prom…

Run Unsloth’s GGUF quantizations of Qwen 3.8 Flash Next for long-context chat and reasoning. Includes optional thinking…

Refusal-reduced variant of Qwen 3.8 27B for long-context chat and coding. It can emit or hide thinking traces and suppo…

Claude Fable 5 is a frontier text model with vision and PDF support. It handles 1M-token context for deep reasoning, co…

Claude Opus 5 is a premium text model with vision for agentic coding and document work. Upload images or PDFs and get l…

Claude Sonnet 5 is an agent-ready text model for coding, reasoning, and long-context work. It can read images and PDFs…

OpenAI’s GPT Realtime Whisper turns live audio into streaming transcript updates. Adjust delay levels to trade latency…
Popular Models
522 models in this collection.

Generate high-resolution images using Seedream v4.5 Uncensored. Supports text-to-image and image-to-image transformatio…
Generates videos from images and reference videos with motion control. Supports custom prompts and character orientatio…
Generate high-quality videos from text prompts using Kling V3. Supports custom frames, duration, and aspect ratios.

An image editing tool designed for quick transformations using reference images and prompts. Supports multi-image mixin…
Seedance Pro v1.5 Uncensored by ByteDance generates short videos from text with optional native audio, strong prompt fo…
Generate 3–10s videos from reference images, or edit an existing clip with one instruction. Built on Google Gemini Omni…
Generate 5s,10s 720p MP4 videos from text, or animate a still image as the opening frame. Google Gemini Omni Flash supp…

Google’s Nano Banana 2 Lite generates and edits 1K images from text, with optional multi-image references. It’s built f…

Google's Gemini 3 Pro Image Preview, also known as Nano Banana, model for text-to-image and image-to-image generation.

Chat model based on Qwen 3.8 27B with refusal behavior reduced through direction removal. Its 262k context fits long do…
Create 4–30s videos from a description, or steer motion using first and last frame images. Choose 480p or 720p output w…

Refusal-reduced variant of Qwen 3.8 27B for long-context chat and coding. It can emit or hide thinking traces and suppo…

Run Unsloth’s GGUF quantizations of Qwen 3.8 Flash Next for long-context chat and reasoning. Includes optional thinking…

Convert a video to MP4, MOV, WebM, MKV, AVI, MPEG, or M4V with adjustable compression. Built by Wiro for clean exports…

Convert an image to JPEG, PNG, WebP, TIFF, or AVIF. Set output quality from 0 to 100 to balance file size and visual fi…
Turn a selfie into a Panini-style player card video. Enter name, height, weight, and birth date, then choose a country…

Create or edit images with OpenAI GPT Image 2 using custom pixel sizes, quality tiers, and format controls. Add optiona…

Smart Resize by Wiro converts one image into multiple exact sizes, using AI recomposition to keep key subjects in frame…
Create 4–30 second videos guided by reference clips, images, and audio. Preserve motion and timing while changing subje…

Generate or edit images with GPT Image 2 from OpenAI. It delivers strong instruction following, sharp text rendering, a…
Generate Videos
165 models in this collection.
Transform a single portrait into a viral cut-your-own-hair video. Say the words, make the first snip, and reveal a perf…
Transform a single portrait into a high-altitude wing-walking stunt video. Step out onto the wing of a plane in flight,…

Burn captions into a video with TikTok-style timing. Transcribe spoken audio word-by-word or overlay fixed text, with f…
Pruna P Video Edit edits a short MP4 clip from one text instruction, with 1-4 reference images to guide identity or sty…
FastH3 is a four-forward distilled MiniMax H3 model that generates synchronized video and stereo audio from a text prom…
Create 4–30s videos from a description, or steer motion using first and last frame images. Choose 480p or 720p output w…
Create 4–30 second videos guided by reference clips, images, and audio. Preserve motion and timing while changing subje…
Create 4–30 second videos from a prompt and optional first and last frame images, with optional synced audio. Powered b…
MiniMax H3 R2V generates 4–15s videos with synced stereo audio from reference images, clips, and optional audio. Use it…
MiniMax H3 I2V generates 4–15s videos from a prompt and optional first or last frame. It outputs MP4 video at 24 FPS wi…
Extend a source video past its final frame. FLUX 3 V2V generates 5–20 seconds of 720p or 1080p footage, with optional s…
FLUX 3 I2V turns 1–10 images and a prompt into a 5–20s MP4 clip at 720p or 1080p, with optional native audio and keyfra…
Create UGC-style product ad videos from a single product photo and a short script. Choose a preset scene, duration, asp…
Runway Gen-4.5 turns text or a first-frame image into 2 to 10s 720p video clips. It follows sequenced actions and camer…
Generate 3–10s videos from reference images, or edit an existing clip with one instruction. Built on Google Gemini Omni…
Generate 5s,10s 720p MP4 videos from text, or animate a still image as the opening frame. Google Gemini Omni Flash supp…
ByteDance’s Seedance 2.0 Mini V2V edits short clips from 1–3 reference videos and a prompt. It outputs 480p or 720p MP4…
Seedance 2.0 Mini by ByteDance generates short MP4 videos from a prompt, with optional start and end frames or referenc…
Kling V3 Turbo turns a prompt, or a still image plus motion instructions, into a short MP4 video with native audio and…
Create a short talking-portrait video from a human photo and a script. Choose an effect preset, aspect ratio, and 5, 10…
Generate Images
86 models in this collection.

Create new images or edit a source photo with xAI’s Grok Imagine Image V2. Choose 1K or 2K output, quality level, and a…

Generate 1K or 2K images with strong typography and reference-guided edits. Built on ByteDance Seedream 5.0 Pro for pos…

Pruna’s P Image Ideogram Custom generates poster-ready images from a short prompt. It’s tuned for readable text, typogr…

P-Image Ideogram by Pruna generates high-quality text-to-image outputs with strong typography, offering 1K or 2K resolu…

ByteDance Seedream V5 Pro generates images from text and edits a base image using up to 10 references. It’s tuned for i…

Google’s Nano Banana 2 Lite generates and edits 1K images from text, with optional multi-image references. It’s built f…

A foundational text-to-image model designed for efficient, high-resolution image generation with strong prompt followin…

Create or edit images with OpenAI GPT Image 2 using custom pixel sizes, quality tiers, and format controls. Add optiona…

SenseNova U1-8B turns text into high-detail images with strong layout control and clearer in-image text. Use it for pos…

Generate or edit images with GPT Image 2 from OpenAI. It delivers strong instruction following, sharp text rendering, a…

xAI’s Grok Imagine Image generates images from prompts and can edit a source image using instructions. Pick an aspect r…

Generate or edit images using text prompts or image edits with GPT Image 1.5. Supports multiple sizes, formats, and qua…

Generate high-resolution images using Seedream v4.5 Uncensored. Supports text-to-image and image-to-image transformatio…

FireRed-Image-Edit-1.1 significantly enhances identity consistency, multi-image conditioning, and domain-specialized ed…

Generate high-quality images from text prompts or image inputs using the Seedream v5 Lite Uncensored model. Supports mu…

An image editing tool designed for quick transformations using reference images and prompts. Supports multi-image mixin…

Generate high-quality images using Seedream V5 Lite, supporting both image-to-image and text-to-image transformations w…

FireRed-Image-Edit is a general-purpose image editing model that delivers high-fidelity and consistent editing across a…

GLM-Image is an image generation model adopts a hybrid autoregressive + diffusion decoder architecture. In general imag…

FLUX.2 [klein] 9B Base is a 9 billion parameter rectified flow transformer capable of generating images from text descr…
Edit Images
81 models in this collection.

Create new images or edit a source photo with xAI’s Grok Imagine Image V2. Choose 1K or 2K output, quality level, and a…

Generate 1K or 2K images with strong typography and reference-guided edits. Built on ByteDance Seedream 5.0 Pro for pos…

ByteDance Seedream V5 Pro generates images from text and edits a base image using up to 10 references. It’s tuned for i…

RMBG-2.0 removes image backgrounds with an 8-bit alpha matte for smooth edges. Built by BRIA AI for production cutouts…

Google’s Nano Banana 2 Lite generates and edits 1K images from text, with optional multi-image references. It’s built f…

Remove image backgrounds and get a clean transparent PNG cutout. Ideogram’s generative matting keeps hair, glass, and f…
Turn your photo into a fun birthday celebration while keeping your face natural and unchanged for instantly share-ready…
Turn your photo into a fun birthday celebration with your age featured on the scene, while keeping your face natural an…

Pruna’s p-image-try-on creates virtual try-on images by applying one or more garment photos to a person photo. Generate…

LocateAnything 3B by Nvidia localizes objects, UI elements, and text in images. It returns normalized boxes or points f…

Convert an image to JPEG, PNG, WebP, TIFF, or AVIF. Set output quality from 0 to 100 to balance file size and visual fi…

Smart Resize by Wiro converts one image into multiple exact sizes, using AI recomposition to keep key subjects in frame…

u1-8b-interleave by SenseNova generates step-by-step text with matching images, optionally conditioned on up to 5 refer…

Create or edit images with OpenAI GPT Image 2 using custom pixel sizes, quality tiers, and format controls. Add optiona…

Generate or edit images with GPT Image 2 from OpenAI. It delivers strong instruction following, sharp text rendering, a…

xAI’s Grok Imagine Image generates images from prompts and can edit a source image using instructions. Pick an aspect r…

Generate or edit images using text prompts or image edits with GPT Image 1.5. Supports multiple sizes, formats, and qua…

Generate stylish Instagram-style pose images with trendy angles, natural expressions, and a modern aesthetic. Built by…

Generate high-resolution images using Seedream v4.5 Uncensored. Supports text-to-image and image-to-image transformatio…

FireRed-Image-Edit-1.1 significantly enhances identity consistency, multi-image conditioning, and domain-specialized ed…
Generate Text
28 models in this collection.

u1-8b-interleave by SenseNova generates step-by-step text with matching images, optionally conditioned on up to 5 refer…

SenseNova U1 8B Visual Understanding creates infographic-style images and prompt-based edits from text and an optional…

Multilingual speech-to-text for 25 European languages with auto language detection, punctuation, capitalization, and op…

CohereLabs cohere-transcribe-03-2026 is a 2B Conformer speech-to-text model for 14 languages. It creates accurate trans…

A lightweight speech-to-text model optimized for fast inference. Converts audio input into text with support for multip…

Voxtral Mini 4B Realtime 2602 is a multilingual, realtime speech-transcription model and among the first open-source so…

Nemotron-Speech-Streaming-En-0.6b is the first unified model in the Nemotron Speech family, engineered to deliver high-…

Speech to text model from ElevenLabs

OpenAI Whisper Medium transcribes speech to text with timestamps and optional speaker diarization. Works across many la…

Moondream3 is a cutting-edge vision-language model that delivers advanced visual reasoning with built-in object detecti…

Moondream3 is a cutting-edge vision-language model that delivers advanced visual reasoning with built-in object detecti…

Moondream3 is a cutting-edge vision-language model that delivers advanced visual reasoning with built-in object detecti…

Moondream3 is a cutting-edge vision-language model that delivers advanced visual reasoning with built-in object detecti…

Whisper is a pre-trained model for automatic speech recognition (ASR) and speech translation. Trained on 680k hours of…

NSFW video detection automatically analyzes video content to identify inappropriate or explicit material, ensuring comp…
VideoLLaMA3-2B is a model designed for video understanding.
VideoLLaMA3-2B-Image is a model designed for image understanding.

VideoLLaMA3-7B-Image is a model designed for image understanding.

VideoLLaMA3-7B is a model designed for video understanding.

BLIP-2 creates captions or detailed descriptions for images. This is BLIP-2 model, leveraging Flan T5-xl.
Generate 3D
4 models in this collection.
Generate Audio
31 models in this collection.




















Generate Music
18 models in this collection.
















Realtime Stream
10 models in this collection.










LLM & Chat
99 models in this collection.




















AI Models for E-commerce
17 models in this collection.








AI Models for Social Media Creators
66 models in this collection.
Everything you need
to ship with confidence.
Wiro is the production layer between your product and every model.
Same auth, same format, every model.
Know your costs before you run.
Inputs, outputs, and cost logged.
Automate completions and build at scale.
Hundreds of models.
One place to find them.
Filter by modality, provider, pricing, context, and more.
Pinpoint the right model.
Evaluate quality, cost, and performance.
Bookmark favorites and build collections.


