Try MiniMax H3 (Text-to-Video) (Image-to-Video) from MiniMax →
Models
Agents
WorkflowsStudioPricingBlogDocs
ExploreDiscover models by categoryBrowse All ModelsBrowse the complete catalogSee FavoritesSign in to view saved models
OverviewThe platform at a glanceLearnSkills, knowledge, guardrailsAnatomyWhat makes agents reasonBuild Your AgentPick skills, set tier, deploy
Pre-built AgentsBrowse the catalog
Agent Usecases
Ad Campaign ManagerApp Event ManagerApp Review RepliesBarber BookingCustomer Win-BackEcommerce ListingsRestaurant Reviews
Sign InStart Building

Task History

Click to see output list

No tasks yet

Go to Models
DISCOVER AI MODELS

Search the full model catalog.

Discover and compare production-ready models. Filter by modality, provider, pricing, and production fit.

7 models1 active filters

Active filters

Wiro AI

Video Generation

Image Generation

Audio & Speech

Realtime Stream

3

Music Generation

3D Generation

LLM & Chat

Workflow

Partners

Need help choosing?

Compare models by capability, cost, and provider fit.

Showing 7 of 7 models
VoxCPM2 - AI model cover image

VoxCPM2

openbmb

A real-time, multilingual text-to-speech system offering expressive voice design and high-fidelity voice cloning through low-latency streaming inference.

Realtime TTStext-to-speechstreaming text input
Run
Voxtral-Mini-4B-Realtime-2602 - AI model cover image

Voxtral-Mini-4B-Realtime-2602

mistralai

Voxtral Mini 4B Realtime 2602 is a multilingual, realtime speech-transcription model and among the first open-source solutions to achieve accuracy comparable to offline systems with a delay of <500ms. It supports 13 languages and outperforms existing open-source baselines across a range of tasks, making it ideal for applications like voice assistants and live subtitling.

Realtime STTtransformersenglish
Run
MOSS-TTS-Realtime - AI model cover image

MOSS-TTS-Realtime

OpenMOSS

Real-time streaming text-to-speech with zero-shot voice cloning. Supports 20 languages including English, Chinese, Japanese, Korean, and more. Audio starts playing immediately — no waiting for full generation. Clone any voice from a short reference clip.

Realtime TTStext-to-speechtransformers
Run
Realtime Conversational AI - AI model cover image

Realtime Conversational AI

elevenlabs

A real-time voice conversation tool using ElevenLabs' AI voice agents. Customize voices, behaviors, and languages for interactive AI experiences.

Realtime Conversation
Run
gpt-realtime-mini - AI model cover image

gpt-realtime-mini

openai

GPT Mini Realtime enables low-latency, bidirectional streaming for voice and text. Build interactive, responsive AI experiences that feel natural and immediate.

Realtime Conversation
Run
gpt-realtime - AI model cover image

gpt-realtime

openai

GPT Realtime enables low-latency, bidirectional streaming for voice and text. Build interactive, responsive AI experiences that feel natural and immediate.

Realtime Conversation
Run
PersonaPlex-Realtime - AI model cover image

PersonaPlex-Realtime

nvidia

Convert speech to speech with customizable voices using PersonaPlex. Supports various audio formats and offers control over text temperature and audio top K settings.

Realtime Conversationspeech-to-textaudio
Run

Discover, test, and run AI models, build workflows and agents with one unified API.

All systems operational
WiroAboutBlogCareersContact
ProductModelsAgentsPricingPartnerChangelogStatusFAQ
Getting StartedIntroductionAuthenticationProjectsCode ExamplesWiro MCP ServerSelf-Hosted MCPn8n IntegrationLLMs.txt
API ReferenceModelsRun a ModelModel ParametersTasksLLM & Chat StreamingWebSocketRealtime VoiceFiles
© 2026 Wiro AI. All rights reserved.
PrivacyTermsData Deletion