Try MiniMax H3 (Text-to-Video) (Image-to-Video) from MiniMax →
Models
Agents
WorkflowsStudioPricingBlogDocs
ExploreDiscover models by categoryBrowse All ModelsBrowse the complete catalogSee FavoritesSign in to view saved models
OverviewThe platform at a glanceLearnSkills, knowledge, guardrailsAnatomyWhat makes agents reasonBuild Your AgentPick skills, set tier, deploy
Pre-built AgentsBrowse the catalog
Agent Usecases
Ad Campaign ManagerApp Event ManagerApp Review RepliesBarber BookingCustomer Win-BackEcommerce ListingsRestaurant Reviews
Sign InStart Building

Task History

Click to see output list

No tasks yet

Go to Models
DISCOVER AI MODELS

Search the full model catalog.

Discover and compare production-ready models. Filter by modality, provider, pricing, and production fit.

26 models2 active filters

Active filters

Wiro AI

Video Generation

Image Generation

Audio & Speech

2

Realtime Stream

Music Generation

3D Generation

LLM & Chat

Workflow

Partners

Need help choosing?

Compare models by capability, cost, and provider fit.

Showing 20 of 26 models
MOSS-TTS-v1.5 - AI model cover image

MOSS-TTS-v1.5

OpenMOSS

OpenMOSS MOSS-TTS v1.5 turns text into natural speech, with optional zero-shot voice cloning from a reference clip. It supports 31 languages plus pause and pronunciation control.

Text to Speechtext-to-speechtransformers
Run
VoxCPM2 - AI model cover image

VoxCPM2

openbmb

A real-time, multilingual text-to-speech system offering expressive voice design and high-fidelity voice cloning through low-latency streaming inference.

Text to Speechtext-to-speechstreaming text input
Run
OmniVoice - AI model cover image

OmniVoice

k2-fsa

OmniVoice by k2-fsa generates 24 kHz speech from text in 600+ languages. Clone a speaker from a short reference clip or design a new voice from attributes.

Text to Speechtext-to-speechtransformers
Run
tada-3b-ml - AI model cover image

tada-3b-ml

humeai

TADA is a unified speech-language model that synchronizes speech and text into a single, cohesive stream via 1:1 alignment. By leveraging a novel tokenizer and architectural design, TADA achieves high-fidelity synthesis and generation with a fraction of the computational overhead required by traditional models.

Text to Speechtext-to-speechtransformers
Run
s2-pro - AI model cover image

s2-pro

fishaudio

Generates high-quality speech from text using advanced TTS technology with support for voice cloning and multi-speaker synthesis.

Text to Speechtext-to-speechtransformers
Run
kani-tts-2-en - AI model cover image

kani-tts-2-en

nineninesix

Generates natural-sounding speech from text with support for multi-speaker voice cloning and fast inference capabilities.

Text to Speechtext-to-speechtransformers
Run
chatterbox-turbo - AI model cover image

chatterbox-turbo

resemble-ai

The fastest open source TTS model without sacrificing quality.

Speech to Speechspeech-to-textspeech-to-speech
Run
chatterbox-multilingual  - AI model cover image

chatterbox-multilingual

resemble-ai

Generate expressive, natural speech in 23 languages. Features instant voice cloning from short audio, emotion control, and seamless cross-language voice transfer.

Speech to Speechspeech-to-textspeech-to-speech
Run
MOSS-TTSD - AI model cover image

MOSS-TTSD

OpenMOSS

MOSS-TTSD is a production long-form dialogue model for expressive multi-speaker conversational audio at scale. It supports long-duration continuity, turn-taking control, and zero-shot voice cloning from short references for podcasts, audiobooks, commentary, dubbing, and entertainment dialogue.

Text to Speechtext-to-speechtransformers
Run
MOSS-TTS-Realtime - AI model cover image

MOSS-TTS-Realtime

OpenMOSS

Real-time streaming text-to-speech with zero-shot voice cloning. Supports 20 languages including English, Chinese, Japanese, Korean, and more. Audio starts playing immediately — no waiting for full generation. Clone any voice from a short reference clip.

Text to Speechtext-to-speechtransformers
Run
Realtime Conversational AI - AI model cover image

Realtime Conversational AI

elevenlabs

A real-time voice conversation tool using ElevenLabs' AI voice agents. Customize voices, behaviors, and languages for interactive AI experiences.

Speech to Speech
Run
gpt-realtime-mini - AI model cover image

gpt-realtime-mini

openai

GPT Mini Realtime enables low-latency, bidirectional streaming for voice and text. Build interactive, responsive AI experiences that feel natural and immediate.

Speech to Speech
Run
gpt-realtime - AI model cover image

gpt-realtime

openai

GPT Realtime enables low-latency, bidirectional streaming for voice and text. Build interactive, responsive AI experiences that feel natural and immediate.

Speech to Speech
Run
Qwen3-TTS-12Hz-1.7B - AI model cover image

Qwen3-TTS-12Hz-1.7B

Qwen

A fast inference text-to-speech model optimized for real-time audio generation with multi-language support.

$0.0107Text to Speechtext-to-speech
Run
PersonaPlex-Realtime - AI model cover image

PersonaPlex-Realtime

nvidia

Convert speech to speech with customizable voices using PersonaPlex. Supports various audio formats and offers control over text temperature and audio top K settings.

Speech to Speechspeech-to-textaudio
Run
VibeVoice-Realtime - AI model cover image

VibeVoice-Realtime

microsoft

VibeVoice-Realtime is a lightweight real‑time text-to-speech model supporting streaming text input and robust long-form speech generation.

$0.0156Text to Speechtext-to-speech
Run
VoxCPM - AI model cover image

VoxCPM

openbmb

Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning

Speech to Speechspeech-to-textspeech-to-speech
Run
gemini-2.5-tts - AI model cover image

gemini-2.5-tts

google

Google's Gemini 2.5 Flash Text To Speech Preview model

Text to Speechgoogle
Run
text-to-speech - AI model cover image

text-to-speech

elevenlabs

Text to speech model from ElevenLabs

Text to Speechelevenlabstext-to-speech
Run

Faceless-Video-Generator

wiro

Create professional short videos (30s) for YouTube Shorts, Instagram Reels, TikTok, and X (Twitter) – from a single prompt. Automatically generate speech, captions, and optional talking head avatars using AI. Perfect for content creators, marketers, and educators looking to grow faster with less effort.

Text to Speechsocial media videosreels generator
Run
Showing 20 of 26 models

Discover, test, and run AI models, build workflows and agents with one unified API.

All systems operational
WiroAboutBlogCareersContact
ProductModelsAgentsPricingPartnerChangelogStatusFAQ
Getting StartedIntroductionAuthenticationProjectsCode ExamplesWiro MCP ServerSelf-Hosted MCPn8n IntegrationLLMs.txt
API ReferenceModelsRun a ModelModel ParametersTasksLLM & Chat StreamingWebSocketRealtime VoiceFiles
© 2026 Wiro AI. All rights reserved.
PrivacyTermsData Deletion