Try MiniMax H3 (Text-to-Video) (Image-to-Video) from MiniMax →
Models
Agents
WorkflowsStudioPricingBlogDocs
ExploreDiscover models by categoryBrowse All ModelsBrowse the complete catalogSee FavoritesSign in to view saved models
OverviewThe platform at a glanceLearnSkills, knowledge, guardrailsAnatomyWhat makes agents reasonBuild Your AgentPick skills, set tier, deploy
Pre-built AgentsBrowse the catalog
Agent Usecases
Ad Campaign ManagerApp Event ManagerApp Review RepliesBarber BookingCustomer Win-BackEcommerce ListingsRestaurant Reviews
Sign InStart Building

Task History

Click to see output list

No tasks yet

Go to Models
Explore models/
Audio & SpeechActive

openbmb / VoxCPM2

VoxCPM2

byopenbmb

A real-time, multilingual text-to-speech system offering expressive voice design and high-fidelity voice cloning through low-latency streaming inference.

Text to SpeechRealtime TTSFast InferenceH200
Model ID
VoxCPM2
Provider
openbmb
Updated
1778075045
VoxCPM2
2
Comments
Average rating : 1 (1 users)
Provideropenbmb
ModelVoxCPM2
Text to SpeechRealtime TTSFast InferenceH200
wiro playground—openbmb/VoxCPM2
Reset to defaults

The text to synthesize into speech. Supports 30+ languages. Keep individual runs under ~500 characters for the best stability; very long inputs may lose prosody consistency.

0 / 1
Maximum 1 audio allowed
Drop audio to upload

OR

Click to browse your device

Supports: MP3, WAV, M4A, WEBM, OPUS

Optional voice cloning sample. The model copies only the speaker's timbre, no transcript needed. Recommended 5–15 s of clean speech; clips longer than 30 s are automatically trimmed. Supported: wav, mp3, flac, ogg, m4a, wma, opus.

Sample outputs
openbmb-voxcpm2-sample-1.mp3
openbmb-voxcpm2-sample-2.mp3
openbmb-voxcpm2-sample-3.mp3
Updated 1778075045

# VoxCPM2

A real-time, multilingual text-to-speech system offering expressive voice design and high-fidelity voice cloning through low-latency streaming inference.

## Overview

VoxCPM2 is a 2B-parameter, diffusion-autoregressive speech model that produces natural, 48 kHz audio across 30+ languages. It supports three complementary modes out of the box:

- Voice Design — describe a voice in natural language (e.g. *"warm female voice, mid-thirties, calm"*) and the model synthesises a matching speaker.

- Voice Cloning — provide a short reference recording and the model reproduces that speaker's timbre on arbitrary text.

- Style Continuation — pair a reference recording with its transcript to carry over the speaker's prosody, pacing, and emotion into the output.

## Use Cases

- Real-time voice agents and conversational assistants

- Long-form narration, audiobooks, and podcasts

- Dubbing, localisation, and cross-lingual voiceover

- Character voices for games and interactive media

- Accessibility tools and screen readers

## Highlights

- Faster-than-real-time generation (RTF ≈ 0.30, or ≈ 0.13 with acceleration)

- Clean, broadcast-quality 48 kHz output

- Robust multilingual coverage, including code-switched text

- Built-in denoising for reliable cloning from noisy inputs

- Streaming API that yields audio as it is generated

Example prompts

Great starting points for VoxCPM2.

I received a birthday gift from a friend who sent it from afar. That unexpected surprise and deep blessing filled my heart with sweet happiness, and my smile bloomed like a flower.Audio & Speech
I completely understand the frustration you're experiencing. Technical issues are never convenient. To help me resolve this for you immediately, could you please confirm the last four digits of your account number?Audio & Speech
Well, we did it. We all accomplished one of the major early milestones of our lives: high school graduation. This is a major step in the journey of our lives, one that should be recognized for its immense significance. It is an act not only of personal commitment, but also one of pride. We all worked hard to get to this day, and our work did not go to waste. A high school diploma is a wonderful tool in this world, one that opens many doors of opportunity for anyone who is lucky enough to have one.Audio & Speech

API quick start

Run VoxCPM2 with a single API call.

POST https://api.wiro.ai/v1/Run/openbmb/VoxCPM2
{
  "prompt": "A small robot learned to paint by watchin…",
  "promptAudioText": "...",
  "cfgValue": 2.0,
  "inputAudio": "https://your-cdn.com/input.mp3"
}
View full API docs

Discover, test, and run AI models, build workflows and agents with one unified API.

All systems operational
WiroAboutBlogCareersContact
ProductModelsAgentsPricingPartnerChangelogStatusFAQ
Getting StartedIntroductionAuthenticationProjectsCode ExamplesWiro MCP ServerSelf-Hosted MCPn8n IntegrationLLMs.txt
API ReferenceModelsRun a ModelModel ParametersTasksLLM & Chat StreamingWebSocketRealtime VoiceFiles
© 2026 Wiro AI. All rights reserved.
PrivacyTermsData Deletion