Try MiniMax H3 (Text-to-Video) (Image-to-Video) from MiniMax →
Models
Agents
WorkflowsStudioPricingBlogDocs
ExploreDiscover models by categoryBrowse All ModelsBrowse the complete catalogSee FavoritesSign in to view saved models
OverviewThe platform at a glanceLearnSkills, knowledge, guardrailsAnatomyWhat makes agents reasonBuild Your AgentPick skills, set tier, deploy
Pre-built AgentsBrowse the catalog
Agent Usecases
Ad Campaign ManagerApp Event ManagerApp Review RepliesBarber BookingCustomer Win-BackEcommerce ListingsRestaurant Reviews
Sign InStart Building

Task History

Click to see output list

No tasks yet

Go to Models
Explore models/
Audio & SpeechActive

OpenMOSS / MOSS-TTS-Realtime

MOSS-TTS-Realtime

byopenmoss

Real-time streaming text-to-speech with zero-shot voice cloning. Supports 20 languages including English, Chinese, Japanese, Korean, and more. Audio starts playing immediately — no waiting for full generation. Clone any voice from a short reference clip.

Text to SpeechRealtime TTSFast InferenceH200
Model ID
MOSS-TTS-Realtime
Provider
openmoss
Updated
1778067056
MOSS-TTS-Realtime
0
Comments
Average rating : 0 (0 users)
Provideropenmoss
ModelMOSS-TTS-Realtime
Text to SpeechRealtime TTSFast InferenceH200
wiro playground—openmoss/MOSS-TTS-Realtime
Reset to defaults

The text to generate audio from.

0 / 1
Maximum 1 audio allowed
Drop audio to upload

OR

Click to browse your device

Supports: MP3, WAV, M4A, WEBM, OPUS

Reference audio for voice cloning (optional). Supported: wav, mp3, flac, ogg, m4a, wma, opus.

Sample outputs

No samples yet

Run this model to create outputs and build up samples.

Updated 1778067056

## Overview

MOSS-TTS-Realtime is a context-aware, real-time streaming text-to-speech model from the OpenMOSS team. Unlike conventional TTS that generates the entire audio before playback, MOSS-TTS-Realtime starts streaming audio as it generates — you hear the speech within milliseconds, even for long paragraphs. It supports zero-shot voice cloning from a short reference audio, producing natural and expressive speech across 20 languages.

## What you can build

- Voice agents and real-time conversational AI

- Audiobook and podcast narration with custom voices

- Live dubbing and voice-over tools

- Multilingual content creation

- Accessibility tools for text-to-speech

- Voice cloning applications with short audio references

## Key features

- Real-time streaming: Audio starts playing immediately as it generates — no waiting for the full text to be synthesized

- Zero-shot voice cloning: Clone any voice from a short reference audio clip (wav, mp3, flac, ogg, m4a, wma, opus)

- 20 languages: Chinese, English, German, Spanish, French, Japanese, Italian, Hebrew, Korean, Russian, Persian, Arabic, Polish, Portuguese, Czech, Danish, Swedish, Hungarian, Greek, Turkish

- Context-aware synthesis: Models dialogue history across text and acoustics for natural multi-turn conversations

- Long-form stability: Supports up to 32K context (~40 minutes) with consistent voice and prosody

- High expressiveness: Trained on 2.5M+ hours of speech data for highly natural prosody

- 2.3B parameters: Powerful yet efficient architecture based on Qwen3 language model backbone

## Inputs

- Text: The text to synthesize into speech (required)

- Reference Audio: A short audio clip of the voice to clone (optional). If not provided, the model uses its default voice. Supported formats: wav, mp3, flac, ogg, m4a, wma, opus

## Outputs

- Streaming audio: Real-time PCM audio stream at 24kHz — playback begins immediately

- Saved audio file: Complete audio saved as MP3 (or WAV fallback) when the session ends

## How it works

1. Click Run and enter your text

2. Optionally upload a reference audio for voice cloning

3. Audio starts streaming immediately — you hear it as it generates

4. Click End Session to stop early, or let it finish naturally

5. Download the complete audio file after the session ends

## Recommended settings

- For voice cloning, use a clear 5-15 second reference clip with minimal background noise

- Long texts work great — the model streams audio progressively without delay

- The model auto-detects language from the input text

## Limitations

- Voice cloning quality depends on the clarity and length of the reference audio

- Very short reference clips (under 2 seconds) may produce lower quality clones

- Some languages may have lower quality than Chinese and English

- Session ends automatically after the full text is synthesized

## Technical details

- Architecture: Qwen3-based language model + local transformer for audio token generation

- Audio codec: MOSS-Audio-Tokenizer at 24kHz

- Parameters: 2.3B (BF16 precision)

- Training data: 2.5M+ hours single-speaker + 1M+ hours multi-speaker conversational

- License: Apache 2.0

- Source: [OpenMOSS-Team/MOSS-TTS-Realtime](https://huggingface.co/OpenMOSS-Team/MOSS-TTS-Realtime)

API quick start

Run MOSS-TTS-Realtime with a single API call.

POST https://api.wiro.ai/v1/Run/OpenMOSS/MOSS-TTS-Realtime
{
  "text": "A small robot learned to paint by watchin…",
  "inputAudio": "https://your-cdn.com/input.mp3",
  "temperature": 0.8,
  "topP": 0.6
}
View full API docs

Discover, test, and run AI models, build workflows and agents with one unified API.

All systems operational
WiroAboutBlogCareersContact
ProductModelsAgentsPricingPartnerChangelogStatusFAQ
Getting StartedIntroductionAuthenticationProjectsCode ExamplesWiro MCP ServerSelf-Hosted MCPn8n IntegrationLLMs.txt
API ReferenceModelsRun a ModelModel ParametersTasksLLM & Chat StreamingWebSocketRealtime VoiceFiles
© 2026 Wiro AI. All rights reserved.
PrivacyTermsData Deletion