Try MiniMax H3 (Text-to-Video) (Image-to-Video) from MiniMax →
Models
Agents
WorkflowsStudioPricingBlogDocs
ExploreDiscover models by categoryBrowse All ModelsBrowse the complete catalogSee FavoritesSign in to view saved models
OverviewThe platform at a glanceLearnSkills, knowledge, guardrailsAnatomyWhat makes agents reasonBuild Your AgentPick skills, set tier, deploy
Pre-built AgentsBrowse the catalog
Agent Usecases
Ad Campaign ManagerApp Event ManagerApp Review RepliesBarber BookingCustomer Win-BackEcommerce ListingsRestaurant Reviews
Sign InStart Building

Task History

Click to see output list

No tasks yet

Go to Models
Explore models/
Audio & SpeechActive

OpenMOSS / MOSS-TTSD

MOSS-TTSD

byopenmoss

MOSS-TTSD is a production long-form dialogue model for expressive multi-speaker conversational audio at scale. It supports long-duration continuity, turn-taking control, and zero-shot voice cloning from short references for podcasts, audiobooks, commentary, dubbing, and entertainment dialogue.

Text to SpeechFast Inference
Model ID
MOSS-TTSD
Provider
openmoss
Updated
1771416975
MOSS-TTSD
0
Comments
Average rating : 0 (0 users)
Provideropenmoss
ModelMOSS-TTSD
Text to SpeechFast Inference
wiro playground—openmoss/MOSS-TTSD
Reset to defaults

The text to generate audio from.

0 / 1
Maximum 1 audio allowed
Drop audio to upload

OR

Click to browse your device

Supports: MP3, WAV, M4A, WEBM, OPUS

Speaker 1 reference audio for voice cloning (optional). Must be provided together with Speaker 1 Text. Supported formats: .wav, .mp3, .m4a and .webm.

Speaker 1 prompt text for voice cloning (optional). Must be provided together with Speaker 1 Reference Audio. The [S1] tag is added automatically if missing.

0 / 1
Maximum 1 audio allowed
Drop audio to upload

OR

Click to browse your device

Supports: MP3, WAV, M4A, WEBM, OPUS

Speaker 2 reference audio for voice cloning (optional). Must be provided together with Speaker 2 Text. Supported formats: .wav, .mp3, .m4a and .webm.

Speaker 2 prompt text for voice cloning (optional). Must be provided together with Speaker 2 Reference Audio. The [S2] tag is added automatically if missing.

Sample outputs
openmoss-moss-ttsd-sample-1.mp3
openmoss-moss-ttsd-sample-2.mp3
Updated 1771416975

# MOSS-TTSD (Multi-Speaker Dialogue TTS)

Long-form spoken dialogue generation model with multi-party conversational speech synthesis. Part of the MOSS-TTS Family — an open-source speech and sound generation model family from MOSI.AI and the OpenMOSS team. MOSS-TTSD supports 1-5 speakers with flexible control, handles natural turn-taking, and features zero-shot voice cloning with only 3-10 seconds of reference audio. Supports Chinese, English, Japanese, and European languages (auto-detected from input text).

dialogue (string, required) — Dialogue text with speaker tags. Use [S1], [S2] etc. to indicate speaker turns. Example: "[S1] Hey, did you finish the report? [S2] Almost done, just need to review the numbers."

s1Audio (file, optional) — Speaker 1 reference audio for voice cloning. Must be provided together with Speaker 1 Text. Supports .wav, .mp3, .m4a and .webm formats.

s1Text (string, optional) — Speaker 1 prompt transcript for voice cloning. Must be provided together with Speaker 1 Audio. The [S1] tag is added automatically if missing.

s2Audio (file, optional) — Speaker 2 reference audio for voice cloning. Must be provided together with Speaker 2 Text. Supports .wav, .mp3, .m4a and .webm formats.

s2Text (string, optional) — Speaker 2 prompt transcript for voice cloning. Must be provided together with Speaker 2 Audio. The [S2] tag is added automatically if missing.

s3Audio (file, optional) — Speaker 3 reference audio for voice cloning. Must be provided together with Speaker 3 Text. Supports .wav, .mp3, .m4a and .webm formats.

s3Text (string, optional) — Speaker 3 prompt transcript for voice cloning. Must be provided together with Speaker 3 Audio. The [S3] tag is added automatically if missing.

s4Audio (file, optional) — Speaker 4 reference audio for voice cloning. Must be provided together with Speaker 4 Text. Supports .wav, .mp3, .m4a and .webm formats.

s4Text (string, optional) — Speaker 4 prompt transcript for voice cloning. Must be provided together with Speaker 4 Audio. The [S4] tag is added automatically if missing.

s5Audio (file, optional) — Speaker 5 reference audio for voice cloning. Must be provided together with Speaker 5 Text. Supports .wav, .mp3, .m4a and .webm formats.

s5Text (string, optional) — Speaker 5 prompt transcript for voice cloning. Must be provided together with Speaker 5 Audio. The [S5] tag is added automatically if missing.

textNormalize (checkbox, optional) — Normalize dialogue text before synthesis. Removes special characters and fixes punctuation for cleaner output.

API quick start

Run MOSS-TTSD with a single API call.

POST https://api.wiro.ai/v1/Run/OpenMOSS/MOSS-TTSD
{
  "dialogue": "[S1] Hey, did you finish the report? [S2]…",
  "inputAudio1": "https://your-cdn.com/input.mp3",
  "inputText1": "...",
  "inputAudio2": "https://your-cdn.com/input.mp3"
}
View full API docs

Discover, test, and run AI models, build workflows and agents with one unified API.

All systems operational
WiroAboutBlogCareersContact
ProductModelsAgentsPricingPartnerChangelogStatusFAQ
Getting StartedIntroductionAuthenticationProjectsCode ExamplesWiro MCP ServerSelf-Hosted MCPn8n IntegrationLLMs.txt
API ReferenceModelsRun a ModelModel ParametersTasksLLM & Chat StreamingWebSocketRealtime VoiceFiles
© 2026 Wiro AI. All rights reserved.
PrivacyTermsData Deletion