Try MiniMax H3 (Text-to-Video) (Image-to-Video) from MiniMax →
Models
Agents
WorkflowsStudioPricingBlogDocs
ExploreDiscover models by categoryBrowse All ModelsBrowse the complete catalogSee FavoritesSign in to view saved models
OverviewThe platform at a glanceLearnSkills, knowledge, guardrailsAnatomyWhat makes agents reasonBuild Your AgentPick skills, set tier, deploy
Pre-built AgentsBrowse the catalog
Agent Usecases
Ad Campaign ManagerApp Event ManagerApp Review RepliesBarber BookingCustomer Win-BackEcommerce ListingsRestaurant Reviews
Sign InStart Building

Task History

Click to see output list

No tasks yet

Go to Models
Explore models/
Audio & SpeechActive

OpenMOSS / MOSS-TTS-v1.5

MOSS-TTS-v1.5

byopenmoss

OpenMOSS MOSS-TTS v1.5 turns text into natural speech, with optional zero-shot voice cloning from a reference clip. It supports 31 languages plus pause and pronunciation control.

Text to SpeechVoice CloneFast Inference
Model ID
MOSS-TTS-v1.5
Provider
openmoss
Updated
1781766389
MOSS-TTS-v1.5
0
Comments
Average rating : 2 (1 users)
Provideropenmoss
ModelMOSS-TTS-v1.5
Text to SpeechVoice CloneFast Inference
wiro playground—openmoss/MOSS-TTS-v1.5
Reset to defaults

The text to synthesize. Supports Pinyin/IPA, inline pause markers like [pause 3.2s], and [S1]/[S2] speaker tags.

Optional language tag. Recommended in v1.5 for non zh/en text; improves multilingual stability.

0 / 1
Maximum 1 audio allowed
Drop audio to upload

OR

Click to browse your device

Supports: MP3, WAV, M4A, WEBM, OPUS

Optional reference clip for zero-shot voice cloning. If empty, the model picks a voice automatically.

Sample outputs
openmoss-moss-tts-v1-5-sample-1.mp3
Updated 1781766389
## Overview MOSS-TTS-v1.5 is an autoregressive text-to-speech model by OpenMOSS and MOSI.AI. It generates discrete audio tokens, then decodes them into speech audio. It supports zero-shot voice cloning from a short reference clip, so you can keep a consistent speaker without training. It also adds explicit pause control and stronger multilingual stability when you provide a language tag. ## What you can build - Voiceovers for product demos, explainers, and slides - Long-form narration, including multi-minute and extended reads - Multilingual IVR prompts and support content across 31 languages - Voice cloning for consistent character voices in apps and games - Pronunciation-controlled learning content using Pinyin or IPA hints - Scripted speech with explicit, timed pauses for pacing ## Inputs - Text to synthesize, provided as plain text. It can mix normal text with Pinyin or IPA for pronunciation control. - Optional inline pause markers embedded in the text, using the form `[pause X.Ys]`. - Optional language tag you select to guide multilingual synthesis. This is recommended for languages beyond Chinese and English. - Optional reference audio clip for zero-shot voice cloning. Provide a single clip when you want speaker similarity. - Optional target duration control, provided as an expected count of audio tokens. The model uses about 12.5 audio tokens per second. - Optional generation length cap, set as a maximum number of generated tokens. Raise it for long passages. - Optional decoding controls for text planning, such as temperature and sampling filters. Lower values make results more deterministic. - Optional decoding controls for acoustic generation, such as temperature and sampling filters. Higher values can add variety but may reduce stability. - Optional anti-repetition control for audio tokens. Slight increases can reduce stutters and loops. ## Outputs The model returns synthesized speech audio as a WAV result. The audio content is the spoken rendition of your input text, including any explicit pauses you inserted. If you provide reference audio, the output aims to match the reference speaker’s timbre and some speaking style. When you use duration control, the output length shifts toward your target token count. ## Recommended settings - General TTS: use an audio temperature of 1.7, top-p of 0.8, top-k of 25, and an audio repetition penalty of 1.0. - Multilingual text: always set the language tag when you know it. - Long passages: increase the generation token cap so the model can finish the full script. - Voice cloning: use a clean reference clip with one speaker and minimal background noise. ## Limitations - The model is sensitive to decoding settings. Aggressive sampling can cause instability or artifacts. - Voice cloning quality depends heavily on the reference clip. Noise, music, or multiple speakers reduce similarity. - Very long texts need a higher generation token cap. Otherwise, audio may stop early. - Language tags matter for non-English and non-Chinese text. Missing tags can reduce stability. - Messy inputs can hurt prosody. This includes unpunctuated text, inconsistent casing, or malformed pause markers. ## Safety & compliance Models in the MOSS-TTS family are released under the Apache License 2.0. Voice cloning can enable impersonation. Only clone voices you own or have explicit permission to use. Don’t use generated audio to mislead people about identity, consent, or endorsements.

API quick start

Run MOSS-TTS-v1.5 with a single API call.

POST https://api.wiro.ai/v1/Run/OpenMOSS/MOSS-TTS-v1.5
{
  "prompt": "Genuine love from a devoted man transform…",
  "language": "None",
  "inputAudio": "https://your-cdn.com/input.mp3",
  "maxNewTokens": 4096
}
View full API docs

Discover, test, and run AI models, build workflows and agents with one unified API.

All systems operational
WiroAboutBlogCareersContact
ProductModelsAgentsPricingPartnerChangelogStatusFAQ
Getting StartedIntroductionAuthenticationProjectsCode ExamplesWiro MCP ServerSelf-Hosted MCPn8n IntegrationLLMs.txt
API ReferenceModelsRun a ModelModel ParametersTasksLLM & Chat StreamingWebSocketRealtime VoiceFiles
© 2026 Wiro AI. All rights reserved.
PrivacyTermsData Deletion