{"id":1638,"date":"2026-03-24T07:51:39","date_gmt":"2026-03-24T07:51:39","guid":{"rendered":"https:\/\/wiro.ai\/blog\/?p=1638"},"modified":"2026-09-27T18:53:25","modified_gmt":"2026-09-27T18:53:25","slug":"moss-ttsd-dialogue-tts-in-6-tests","status":"publish","type":"post","link":"https:\/\/wiro.ai\/blog\/moss-ttsd-dialogue-tts-in-6-tests\/","title":{"rendered":"MOSS-TTSD: Dialogue TTS in 6 Tests"},"content":{"rendered":"<p><strong>MOSS-TTSD dialogue TTS<\/strong> was tested here as a script-to-conversation tool, not as a generic read-aloud engine. The six clips ask a narrow question: can one tagged dialogue string keep two speakers distinct while changing pace, energy, language, and emotional weight? Each output below is the original Wiro result, kept intact. The runs use the <a href=\"https:\/\/wiro.ai\/models\/openmoss\/moss-ttsd\">MOSS-TTSD model on Wiro<\/a> with only the <code>dialogue<\/code> field populated, <code>textNormalize<\/code> enabled, and no reference audio or speaker transcript fields supplied.<\/p>\n<nav><strong>On this page<\/strong><\/p>\n<ul>\n<li><a href=\"#setup\">Test setup<\/a><\/li>\n<li><a href=\"#results\">What the six outputs show<\/a><\/li>\n<li><a href=\"#speed\">Run time and cost record<\/a><\/li>\n<li><a href=\"#choose\">When to choose MOSS-TTSD<\/a><\/li>\n<\/ul>\n<\/nav>\n<h2 id=\"setup\">MOSS-TTSD dialogue TTS test setup<\/h2>\n<p>MOSS-TTSD accepts a dialogue string with speaker tags such as <code>[S1]<\/code> and <code>[S2]<\/code>. Wiro also exposes optional reference-audio and reference-text pairs for up to five speakers, but this set deliberately leaves those fields empty. That makes the clips a test of the model&#8217;s default speaker separation and prosody rather than a voice-cloning test. Text normalization stayed on, so punctuation and special characters could be cleaned before synthesis.<\/p>\n<p>The setup matters. A short clip cannot prove the model&#8217;s published long-context claims, but it can reveal whether a scripted exchange has clean boundaries, usable pacing, and enough variation to edit into a podcast, commentary segment, or prototype. The model documentation positions MOSS-TTSD for one to five speakers and long-form dialogue, while the <a href=\"https:\/\/github.com\/OpenMOSS\/MOSS-TTSD\" target=\"_blank\" rel=\"noopener\">OpenMOSS repository<\/a> describes reference-driven continuation for controlled speaker identity. The <a href=\"https:\/\/huggingface.co\/OpenMOSS-Team\/MOSS-TTSD-v1.0\" target=\"_blank\" rel=\"noopener\">MOSS-TTSD v1.0 model card<\/a> lists an 8B model and recommends short 3-10 second references when cloning is needed.<\/p>\n<figure>\n  <img loading=\"lazy\" decoding=\"async\" width=\"2528\" height=\"1696\" class=\"wp-image-1630\" src=\"https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/03\/moss-ttsd-hero.jpg\" alt=\"MOSS-TTSD dialogue TTS podcast studio with microphones and waveform\" srcset=\"https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/03\/moss-ttsd-hero.jpg 2528w, https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/03\/moss-ttsd-hero-510x342.jpg 510w, https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/03\/moss-ttsd-hero-900x604.jpg 900w, https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/03\/moss-ttsd-hero-768x515.jpg 768w, https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/03\/moss-ttsd-hero-1536x1030.jpg 1536w, https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/03\/moss-ttsd-hero-2048x1374.jpg 2048w\" sizes=\"auto, (max-width: 2528px) 100vw, 2528px\" \/><figcaption>Prompt: Photorealistic podcast studio desk with two microphones and headphones. Soft warm lighting. A floating translucent audio waveform and subtitle lines in the background. Shallow depth of field.<\/figcaption><\/figure>\n<h2 id=\"results\">Results: six MOSS-TTSD dialogue TTS checks<\/h2>\n<h3>Test 1: office back and forth<\/h3>\n<p><strong>Dialogue:<\/strong> [S1] Morning. The numbers from yesterday look off. [S2] Yep. The export rounded decimals. [S1] Fix it and resend in ten minutes. [S2] On it.<\/p>\n<figure>\n  <audio controls src=\"https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/03\/moss-ttsd-test-1.mp3\"><\/audio><figcaption>Prompt: the dialogue above. Parameters: dialogue only; no speaker references; text normalization enabled.<\/figcaption><\/figure>\n<p>This is the cleanest baseline because every turn is short and functional. The clip shows whether tags produce an audible change in speaker rather than a single narrator with pauses. The fast handoff works well for an operational exchange: the reply lands after the request, and the compact lines do not invite dramatic over-reading. Pick this pattern for support simulations, explainer dialogue, or rough podcast scripting. Keep turns short when clarity matters more than performance.<\/p>\n<h3>Test 2: podcast intro pacing<\/h3>\n<p><strong>Dialogue:<\/strong> [S1] Welcome back to the show. Today: why latency matters. [S2] And why everyone notices bad timing. [S1] First question. What makes a voice feel real. [S2] Pauses, breaths, and turn taking.<\/p>\n<figure>\n  <audio controls src=\"https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/03\/moss-ttsd-test-2.mp3\"><\/audio><figcaption>Prompt: the dialogue above. Parameters: dialogue only; no speaker references; text normalization enabled.<\/figcaption><\/figure>\n<p>This output checks cadence rather than speed. The punctuation gives the model places to reset, and the turns feel more like an introduction than a list read by two alternating voices. It is a better fit for a host-and-guest outline than Test 1. It also shows why script punctuation deserves an editing pass before generation: short clauses and explicit sentence stops create the timing that a listener hears.<\/p>\n<h3>Test 3: sports commentary energy<\/h3>\n<p><strong>Dialogue:<\/strong> [S1] Goal. Goal. Listen to the crowd. [S2] The pass was perfect. [S1] The striker did not hesitate. [S2] Replay it. Slow. The timing is everything.<\/p>\n<figure>\n  <audio controls src=\"https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/03\/moss-ttsd-test-3.mp3\"><\/audio><figcaption>Prompt: the dialogue above. Parameters: dialogue only; no speaker references; text normalization enabled.<\/figcaption><\/figure>\n<p>The contrast between the repeated opening and the measured replay note makes this a useful energy test. The result has more urgency than the office clip without losing the alternating structure. That does not make it a substitute for a live call, crowd bed, or a carefully directed human commentator. It does make it useful for highlight mockups and pre-produced analysis, where the editor can choose a restrained script and add ambience later.<\/p>\n<h3>Test 4: code-switch lines<\/h3>\n<p><strong>Dialogue:<\/strong> [S1] Quick check. Are we live. [S2] Yes. Ses iyi mi. [S1] Great. Start with the headline. [S2] Tamam. Today the update ships at noon.<\/p>\n<figure>\n  <audio controls src=\"https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/03\/moss-ttsd-test-4.mp3\"><\/audio><figcaption>Prompt: the dialogue above. Parameters: dialogue only; no speaker references; text normalization enabled.<\/figcaption><\/figure>\n<p>This is the caution clip. The model documentation lists Turkish among its supported languages, and the result demonstrates a mixed English-Turkish script rather than a claim of flawless bilingual pronunciation. The handoffs remain intelligible, but code-switched proper nouns, brand names, and short colloquial phrases should always get a native-speaker listening pass. Choose MOSS-TTSD for multilingual drafts and controlled review workflows, not for unattended publication where a single pronunciation mistake is costly.<\/p>\n<h3>Test 5: emotional tone shift<\/h3>\n<p><strong>Dialogue:<\/strong> [S1] I am sorry. I should have called. [S2] You left the room and never came back. [S1] I froze. I did not know what to say. [S2] Say it now. Slowly.<\/p>\n<figure>\n  <audio controls src=\"https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/03\/moss-ttsd-test-5.mp3\"><\/audio><figcaption>Prompt: the dialogue above. Parameters: dialogue only; no speaker references; text normalization enabled.<\/figcaption><\/figure>\n<p>Here the useful signal is restraint. A quieter exchange can easily flatten into one neutral tone or become theatrical. This output keeps the conversation legible without treating every sentence as a climax. Use this shape for audiobook dialogue, narrative previews, and dramatic reads where the words already carry the scene. For a named character or a defined cast voice, add the relevant short reference clip and matching speaker text, then check consent and usage rights before cloning any voice.<\/p>\n<h3>Test 6: production notes debate<\/h3>\n<p><strong>Dialogue:<\/strong> [S1] Step one. Read the script. [S2] Step two. Record clean takes. [S1] Step three. Cut the breaths. [S2] No. Keep some breaths. [S1] Fine. But remove the clicks. [S2] Deal.<\/p>\n<figure>\n  <audio controls src=\"https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/03\/moss-ttsd-test-6.mp3\"><\/audio><figcaption>Prompt: the dialogue above. Parameters: dialogue only; no speaker references; text normalization enabled.<\/figcaption><\/figure>\n<p>The longest exchange in the set tests repeated turns and a small disagreement. The output is useful because the speakers do not need a new label for every sentence: the tags keep the interaction organized. This is the strongest match for tutorial banter and podcast planning segments. It also points to a practical script rule: group each speaker&#8217;s thought into one concise turn instead of bouncing single words back and forth.<\/p>\n<h2 id=\"speed\">Run time and cost record<\/h2>\n<p>The original Wiro task records show elapsed task time, not audio duration. They do not contain a billable cost field, and the current Wiro model documentation does not publish a price for this endpoint. No cost has been inferred or invented. The elapsed times below include task processing, so they are useful as a workflow snapshot rather than a latency guarantee.<\/p>\n<table>\n<tr>\n<th>Test<\/th>\n<th>Elapsed task time<\/th>\n<th>What it checked<\/th>\n<\/tr>\n<tr>\n<td>1<\/td>\n<td>72 seconds<\/td>\n<td>Short operational turn-taking<\/td>\n<\/tr>\n<tr>\n<td>2<\/td>\n<td>87 seconds<\/td>\n<td>Intro pacing and pauses<\/td>\n<\/tr>\n<tr>\n<td>3<\/td>\n<td>82 seconds<\/td>\n<td>High-energy commentary<\/td>\n<\/tr>\n<tr>\n<td>4<\/td>\n<td>76 seconds<\/td>\n<td>English-Turkish code switching<\/td>\n<\/tr>\n<tr>\n<td>5<\/td>\n<td>51 seconds<\/td>\n<td>Restrained emotional delivery<\/td>\n<\/tr>\n<tr>\n<td>6<\/td>\n<td>67 seconds<\/td>\n<td>Repeated two-speaker turns<\/td>\n<\/tr>\n<\/table>\n<h2 id=\"choose\">When to choose MOSS-TTSD<\/h2>\n<p>Choose MOSS-TTSD when the output is a conversation: two or more speakers, an interview, narrated character exchanges, sports analysis, or a podcast segment. It is most convincing when the script makes turn boundaries obvious and each line has a job. Use reference audio only when a consistent, permitted voice identity is important. Leave it out for a quick structural prototype like this set.<\/p>\n<p>Choose a single-speaker TTS model instead when one narrator carries the entire piece. Choose a realtime voice stack when the primary need is live response rather than a finished audio file. For adjacent workflow ideas, see <a href=\"https:\/\/wiro.ai\/blog\/chatterbox-multilingual-voice-cloning-in-23-languages\/\">Chatterbox Multilingual: Voice Cloning in 23 Languages<\/a>, <a href=\"https:\/\/wiro.ai\/blog\/chatterbox-multilingual-5-language-tts-samples\/\">Chatterbox Multilingual: 5 Language TTS Samples<\/a>, and <a href=\"https:\/\/wiro.ai\/blog\/realtime-voice-conversation-wiro-2026\/\">Realtime Voice Conversation: 3 Smart Wiro Setups in 2026<\/a>.<\/p>\n<h2>Try the model<\/h2>\n<p><a href=\"https:\/\/wiro.ai\/models\/openmoss\/moss-ttsd\">Run MOSS-TTSD on Wiro<\/a> with a short tagged scene first. Listen for turn boundaries before moving on to reference voices or a longer script.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>MOSS-TTSD dialogue TTS was tested here as a script-to-conversation tool, not as a generic read-aloud engine. The six clips ask a narrow&hellip;<\/p>\n","protected":false},"author":4,"featured_media":1639,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[52],"tags":[62],"class_list":["post-1638","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-model-reviews","tag-text-to-speech"],"_links":{"self":[{"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/posts\/1638","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/users\/4"}],"replies":[{"embeddable":true,"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/comments?post=1638"}],"version-history":[{"count":2,"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/posts\/1638\/revisions"}],"predecessor-version":[{"id":4218,"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/posts\/1638\/revisions\/4218"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/media\/1639"}],"wp:attachment":[{"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/media?parent=1638"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/categories?post=1638"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/tags?post=1638"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}