{"id":1754,"date":"2026-04-01T09:48:11","date_gmt":"2026-04-01T09:48:11","guid":{"rendered":"https:\/\/wiro.ai\/blog\/?p=1754"},"modified":"2026-09-27T18:46:47","modified_gmt":"2026-09-27T18:46:47","slug":"8-multi-speaker-dialogue-prompts-for-fishaudio-s2-pro","status":"publish","type":"post","link":"https:\/\/wiro.ai\/blog\/8-multi-speaker-dialogue-prompts-for-fishaudio-s2-pro\/","title":{"rendered":"8 Multi-Speaker Dialogue Prompts for FishAudio S2 Pro"},"content":{"rendered":"<p><strong>FishAudio S2 Pro multi-speaker dialogue prompts<\/strong> work because the speaker tokens and delivery tags live in the same text stream. This set checks the practical parts of that claim: turn changes, short pauses, mood changes, a third speaker, mixed-language copy, and script-like pacing. The eight embedded audio files are the original Wiro outputs, kept in place so each prompt can be judged by listening rather than by a written description alone.<\/p>\n<div class=\"wp-block-group\">\n<p><strong>Contents<\/strong><\/p>\n<ul>\n<li><a href=\"#test\">What the test checks<\/a><\/li>\n<li><a href=\"#settings\">Prompt format and parameters<\/a><\/li>\n<li><a href=\"#outputs\">What the eight outputs show<\/a><\/li>\n<li><a href=\"#selection\">When to pick FishAudio S2 Pro<\/a><\/li>\n<\/ul>\n<\/div>\n<h2 id=\"test\">What this FishAudio S2 Pro test checks<\/h2>\n<p>The test is deliberately narrow. It does not try to score a single narrator or clone a supplied voice. Each prompt asks one generation to carry a conversation. The important question is whether the result makes turn-taking intelligible while still giving each line a useful delivery cue.<\/p>\n<p>On Wiro, the model accepts a text prompt, optional reference audio plus its matching transcription for cloning, and generation controls. The model page documents <code>maxNewTokens<\/code>, <code>temperature<\/code>, <code>topP<\/code>, <code>topK<\/code>, <code>chunkLength<\/code>, and <code>seed<\/code>. It recommends 10-30 seconds of reference audio when cloning. None of the eight archived outputs includes a reference clip, reference transcript, or a saved parameter payload, so this post does not claim unrecorded voices, settings, elapsed times, or charges.<\/p>\n<p>The documented defaults are <code>maxNewTokens=0<\/code>, <code>temperature=1.0<\/code>, <code>topP=0.9<\/code>, <code>topK=30<\/code>, <code>chunkLength=300<\/code>, and <code>seed=0<\/code>. Use the same seed with the same inputs when a repeatable revision matters. Lower temperature, such as 0.5, should favor consistency; higher values increase variation. Smaller chunks can help longer scripts, but add processing overhead. Wiro&#8217;s model documentation does not publish a fixed runtime or price for these individual outputs, so no cost estimate is presented here.<\/p>\n<p>FishAudio documents speaker switches with <code>&lt;|speaker:N|&gt;<\/code> and supports inline bracketed delivery instructions. Its public model card describes free-form controls such as <code>[whisper]<\/code>, <code>[excited]<\/code>, <code>[short pause]<\/code>, and longer natural-language cues. That makes a script easier to edit than a collection of separate one-line renders. For the full input surface, see <a href=\"https:\/\/wiro.ai\/models\/fishaudio\/s2-pro\">FishAudio S2 Pro on Wiro<\/a>.<\/p>\n<h2 id=\"settings\">Prompt format and practical setup<\/h2>\n<p>Put a speaker token before every turn, including a return to a previous speaker. Keep the tag close to the words it should affect. Short lines make errors easier to isolate: if a delivery cue misses, change that line instead of rerendering a long scene. The eight prompts use no invented speaker biographies or claims of cloned identities.<\/p>\n<p>For a real production pass, start with the documented defaults and set a nonzero seed before comparing variants. Add a reference clip only when voice cloning is needed, then supply its exact transcription. Do not treat a bracketed cue as a guarantee. It is an instruction, and the audio is the result to inspect.<\/p>\n<h2 id=\"outputs\">Eight multi-speaker dialogue prompts and what the outputs show<\/h2>\n<h3>1. Customer support refund: calm reassurance after a hesitant response<\/h3>\n<p>Prompt: <code>&lt;|speaker:0|&gt;[calm]Thanks for calling. Please say the order number. [pause]&lt;|speaker:1|&gt;[nervous]Uh. It is seven one two nine. [short pause]&lt;|speaker:0|&gt;[reassuring]Got it. A refund request is now submitted.<\/code><\/p>\n<figure><audio controls src=\"https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/03\/fishaudio-dialogue-1.mp3\"><\/audio><figcaption>The output tests a two-way turn change, a filled hesitation, and a pause before the response.<\/figcaption><\/figure>\n<h3>2. Product ad: an excited pitch interrupted by a deadpan line<\/h3>\n<p>Prompt: <code>&lt;|speaker:0|&gt;[excited]Quick update. The NovaCell Pro just dropped. Ultra thin. No buttons. It unlocks when you look at it. &lt;|speaker:1|&gt;[deadpan]So it is face unlock. &lt;|speaker:0|&gt;[laugh]Yes. Want to see the colors.<\/code><\/p>\n<figure><audio controls src=\"https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/03\/fishaudio-dialogue-2.mp3\"><\/audio><figcaption>This output checks whether a tonal contrast survives a quick switch back to the first speaker.<\/figcaption><\/figure>\n<h3>3. Whisper scene: quiet cues and a low-voice close<\/h3>\n<p>Prompt: <code>&lt;|speaker:0|&gt;[whisper]Do not run. &lt;|speaker:1|&gt;[hushed]The camera is on us. &lt;|speaker:0|&gt;[pause][low voice]Keep breathing. Act normal.<\/code><\/p>\n<figure><audio controls src=\"https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/03\/fishaudio-dialogue-3.mp3\"><\/audio><figcaption>The short scene makes it easy to hear whether quiet delivery and the inserted pause change the reading.<\/figcaption><\/figure>\n<h3>4. Bilingual demo: Turkish and English in one exchange<\/h3>\n<p>Prompt: <code>&lt;|speaker:0|&gt;[neutral]Merhaba. Today is a quick demo. &lt;|speaker:1|&gt;[friendly]Hello. &lt;|speaker:0|&gt;Then say: WIRO API. &lt;|speaker:1|&gt;[cheerful]WIRO API. &lt;|speaker:0|&gt;Now a warm goodbye in Turkish: gorusuruz.<\/code><\/p>\n<figure><audio controls src=\"https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/03\/fishaudio-dialogue-4.mp3\"><\/audio><figcaption>This clip probes language switching and branded-letter pronunciation in a compact script.<\/figcaption><\/figure>\n<h3>5. Three-speaker standup: role changes without narration<\/h3>\n<p>Prompt: <code>&lt;|speaker:0|&gt;[serious]Standup starts now. What is blocked. &lt;|speaker:1|&gt;[tired]The build is failing. &lt;|speaker:2|&gt;[focused]A dependency update broke tests. Fix is ready. &lt;|speaker:0|&gt;[short pause]Ship it after CI is green.<\/code><\/p>\n<figure><audio controls src=\"https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/03\/fishaudio-dialogue-5.mp3\"><\/audio><figcaption>Three speaker IDs test whether a brief meeting can remain legible without speaker labels read aloud.<\/figcaption><\/figure>\n<h3>6. Audiobook tension: narration-like lines with an urgent warning<\/h3>\n<p>Prompt: <code>&lt;|speaker:0|&gt;[narration]The elevator stops. The doors open. &lt;|speaker:1|&gt;[confused]Wait. This is not our floor. &lt;|speaker:0|&gt;[urgent]Do not step out. &lt;|speaker:1|&gt;[shaky]Did you hear that.<\/code><\/p>\n<figure><audio controls src=\"https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/03\/fishaudio-dialogue-6.mp3\"><\/audio><figcaption>The output checks fast alternation and a shift from narration to urgency.<\/figcaption><\/figure>\n<h3>7. Technical explainer: robotic framing and a patient correction<\/h3>\n<p>Prompt: <code>&lt;|speaker:0|&gt;[robotic]An API gateway checks auth. It applies rate limits. It routes traffic. &lt;|speaker:1|&gt;[patient]That is the simple version. &lt;|speaker:0|&gt;[curious]What about retries. &lt;|speaker:1|&gt;[calm]Retries belong in the client and the queue.<\/code><\/p>\n<figure><audio controls src=\"https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/03\/fishaudio-dialogue-7.mp3\"><\/audio><figcaption>This output is useful for checking whether delivery tags serve a technical script instead of overwhelming it.<\/figcaption><\/figure>\n<h3>8. Coaching exchange: a deliberate pause before a quiet answer<\/h3>\n<p>Prompt: <code>&lt;|speaker:0|&gt;[gentle]Take a slow breath in. &lt;|speaker:1|&gt;[anxious]I cannot stop thinking about it. &lt;|speaker:0|&gt;[steady]Name one small thing that is under control today. &lt;|speaker:1|&gt;[pause][quiet]Drink water. &lt;|speaker:0|&gt;[warm]Good. Start there.<\/code><\/p>\n<figure><audio controls src=\"https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/03\/fishaudio-dialogue-8.mp3\"><\/audio><figcaption>The last clip tests paced dialogue and local control on one short answer.<\/figcaption><\/figure>\n<h2 id=\"selection\">When to pick FishAudio S2 Pro<\/h2>\n<p>Pick FishAudio S2 Pro when the script needs speaker turns, local emotional direction, or a cloned reference voice in the same job. The model is a better fit for dialogue previews, narrative scenes, support simulations, and structured explainers than for a single neutral announcement with no need for control. Keep dialogue turns concise, listen for boundary drift, and revise the exact line that needs work.<\/p>\n<p>For more voice-model context, see <a href=\"https:\/\/wiro.ai\/blog\/fishaudio-s2-pro-vs-qwen3-tts-6-audio-tests\/\">FishAudio S2 Pro vs Qwen3-TTS<\/a>, <a href=\"https:\/\/wiro.ai\/blog\/chatterbox-multilingual-5-language-tts-samples\/\">Chatterbox Multilingual: 5 Language TTS Samples<\/a>, and <a href=\"https:\/\/wiro.ai\/blog\/chatterbox-turbo-fast-tts-with-paralinguistic-tags-in-6-tests\/\">Chatterbox Turbo: Fast TTS with Paralinguistic Tags<\/a>.<\/p>\n<p>Read the maker&#8217;s <a href=\"https:\/\/huggingface.co\/fishaudio\/s2-pro\" target=\"_blank\" rel=\"noopener\">Fish Audio S2 Pro model card<\/a> and the <a href=\"https:\/\/github.com\/fishaudio\/fish-speech\" target=\"_blank\" rel=\"noopener\">Fish Speech GitHub repository<\/a> for the published model details and implementation material. Then run the same scripts on <a href=\"https:\/\/wiro.ai\/models\/fishaudio\/s2-pro\">FishAudio S2 Pro on Wiro<\/a> with a fixed seed before choosing a production voice.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>FishAudio S2 Pro multi-speaker dialogue prompts work because the speaker tokens and delivery tags live in the same text stream. This set&hellip;<\/p>\n","protected":false},"author":4,"featured_media":1759,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[55],"tags":[143,144,62],"class_list":["post-1754","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-prompt-guides","tag-fishaudio","tag-multi-speaker","tag-text-to-speech"],"_links":{"self":[{"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/posts\/1754","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/users\/4"}],"replies":[{"embeddable":true,"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/comments?post=1754"}],"version-history":[{"count":2,"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/posts\/1754\/revisions"}],"predecessor-version":[{"id":4215,"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/posts\/1754\/revisions\/4215"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/media\/1759"}],"wp:attachment":[{"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/media?parent=1754"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/categories?post=1754"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/tags?post=1754"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}