Skip to content
Prompt Guides

8 Multi-Speaker Dialogue Prompts for FishAudio S2 Pro

FishAudio S2 Pro multi-speaker dialogue prompts work because the speaker tokens and delivery tags live in the same text stream. This set checks the practical parts of that claim: turn changes, short pauses, mood changes, a third speaker, mixed-language copy, and script-like pacing. The eight embedded audio files are the original Wiro outputs, kept in place so each prompt can be judged by listening rather than by a written description alone.

What this FishAudio S2 Pro test checks

The test is deliberately narrow. It does not try to score a single narrator or clone a supplied voice. Each prompt asks one generation to carry a conversation. The important question is whether the result makes turn-taking intelligible while still giving each line a useful delivery cue.

On Wiro, the model accepts a text prompt, optional reference audio plus its matching transcription for cloning, and generation controls. The model page documents maxNewTokens, temperature, topP, topK, chunkLength, and seed. It recommends 10-30 seconds of reference audio when cloning. None of the eight archived outputs includes a reference clip, reference transcript, or a saved parameter payload, so this post does not claim unrecorded voices, settings, elapsed times, or charges.

The documented defaults are maxNewTokens=0, temperature=1.0, topP=0.9, topK=30, chunkLength=300, and seed=0. Use the same seed with the same inputs when a repeatable revision matters. Lower temperature, such as 0.5, should favor consistency; higher values increase variation. Smaller chunks can help longer scripts, but add processing overhead. Wiro’s model documentation does not publish a fixed runtime or price for these individual outputs, so no cost estimate is presented here.

FishAudio documents speaker switches with <|speaker:N|> and supports inline bracketed delivery instructions. Its public model card describes free-form controls such as [whisper], [excited], [short pause], and longer natural-language cues. That makes a script easier to edit than a collection of separate one-line renders. For the full input surface, see FishAudio S2 Pro on Wiro.

Prompt format and practical setup

Put a speaker token before every turn, including a return to a previous speaker. Keep the tag close to the words it should affect. Short lines make errors easier to isolate: if a delivery cue misses, change that line instead of rerendering a long scene. The eight prompts use no invented speaker biographies or claims of cloned identities.

For a real production pass, start with the documented defaults and set a nonzero seed before comparing variants. Add a reference clip only when voice cloning is needed, then supply its exact transcription. Do not treat a bracketed cue as a guarantee. It is an instruction, and the audio is the result to inspect.

Eight multi-speaker dialogue prompts and what the outputs show

1. Customer support refund: calm reassurance after a hesitant response

Prompt: <|speaker:0|>[calm]Thanks for calling. Please say the order number. [pause]<|speaker:1|>[nervous]Uh. It is seven one two nine. [short pause]<|speaker:0|>[reassuring]Got it. A refund request is now submitted.

The output tests a two-way turn change, a filled hesitation, and a pause before the response.

2. Product ad: an excited pitch interrupted by a deadpan line

Prompt: <|speaker:0|>[excited]Quick update. The NovaCell Pro just dropped. Ultra thin. No buttons. It unlocks when you look at it. <|speaker:1|>[deadpan]So it is face unlock. <|speaker:0|>[laugh]Yes. Want to see the colors.

This output checks whether a tonal contrast survives a quick switch back to the first speaker.

3. Whisper scene: quiet cues and a low-voice close

Prompt: <|speaker:0|>[whisper]Do not run. <|speaker:1|>[hushed]The camera is on us. <|speaker:0|>[pause][low voice]Keep breathing. Act normal.

The short scene makes it easy to hear whether quiet delivery and the inserted pause change the reading.

4. Bilingual demo: Turkish and English in one exchange

Prompt: <|speaker:0|>[neutral]Merhaba. Today is a quick demo. <|speaker:1|>[friendly]Hello. <|speaker:0|>Then say: WIRO API. <|speaker:1|>[cheerful]WIRO API. <|speaker:0|>Now a warm goodbye in Turkish: gorusuruz.

This clip probes language switching and branded-letter pronunciation in a compact script.

5. Three-speaker standup: role changes without narration

Prompt: <|speaker:0|>[serious]Standup starts now. What is blocked. <|speaker:1|>[tired]The build is failing. <|speaker:2|>[focused]A dependency update broke tests. Fix is ready. <|speaker:0|>[short pause]Ship it after CI is green.

Three speaker IDs test whether a brief meeting can remain legible without speaker labels read aloud.

6. Audiobook tension: narration-like lines with an urgent warning

Prompt: <|speaker:0|>[narration]The elevator stops. The doors open. <|speaker:1|>[confused]Wait. This is not our floor. <|speaker:0|>[urgent]Do not step out. <|speaker:1|>[shaky]Did you hear that.

The output checks fast alternation and a shift from narration to urgency.

7. Technical explainer: robotic framing and a patient correction

Prompt: <|speaker:0|>[robotic]An API gateway checks auth. It applies rate limits. It routes traffic. <|speaker:1|>[patient]That is the simple version. <|speaker:0|>[curious]What about retries. <|speaker:1|>[calm]Retries belong in the client and the queue.

This output is useful for checking whether delivery tags serve a technical script instead of overwhelming it.

8. Coaching exchange: a deliberate pause before a quiet answer

Prompt: <|speaker:0|>[gentle]Take a slow breath in. <|speaker:1|>[anxious]I cannot stop thinking about it. <|speaker:0|>[steady]Name one small thing that is under control today. <|speaker:1|>[pause][quiet]Drink water. <|speaker:0|>[warm]Good. Start there.

The last clip tests paced dialogue and local control on one short answer.

When to pick FishAudio S2 Pro

Pick FishAudio S2 Pro when the script needs speaker turns, local emotional direction, or a cloned reference voice in the same job. The model is a better fit for dialogue previews, narrative scenes, support simulations, and structured explainers than for a single neutral announcement with no need for control. Keep dialogue turns concise, listen for boundary drift, and revise the exact line that needs work.

For more voice-model context, see FishAudio S2 Pro vs Qwen3-TTS, Chatterbox Multilingual: 5 Language TTS Samples, and Chatterbox Turbo: Fast TTS with Paralinguistic Tags.

Read the maker’s Fish Audio S2 Pro model card and the Fish Speech GitHub repository for the published model details and implementation material. Then run the same scripts on FishAudio S2 Pro on Wiro with a fixed seed before choosing a production voice.