FishAudio S2 Pro multi-speaker dialogue prompts work because the speaker tokens and delivery tags live in the same text stream. This set checks the practical parts of that claim: turn changes, short pauses, mood changes, a third speaker, mixed-language copy, and script-like pacing. The eight embedded audio files are the original Wiro outputs, kept in place so each prompt can be judged by listening rather than by a written description alone.
Contents
What this FishAudio S2 Pro test checks
The test is deliberately narrow. It does not try to score a single narrator or clone a supplied voice. Each prompt asks one generation to carry a conversation. The important question is whether the result makes turn-taking intelligible while still giving each line a useful delivery cue.
On Wiro, the model accepts a text prompt, optional reference audio plus its matching transcription for cloning, and generation controls. The model page documents maxNewTokens, temperature, topP, topK, chunkLength, and seed. It recommends 10-30 seconds of reference audio when cloning. None of the eight archived outputs includes a reference clip, reference transcript, or a saved parameter payload, so this post does not claim unrecorded voices, settings, elapsed times, or charges.
The documented defaults are maxNewTokens=0, temperature=1.0, topP=0.9, topK=30, chunkLength=300, and seed=0. Use the same seed with the same inputs when a repeatable revision matters. Lower temperature, such as 0.5, should favor consistency; higher values increase variation. Smaller chunks can help longer scripts, but add processing overhead. Wiro’s model documentation does not publish a fixed runtime or price for these individual outputs, so no cost estimate is presented here.
FishAudio documents speaker switches with <|speaker:N|> and supports inline bracketed delivery instructions. Its public model card describes free-form controls such as [whisper], [excited], [short pause], and longer natural-language cues. That makes a script easier to edit than a collection of separate one-line renders. For the full input surface, see FishAudio S2 Pro on Wiro.
Prompt format and practical setup
Put a speaker token before every turn, including a return to a previous speaker. Keep the tag close to the words it should affect. Short lines make errors easier to isolate: if a delivery cue misses, change that line instead of rerendering a long scene. The eight prompts use no invented speaker biographies or claims of cloned identities.
For a real production pass, start with the documented defaults and set a nonzero seed before comparing variants. Add a reference clip only when voice cloning is needed, then supply its exact transcription. Do not treat a bracketed cue as a guarantee. It is an instruction, and the audio is the result to inspect.
Eight multi-speaker dialogue prompts and what the outputs show
1. Customer support refund: calm reassurance after a hesitant response
Prompt: <|speaker:0|>[calm]Thanks for calling. Please say the order number. [pause]<|speaker:1|>[nervous]Uh. It is seven one two nine. [short pause]<|speaker:0|>[reassuring]Got it. A refund request is now submitted.
2. Product ad: an excited pitch interrupted by a deadpan line
Prompt: <|speaker:0|>[excited]Quick update. The NovaCell Pro just dropped. Ultra thin. No buttons. It unlocks when you look at it. <|speaker:1|>[deadpan]So it is face unlock. <|speaker:0|>[laugh]Yes. Want to see the colors.
3. Whisper scene: quiet cues and a low-voice close
Prompt: <|speaker:0|>[whisper]Do not run. <|speaker:1|>[hushed]The camera is on us. <|speaker:0|>[pause][low voice]Keep breathing. Act normal.
4. Bilingual demo: Turkish and English in one exchange
Prompt: <|speaker:0|>[neutral]Merhaba. Today is a quick demo. <|speaker:1|>[friendly]Hello. <|speaker:0|>Then say: WIRO API. <|speaker:1|>[cheerful]WIRO API. <|speaker:0|>Now a warm goodbye in Turkish: gorusuruz.
5. Three-speaker standup: role changes without narration
Prompt: <|speaker:0|>[serious]Standup starts now. What is blocked. <|speaker:1|>[tired]The build is failing. <|speaker:2|>[focused]A dependency update broke tests. Fix is ready. <|speaker:0|>[short pause]Ship it after CI is green.
6. Audiobook tension: narration-like lines with an urgent warning
Prompt: <|speaker:0|>[narration]The elevator stops. The doors open. <|speaker:1|>[confused]Wait. This is not our floor. <|speaker:0|>[urgent]Do not step out. <|speaker:1|>[shaky]Did you hear that.
7. Technical explainer: robotic framing and a patient correction
Prompt: <|speaker:0|>[robotic]An API gateway checks auth. It applies rate limits. It routes traffic. <|speaker:1|>[patient]That is the simple version. <|speaker:0|>[curious]What about retries. <|speaker:1|>[calm]Retries belong in the client and the queue.
8. Coaching exchange: a deliberate pause before a quiet answer
Prompt: <|speaker:0|>[gentle]Take a slow breath in. <|speaker:1|>[anxious]I cannot stop thinking about it. <|speaker:0|>[steady]Name one small thing that is under control today. <|speaker:1|>[pause][quiet]Drink water. <|speaker:0|>[warm]Good. Start there.
When to pick FishAudio S2 Pro
Pick FishAudio S2 Pro when the script needs speaker turns, local emotional direction, or a cloned reference voice in the same job. The model is a better fit for dialogue previews, narrative scenes, support simulations, and structured explainers than for a single neutral announcement with no need for control. Keep dialogue turns concise, listen for boundary drift, and revise the exact line that needs work.
For more voice-model context, see FishAudio S2 Pro vs Qwen3-TTS, Chatterbox Multilingual: 5 Language TTS Samples, and Chatterbox Turbo: Fast TTS with Paralinguistic Tags.
Read the maker’s Fish Audio S2 Pro model card and the Fish Speech GitHub repository for the published model details and implementation material. Then run the same scripts on FishAudio S2 Pro on Wiro with a fixed seed before choosing a production voice.