Skip to content
Comparisons

FishAudio S2 Pro vs Qwen3-TTS: 6 Audio Tests

FishAudio S2 Pro vs Qwen3-TTS is a practical six-clip check of the things that usually break a text-to-speech workflow: spoken digits, short sales copy, quiet atmosphere, a language switch, dialogue pacing, and technical terms. The same scripts were sent to both models and the original MP3 outputs remain below. This is not a lab benchmark. It is a listening set designed to expose where a voice sounds natural, where it sounds overly even, and where a production team needs more control.

What this FishAudio S2 Pro vs Qwen3-TTS test checks

The pair takes different control paths. FishAudio S2 Pro on Wiro accepts text plus optional reference audio and matching reference text for cloning. Its exposed controls include maximum new tokens, temperature, top P, top K, chunk length, and seed. The prompt field also accepts speaker markers such as <|speaker:0|> and inline emotional or prosody cues.

Qwen3-TTS 12Hz 1.7B on Wiro uses a simpler request shape for this comparison: prompt, instruction, language, and a named speaker. The selectable voices include Aiden, Dylan, Eric, Ono Anna, Ryan, Serena, Sohee, Uncle Fu, Vivian, Russian, and Spanish. Its language options include Auto, English, Chinese, French, German, Italian, Japanese, Korean, Portuguese, Russian, and Spanish.

That difference matters. FishAudio’s controls suit a script that needs local changes inside one passage. Qwen’s instruction and speaker fields make a cleaner starting point when a fixed voice and a single broad delivery direction are enough. Both models are documented for text-to-speech; neither clip should be treated as proof of every voice, language, or clone scenario.

Test setup and what is documented

Item FishAudio S2 Pro Qwen3-TTS 12Hz 1.7B
Output preserved here MP3 MP3
Script Same text as the paired Qwen clip Same text as the paired FishAudio clip
Relevant controls Inline cues, temperature, top P, top K, chunk length, seed, optional reference audio Instruction, language, named speaker
Run time per output Not recorded for these original clips Not recorded for these original clips
Cost per output Not shown in the model documentation or the saved post data Not shown in the model documentation or the saved post data

No runtime or cost figure is invented here. The Wiro model pages document inputs, not a fixed per-clip price or an elapsed time for these six historical runs. Text length, queue state, selected settings, and any post-processing can change the observed time. FishAudio’s public model card reports a 0.195 real-time factor and about 100 ms time to first audio on a single H200, but that is a vendor hardware measurement, not a measured Wiro result and should not be used to estimate these clips.

Six audio tests, with the original outputs

1. Account digits: intelligibility under pressure

Script: Hello. Sorry about the issue. To fix this fast, please confirm the last four digits of the account number: 3 4 2 9.

This is the least glamorous test and often the most useful. It asks the listener to check whether the apology, instruction, and four isolated digits stay distinct. Listen for merged digit groups, an unclear final consonant, or a rushed transition into the number. It also shows whether a support-style line can keep a calm pace without sounding flat.

FishAudio S2 Pro – account-digits output.
Qwen3-TTS – account-digits output.

2. Product ad: short-form energy

Script: Quick update. NovaCell Pro just dropped. Ultra thin. No buttons. It unlocks when you look at it. Want to see the colors.

The sentence fragments are deliberate. A good output gives each feature its own beat, then turns the last question upward without making it sound like a different speaker. FishAudio’s inline-control design is relevant when a line needs an explicit energy shift on a word or phrase. Qwen’s instruction field is the faster way to try one consistent ad-read direction across the whole clip.

FishAudio S2 Pro – product-ad output.
Qwen3-TTS – product-ad output.

3. Whisper and atmosphere: localized expression

Script: Tonight the city sounded like rain on glass. The train doors closed. The lights flickered. A message appeared: DO NOT RUN. Nobody moved.

This clip tests pacing, contrast, and emphasis. The capitalized warning should land as a change in tension, not merely as louder text. The useful question is whether the buildup remains understandable while the delivery becomes more dramatic. FishAudio explicitly supports free-form inline descriptions for local prosody control, so it is the natural candidate when a producer wants to mark a whisper, pause, or emphasis inside a longer script. Qwen can instead take a whole-clip emotional instruction.

FishAudio S2 Pro – atmospheric narration output.
Qwen3-TTS – atmospheric narration output.

4. Bilingual demo: handoff between languages

Script: Merhaba. Today is a quick demo. First, say hello. Then say: WIRO API. Then add a warm goodbye in Turkish: gorusuruz.

This is a compact switch test, not a language-coverage claim. It places Turkish beside English and an acronym-like product name. Listen for the boundary between the languages, the spelling-out or reading of WIRO API, and whether the final Turkish phrase feels attached to the rest of the sentence. FishAudio’s public documentation lists Turkish among its supported languages. The Wiro Qwen control shown here exposes an Auto setting but does not list Turkish as a selectable language, so the Qwen clip should be judged as an Auto-mode result rather than proof of a dedicated Turkish mode.

FishAudio S2 Pro – bilingual demo output.
Qwen3-TTS – bilingual demo output.

5. Short dialogue: turn-taking without speaker changes

Script: Are we recording. Yes. Keep it short and clear. Got it.

The four lines create a conversation without supplying a separate voice for each turn. That makes it a useful check for question intonation, brief replies, and pauses. Neither audio pair is a full multi-speaker benchmark. Still, FishAudio’s documented speaker markers make it the more direct model to investigate when a script genuinely needs multiple labelled speakers. This particular clip asks a narrower question: does a single generated voice make the exchange readable?

FishAudio S2 Pro – short-dialogue output.
Qwen3-TTS – short-dialogue output.

6. Technical explainer: plain language and jargon

Script: An API gateway sits in front of services. It checks auth. It applies rate limits. It routes traffic. That is it. Keep the rules boring.

This is a narration test. Short statements expose unnatural pauses more clearly than a long paragraph. The key details are API, auth, rate limits, and the dry closing line. For product tours, documentation videos, and onboarding audio, this type of script often matters more than a cinematic demo. The output that makes each technical unit easy to follow without over-performing it is the better fit.

FishAudio S2 Pro – technical-explainer output.
Qwen3-TTS – technical-explainer output.

When to pick each model

Choose When the job needs
FishAudio S2 Pro Voice cloning with a 10-30 second reference sample and matching transcript, multi-speaker markers, or fine-grained expression inside one script. Its Wiro controls also expose sampling and reproducibility settings for teams that need to tune a run.
Qwen3-TTS 12Hz 1.7B A named voice, a selected language, and a broad instruction such as an emotional delivery. The compact control surface is a practical fit for straightforward narration and repeated templated requests.

Start with the script, not the model label. Choose FishAudio when the script needs localized acting direction or a reference voice. Choose Qwen when a fixed speaker plus a single instruction is enough. Then test the actual copy, especially numbers, product names, abbreviations, and every language switch. For other speech-model comparisons, see Chatterbox Multilingual: 5 Language TTS Samples, Chatterbox Turbo: Fast TTS with Paralinguistic Tags, and Top 5 Text-to-Speech APIs in 2026.

Sources

Try the models

Run FishAudio S2 Pro on Wiro for inline control, cloning inputs, and speaker markers. Run Qwen3-TTS 12Hz 1.7B on Wiro for named voices, selectable languages, and instruction-led delivery.