Skip to content
Model Reviews

Chatterbox Turbo: Fast TTS with Paralinguistic Tags in 6 Tests

Chatterbox Turbo is the fast, English-first member of Resemble AI’s open-source Chatterbox family. This six-sample set checks a practical question: can a low-latency voice model keep a cloned speaker understandable when the script stops being plain narration? The tests use short reference-audio cloning, tight scripts, punctuation-driven pacing, and tags such as [sigh], [chuckle], [whisper], and [pause].

What the Chatterbox Turbo test set checks

These are not benchmark scores. They are short production-style reads designed to expose the problems that matter in a voice agent, ad, narration pass, or explainer: transitions into a non-speech cue, consonant clarity, pauses, emphasis, proper nouns, and language switches. Every output below remains available as an audio player so the claims can be checked by ear rather than taken on trust.

The same reference clip was supplied as inputAudio for the six generated samples. That keeps the speaker target consistent and makes differences in wording and delivery easier to hear. The reference itself is retained below. The Wiro model page lists MP3 among the supported input formats; the article’s original test notes identify the delivered test files as MP3.

Reference audio used for voice cloning

The short voice reference used as inputAudio for all six tests.

Chatterbox Turbo is positioned for low-latency English voice work. The upstream Chatterbox repository describes Turbo as a 350M-parameter model with a distilled, single-step speech-token-to-mel decoder and native paralinguistic tags. That design goal explains the test choices: short turns and interruptions are more revealing here than a long audiobook chapter.

What each audio output shows

1. Calm support apology: pacing, digits, and a sigh

Prompt: I completely understand the frustration you are experiencing. [sigh] To help fix this fast, please confirm the last four digits of your account number.

This recording puts a soft non-speech cue between two support sentences, then asks for four digits. It is the most useful sample for support flows because it makes three things audible in one short turn: whether the tag gets a clean boundary, whether the voice resumes without a jump in identity, and whether the final number request stays intelligible. The sample is a better fit for a scripted apology than a free-form conversation, since the text is fixed and the voice has no need to decide what to say next.

2. Product ad: a chuckle beside clipped copy

Prompt: Hey! [chuckle] Quick update: the new NovaCell Pro just dropped. Ultra thin. No buttons. It unlocks when you look at it. Want to see the colors?

The second output tests a common ad-read pattern: a quick informal cue followed by short fragments. Listen for whether [chuckle] stays separate from “Quick update,” and whether the clipped product claims retain stops between sentences. It shows why Turbo suits brief social, product, and agent prompts. The tag is part of the text, not a guaranteed direction to an actor, so it should be auditioned for each production voice instead of assumed to sound the same across references.

3. Suspense narration: whisper placement and emphasis

Prompt: Tonight the city sounded like rain on glass. The train doors closed, the lights flickered, and a single message appeared on the screen: DO NOT RUN. [whisper] Nobody moved.

This is the stress test for expressive direction. The script asks for normal narration, all-caps emphasis, then a whisper cue. It makes sibilance, breath noise, and a sudden register change easy to spot. The recording is useful as a warning as well as a demo: dramatic tags can work, but a noisy reference or a more extreme setting can make them less clean than neutral speech.

4. Turkish and English: language-switch risk

Prompt: Merhaba! Today is a quick demo. First, say hello. Then, say: WIRO API. Then, add a warm goodbye in Turkish: gorusuruz.

The fourth player deliberately mixes Turkish and English with an all-caps product term. It does not establish that Turbo is a multilingual model. Instead, it exposes pronunciation drift that can appear when an English-focused model meets mixed-language text, transliterations, and acronyms. For a global assistant, use a model built for multilingual voice cloning rather than treating this one sample as language support. The related Chatterbox Multilingual voice-cloning test covers that use case.

5. Empathetic coaching: a pause without dead air

Prompt: Family can feel complicated when everything changes. [pause] If today feels heavy, pick one small thing you can control. Drink water. Step outside. Text one person you trust.

This output checks a quieter form of control. The pause sits after an emotionally loaded sentence, followed by three short instructions. The useful signal is whether the break sounds intentional and whether the final commands keep one calm voice rather than accelerating into a list. That makes this style suitable for guided scripts and reminders, provided the copy is reviewed for the context in which it will be heard.

6. Technical explainer: articulation under jargon

Prompt: Here is the simple version. An API gateway sits in front of your services. It checks auth, applies rate limits, and routes traffic. That is it. Keep the rules boring.

The last output removes emotion and focuses on technical words: “API gateway,” “auth,” and “rate limits.” This is the cleanest fit for Turbo’s speed-oriented role. It tests whether short engineering copy remains crisp without adding theatrical delivery. It also pairs naturally with a live conversation stack; see three Wiro setups for realtime voice conversation and realtime speech-to-text models on Wiro for the adjacent input and orchestration side.

Parameters, run time, and cost

The preserved test record confirms the shared voice reference, the six prompts, and MP3 outputs. It does not retain the original Wiro task receipts, so it would be inaccurate to assign an exact historical duration or price to each embedded player. The Wiro documentation lists the model controls and defaults: exaggeration 0.5, cfg_weight 0.5, temperature 0.8, topP 1.0, and topK 1000. These are documented defaults, not a claim that every original submission used them.

The same documentation says exaggeration ranges from 0.25 to 2.0, with 0.5 marked neutral; high values may be unstable. cfg_weight ranges from 0.2 to 1.0, with 0.5 suggested for balance and 0 for language transfer. Temperature ranges from 0.05 to 5.0, while top-p and top-k narrow token selection. For a repeatable production test, start with those defaults, keep the reference unchanged, alter one parameter at a time, and save the task receipt with the output.

No per-output price is published in the model documentation, and the original receipts are not available in this post, so no cost is invented here. One documentation example shows a completed task with six seconds of elapsed processing, but it is an example response rather than a measurement for these six clips. Treat it as an illustration of the task record format, not a service-level promise. For the underlying open-source model and implementation notes, see the checked Chatterbox page on Hugging Face and the project repository linked above.

When to pick Chatterbox Turbo

Pick Chatterbox Turbo when the work is English-first, the turn is short, and response speed matters: voice agents, product prompts, compact narrations, and scripted explainers. The six players show the kinds of text worth auditioning before launch: tags, acronyms, digits, whisper-like direction, and pauses. Keep expressiveness near neutral first. Increase it only after listening to the exact reference voice and script.

Choose a multilingual Chatterbox variant when language coverage or cross-language cloning matters. Choose a smaller on-device option when CPU and memory limits dominate. Above all, test the final reference clip, punctuation, tags, and target device together. A voice model can sound convincing on a demo sentence and still miss a name, acronym, or code in the production script.

Run Chatterbox Turbo on Wiro with the same kinds of short prompts, then keep the output that matches the use case rather than the one that merely sounds dramatic.