Skip to content
Model Reviews

Chatterbox Multilingual: 5 Language TTS Samples

Chatterbox Multilingual was tested here with one practical task: speak a short delivery update in English, Turkish, Spanish, Japanese, and Arabic without changing the core message. That makes the differences easy to hear. Each line includes a time reference, a shipment term, and everyday sentence rhythm rather than a single isolated phrase. The five players below are the original Wiro-hosted outputs from the test.

Chatterbox Multilingual on Wiro supports 23 language choices, optional reference audio, and controls for expression and sampling. The model documentation describes it as multilingual text-to-speech with instant voice cloning and cross-language voice transfer. This post uses text-only generation, so no reference clip was supplied and no claim is made about how a cloned speaker transfers between languages.

What this multilingual TTS test checks

A useful multilingual TTS check has to separate language coverage from voice-cloning quality. This set checks the first part: can the model turn a comparable, short customer-facing update into speech in five writing systems and language families? English, Turkish, and Spanish expose normal Latin-script phrasing. Japanese tests a non-Latin script with different sentence structure. Arabic tests right-to-left script and a different phonetic inventory.

The samples do not prove accent accuracy for every region, long-form consistency, or speaker identity transfer. They are short by design. A short prompt makes pauses, pronunciation choices, and unstable syllables easier to spot. It also avoids pretending that five clips can stand in for a full localization review. Production teams should still have a fluent listener review names, dates, numbers, and regulated wording.

Parameters used for every Chatterbox Multilingual output

Parameter Value Why it matters
inputAudio Not supplied These are text-only samples, not voice-clone tests.
language en, tr, es, ja, ar The selected language changes for each matching line.
exaggeration 0.5 The documentation calls 0.5 neutral; the supported range is 0.25 to 2.0.
cfg_weight 0.5 The documented balanced setting. The docs reserve 0 for language transfer.
temperature 0.8 A moderate sampling setting within the documented 0.05 to 5.0 range.
topP 1.0 No additional nucleus restriction was applied.

Keeping these settings fixed matters more than it may seem. Changing exaggeration or temperature mid-test would mix language differences with style differences. At 0.5 exaggeration, these clips aim for a neutral delivery update rather than dramatic narration. A higher value may suit character reads or promos, but the docs warn that extreme values can become unstable.

Five original language outputs

1. English: baseline for pace and sentence breaks

Text: Your package is on the way. Delivery is expected tomorrow morning.

This clip supplies the baseline. The two-sentence wording checks whether the model gives the delivery statement and the arrival estimate distinct beats. It is the cleanest place to judge pacing before moving to translated copy. It does not test a cloned voice because inputAudio was absent.

2. Turkish: agglutinative phrasing

Text: Paketiniz yolda. Teslimat yarin sabah bekleniyor.

The Turkish line keeps the same delivery intent in a more compact structure. This output is useful for checking how the selected Turkish mode handles a customer-facing statement rather than an English sentence read with a different accent. The test text uses plain ASCII spelling, including yarin; that means this clip cannot validate how the model handles the Turkish dotless i in correctly accented source text. Use correct local spelling in production prompts.

3. Spanish: a longer everyday delivery line

Text: Tu pedido ya va en camino. La entrega esta prevista para manana por la manana.

The Spanish sample has more words and repeats the time-of-day idea. It checks whether a longer service update stays coherent across two sentences. The stored prompt omits Spanish accents in esta, manana, and manana. That is a real limitation of this test set: the audio demonstrates the supplied text, not a linguistic pass on fully accented Spanish copy. A localized prompt should preserve accents and punctuation.

4. Japanese: native-script delivery notice

Text: ご注文の商品は発送されました。配達は明日の午前中の予定です。

This sample shifts from Latin letters to Japanese script while retaining the same shipment and next-morning message. It is the strongest check here for whether the chosen language field and written input agree. It is not a test of romanization, dialect preference, or formal customer-service style beyond this one sentence pair. Japanese teams should test their own honorifics, product names, and date formats.

5. Arabic: right-to-left script

Text: طلبك في الطريق. من المتوقع التسليم صباح الغد.

The Arabic output adds a right-to-left script and a distinct sound system. It checks the model on a short, direct fulfillment message, not on vowel-marked text, names, or regional dialect. Arabic localization needs an extra review because the same written phrase can vary by market and because delivery copy often contains numbers, addresses, and brand terms that deserve their own tests.

Run time and cost: what the saved test can support

The post’s saved output records contain the five audio files but no task receipts, elapsed-time values, or cost fields for these specific runs. The current Wiro documentation lists input controls and language support, but it does not publish a fixed run time or fixed price per generated clip. For that reason, this article does not assign a made-up time or cost to any of the five players. Actual time and cost can vary with prompt length, queue conditions, and any reference-audio processing. Check the task result for a live run when a per-output figure is required.

When to pick Chatterbox Multilingual

Pick this model when one workflow needs text-to-speech across several supported languages and a future pass may need a reference voice. The language selector makes the intended target explicit, and the control set gives room to tune expression after a neutral first pass. Start with the settings used here, then change one control at a time. For a calm customer notice, keep exaggeration near neutral. For a more animated read, raise it carefully and listen for instability.

Choose a faster, narrowly targeted TTS model when low latency matters more than multilingual coverage or voice transfer. For expressive delivery in one language, compare a short set of identical prompts before committing. Related tests on this blog include Chatterbox Turbo with paralinguistic tags, VibeVoice Realtime, and five text-to-speech APIs.

Sources and practical takeaway

For model background, see the official Chatterbox page from Resemble AI, the open-source Chatterbox repository on GitHub, and the Chatterbox model card on Hugging Face. Each source was checked successfully before this update.

The useful result is modest but clear: these five existing clips make it possible to hear Chatterbox Multilingual on comparable delivery text in five languages. They are a starting point for a real localization test, not a substitute for it. Run your own copy with correct spelling, representative names and numbers, then have native listeners approve the final audio.