Eleven V4 Turbo Dialogue by ElevenLabs
Turn a scripted conversation into expressive speech with ElevenLabs Eleven v4 Turbo. Provide a multi-voice dialogue JSON and get a single MP3 scene.
Overview
Eleven V4 Turbo Dialogue generates spoken conversations from a turn-by-turn script. You provide each line of dialogue plus the voice to speak it. The model renders the full exchange as one continuous audio clip. It’s useful when you need natural pacing across speakers without hand-stitching separate files.
What you can build
- NPC conversations for games, with fast iteration on character delivery
- Multi-character scenes for animatics, previsualization, and table reads
- Podcast skits and cold opens with clear speaker switching
- Voice agent demo dialogues for support, booking, or sales flows
- Role-play language learning dialogues with emotion and tone cues
Inputs
- A dialogue script as a JSON array, with one object per turn in the conversation.
- For each turn, you provide the speaker’s voice and the exact text that speaker should say.
- Voices can be one of the preset names (Adam, Alice, Bella, Bill, Brian, Callum, Charlie, Chris, Daniel, Eric, George, Harry, Jessica, Laura, Liam, Lily, Matilda, River, Roger, Sarah, Will) or a custom ElevenLabs voice ID.
- Keep the total dialogue text to 2,000 characters or less for reliable generation.
- Use no more than 10 distinct voices across the full script.
- Optional performance cues inside the text using bracketed audio tags such as [curious], [nervous], [laughing], or [excited].
- An MP3 output choice that sets sample rate and bitrate (22.05 kHz at 32 kbps, or 44.1 kHz at 32, 64, 96, or 128 kbps).
- Optional voice consistency controls on a 0 to 1 scale.
- Lower stability tends to sound more expressive and variable.
- Higher similarity tends to match the target voice more closely.
Outputs
The model returns a single MP3 audio file.
The audio contains every dialogue turn in order, with speaker changes applied between turns. The encoding matches the MP3 format you selected, including sample rate and bitrate.
Recommended settings
- Start with stability at 0.5 for balanced delivery.
- Start with similarity at 0.75 for a strong voice match without harsh artifacts.
- For more acting range, try stability from 0.2 to 0.4.
- For steadier narration, try stability from 0.7 to 0.9.
- If the voice sounds strained or unnatural, reduce similarity by 0.1 to 0.2.
Limitations
- Total dialogue text over 2,000 characters can fail or cut off. Split long scripts into smaller scenes.
- Use up to 10 distinct voices per request.
- Results vary between runs, especially at low stability.
- Audio tags are not guaranteed. Some voices follow tags better than others.
- Low-quality inputs reduce quality. This includes invalid JSON, unclear turn ordering, and messy punctuation.
- Very long single turns can cause pacing issues. Break long monologues into smaller turns.
Safety & compliance
- Only generate speech for voices you have rights to use.
- Don’t clone or imitate a real person’s voice without clear consent or legal authority.
- Don’t use generated audio for fraud, scams, or deceptive impersonation.
- Follow telemarketing and robocalling rules if you use the output in outbound calls.
- Expect platform enforcement. ElevenLabs monitors misuse and provides reporting and detection tools.
Example prompts
Great starting points for eleven v4 turbo dialogue.
API quick start
Run eleven v4 turbo dialogue with a single API call.
{
"prompt": "[{\"voice\": \"matilda\", \"text\": \"[curious] …",
"outputFormat": "mp3_44100_128",
"languageCode": "auto",
"timestamps": "false"
}curl -X POST "https://api.wiro.ai/v1/Run/elevenlabs/eleven-v4-turbo-dialogue" \
-H "Content-Type: application/json" \
-H "x-api-key: YOUR_WIRO_API_KEY" \
--data-binary @- <<'JSON'
{
"prompt": "[{\"voice\": \"matilda\", \"text\": \"[curious] …",
"outputFormat": "mp3_44100_128",
"languageCode": "auto",
"timestamps": "false"
}
JSON