Audio & SpeechActive
Eleven v4 Dialogue by ElevenLabs
Generate natural multi-speaker dialogue with ElevenLabs Eleven v4. Provide a turn-by-turn script and receive a single MP3 track with expressive delivery cues.
Text to Speech
Model ID
eleven v4 dialogue
Provider
elevenlabs
Updated
1790690898
wiro playground—elevenlabs/eleven v4 dialogue
Updated 1790690898
Overview
Eleven v4 Dialogue by ElevenLabs turns a scripted conversation into a single spoken audio track. You provide the dialogue as ordered turns, and the model synthesizes each line in the selected speaker voice. It can react to bracketed direction tags inside the text, like whispering or laughing, to shape delivery. This is useful when you need consistent, expressive dialogue without recording sessions.
What you can build
- Character scenes for games, with distinct voices per character
- Two-person podcast cold opens and skits
- Audiobook scenes with alternating speakers
- Training role-plays for support, sales, or compliance
- Social video voiceovers with quick back-and-forth pacing
- Table reads for scripts before hiring actors
Inputs
- A dialogue script as a JSON array, with one object per turn that includes a speaker choice and that turn’s text
- Speaker selection per turn using either a built-in voice name available on Wiro or an ElevenLabs voice ID
- Dialogue text that can include bracketed direction tags inside the line, such as whispering, laughing, or excited
- A cap on total script size per request; keep the combined dialogue text short for reliable results
- A limit on speaker variety per request; keep the number of distinct voices small within one scene
- An audio output choice that sets MP3 sample rate and bitrate for file size versus quality
- An optional control that trades expressiveness for steadier delivery across regenerations
- An optional control that pushes the output closer to the reference voice, which can reduce naturalness
Outputs
- A single MP3 audio file that contains the full conversation in order
- One continuous track with each turn rendered sequentially, suitable for direct playback or editing
- Audio fidelity that matches the selected MP3 sample rate and bitrate option
Recommended settings
- For dramatic dialogue, use a lower consistency setting so the delivery varies more across lines
- For narration-like steadiness, use a higher consistency setting to reduce emotional swings
- If a cloned or custom voice starts drifting, increase the voice similarity control in small steps
- If the audio sounds strained or less human, back off the voice similarity control slightly
Limitations
- Very long scenes can fail or sound less consistent. Split long scripts into smaller chunks.
- A single request supports only a limited number of distinct speakers. Reuse voices when possible.
- The model is nondeterministic. The same script can produce different reads across runs.
- Bracketed direction tags are still being tuned. Some tags may be ignored or over-applied.
- Eleven v4 focuses on dialogue realism. Some classic markup styles, like SSML, aren’t supported.
- Messy input hurts quality. Unpunctuated text, unclear turn-taking, or inconsistent speaker labels can degrade pacing.
- Poorly structured JSON will fail validation. Validate your dialogue structure before running it.
Safety & compliance
- Don’t impersonate real people without clear consent and legal rights.
- Don’t use generated audio for scams, fraud, or deceptive identity claims.
- Don’t use synthetic voices for unauthorized robocalls or spam outreach.
- Only use scripts and character voices you have rights to use.
- If you use voice cloning, clone only voices you own or are authorized to clone.
Example prompts
Great starting points for eleven v4 dialogue.
[{"voice": "sarah", "text": "[excited] Did you hear the news?"}, {"voice": "george", "text": "[curious] No, what happened?"}, {"voice": "sarah", "text": "[whispering] Wiro is launching Eleven v4 today."}]Audio & Speech
[{"voice": "jessica", "text": "[annoyed] You forgot the keys again?"}, {"voice": "eric", "text": "[nervous] Maybe... [laughing] okay, yes."}]Audio & Speech
[{"voice": "liam", "text": "[whispering] Is the baby finally asleep?"}, {"voice": "lily", "text": "[softly] Yes. [sighs] Don't you dare sneeze."}]Audio & Speech
[{"voice": "brian", "text": "[casual] Pizza or burgers tonight?"}, {"voice": "alice", "text": "[cheerful] Both. [laughing] It's Friday!"}]Audio & Speech
API quick start
Run eleven v4 dialogue with a single API call.
POST https://api.wiro.ai/v1/Run/elevenlabs/eleven-v4-dialogue
{
"prompt": "[{\"voice\": \"sarah\", \"text\": \"[excited] Di…",
"outputFormat": "mp3_44100_128",
"stability": 0.5,
"similarityBoost": 0.75
}curl
curl -X POST "https://api.wiro.ai/v1/Run/elevenlabs/eleven-v4-dialogue" \
-H "Content-Type: application/json" \
-H "x-api-key: YOUR_WIRO_API_KEY" \
--data-binary @- <<'JSON'
{
"prompt": "[{\"voice\": \"sarah\", \"text\": \"[excited] Di…",
"outputFormat": "mp3_44100_128",
"stability": 0.5,
"similarityBoost": 0.75
}
JSON