Eleven V4 by ElevenLabs (Text-to-Speech)
Eleven V4 by ElevenLabs turns scripts into emotionally rich speech with strong speaker identity and multilingual coverage. Add audio tags and export MP3 audio.
Overview
Eleven V4 is an audio model from ElevenLabs for expressive text-to-speech. It reads your script, infers intent from context, then performs it like a voice actor. You pick a voice, and the model keeps that voice identity stable across long passages. This is useful when you need natural narration or character dialogue without recording sessions.
What you can build
- Audiobook and long-form narration with emotional range
- Character voiceovers for games, animation, and interactive stories
- Multilingual voiceovers that keep the same speaker identity across languages
- Podcast-style segments, intros, and ad reads from a written script
- Dialogue scenes that include reactions like laughter, sighs, or whispering
Inputs
- A script you want spoken, as plain text, up to 10,000 characters per request
- Optional bracketed direction tags inside the script, like [whispering], [laughing], or [shouting]
- A voice choice from the included preset voices
- An MP3 export quality choice, based on sample rate and bitrate
- An optional stability control from 0 to 1 to trade expressiveness for consistency
- An optional similarity control from 0 to 1 to keep closer to the chosen voice
Outputs
The model returns a single spoken-audio file in MP3 format. The file contains the rendered performance of your script in the selected voice. If you include bracketed direction tags, the audio may include delivery changes and non-verbal reactions where they appear.
Limitations
- One request supports up to 10,000 characters, which is roughly 10 minutes of audio
- Output can vary between runs, even with the same script and settings
- Audio tags are best-effort. Some tags may be ignored or applied inconsistently
- Names, numbers, and abbreviations can be misread. Write the spoken form when it matters
- Voice quality depends on the selected voice. Noisy or inconsistent cloned samples can add artifacts
- Poorly structured text can reduce quality, including messy punctuation and copied PDF formatting
Safety & compliance
- Only clone voices when you have the right to use that voice. Professional voice cloning requires verification
- ElevenLabs provides detection tooling for ElevenLabs-generated audio, including watermark-based detection
- Generated audio may still enable impersonation misuse. Use clear disclosure where required by policy or law
- You retain ownership of generated audio, but commercial use depends on your ElevenLabs plan terms
Example prompts
Great starting points for eleven v4.
API quick start
Run eleven v4 with a single API call.
{
"prompt": "[measured] Every great story begins with …",
"voice": "george",
"outputFormat": "mp3_44100_128",
"stability": 0.5
}curl -X POST "https://api.wiro.ai/v1/Run/elevenlabs/eleven-v4" \
-H "Content-Type: application/json" \
-H "x-api-key: YOUR_WIRO_API_KEY" \
--data-binary @- <<'JSON'
{
"prompt": "[measured] Every great story begins with …",
"voice": "george",
"outputFormat": "mp3_44100_128",
"stability": 0.5
}
JSON