Audio & SpeechActive
Google Gemini 3.1 Flash TTS
Generate natural speech from text with Gemini 3.1 Flash TTS. Use voice options and expressive tags to control tone and pacing for narration, apps, and accessibility.
Text to Speech
Model ID
gemini-3.1-tts
Provider
google
Updated
1788788171
wiro playground—google/gemini-3.1-tts
Updated 1788788171
Overview
Gemini 3.1 Flash TTS is a text-to-speech model from Google. It uses a large language model to decide both the words and the delivery. You steer performance with plain text direction plus inline square-bracket tags. It outputs spoken audio that’s suitable for voiceovers, narration, and spoken UI.
The model is built for controllable speech. You can change pacing, emotion, and delivery mid-sentence. This helps you ship audio that sounds intentional, not flat.
What you can build
- App voiceovers with a consistent “brand voice”
- Audiobook-style narration with pacing and pause control
- Character dialogue for games (single-voice on this Wiro page)
- Accessibility narration for articles, screens, and UI flows
- Multilingual announcements and help prompts
- Podcast-style scripted segments where exact wording matters
Inputs
- The script you want spoken, provided as plain text. Keep it under roughly 16K tokens of text for the Gemini 3.1 Flash TTS model.
- A style direction written directly into the script, like “Say cheerfully: …” so the model knows how to perform.
- Inline audio tags in square brackets, placed inside the script where you want the change. Examples include tags for laughter, whispering, emotion, and pauses. Google documents 200+ supported tags.
- One prebuilt voice, selected from 30 options: Zephyr, Puck, Charon, Kore, Fenrir, Leda, Orus, Aoede, Callirrhoe, Autonoe, Enceladus, Iapetus, Umbriel, Algieba, Despina, Erinome, Algenib, Rasalgethi, Laomedeia, Achernar, Alnilam, Schedar, Gacrux, Pulcherrima, Achird, Zubenelgenubi, Vindemiatrix, Sadachbia, Sadaltager, Sulafat.
Outputs
- A single speech audio result that reads your script in the selected voice.
- The output is delivered as a playable audio file on Wiro.
- The underlying Gemini TTS examples commonly use mono 24,000 Hz, 16-bit PCM audio wrapped in a WAV container.
- Audio generated by Gemini 3.1 Flash TTS is watermarked with SynthID.
Limitations
- This model is in Preview on Google’s side. Behavior and limits can change.
- It’s text-to-speech. It’s meant to read provided text, not invent a new script.
- Multi-speaker dialogue exists in Google’s Gemini TTS offering, but it’s limited to 2 speakers in the Gemini API. This Wiro page exposes a single voice choice.
- If you use Gemini TTS through Google Cloud Text-to-Speech, the “text” and “prompt” fields can each be at most 4,000 bytes, with 8,000 bytes combined. That path can truncate output around 655 seconds.
- Very long scripts can drift in prosody or volume over time. Split long narration into shorter chunks.
- Messy input text can sound bad. Expect problems with OCR errors, missing punctuation, or inconsistent casing.
Safety & compliance
- Gemini 3.1 Flash TTS audio includes SynthID watermarking to help detect AI-generated speech.
- Don’t use it for deception. Don’t impersonate real people, mislead listeners, or hide that audio is synthetic.
- Follow Google’s generative AI policies. Avoid disallowed content such as hate, harassment, sexual content, or instructions for wrongdoing.
Example prompts
Great starting points for gemini-3.1-tts.
Synthesize the speech below as a live horse-racing commentator.
AUDIO PROFILE: Ray Tulloch, veteran racecourse commentator
SCENE: The final furlong at Doncaster. Forty thousand people on their feet, hooves thundering, the two leaders inseparable.
DIRECTOR'S NOTES
Style: Controlled chaos. Rising urgency that cracks at the peak. He is calling names faster than he can breathe.
Pace: Very fast and accelerating. No gaps, no dead air. Compress the words together at the finish.
Accent: Northern English, Doncaster.
TRANSCRIPT:
And they're into the final furlong, [very fast] Kestrel Lane and Marble Arch stride for stride, nothing between them, nothing at all — [shouting] Marble Arch comes again! Marble Arch on the far rail! Kestrel Lane will not lie down, they are locked together at the line and — [gasp] oh, that is too close to call.Audio & Speech
Synthesize the speech below as a hard-boiled detective's voiceover.
AUDIO PROFILE: Frank Doyle, private investigator, fifty-one, twenty years too tired
SCENE: A one-room office above a laundromat at two in the morning. Rain on the window, a bottle at his elbow, one lamp still burning.
DIRECTOR'S NOTES
Style: Dry, worn down, faintly amused at his own bad luck. He is talking to himself, not to an audience.
Pace: Slow and deliberate, with long pauses between sentences. Let each line land before starting the next.
Accent: Mid-century American, working-class Chicago.
TRANSCRIPT:
She walked in at a quarter past midnight wearing a coat worth more than my car. [sighs] They always come at midnight. That's when the money gets nervous. She told me her husband was missing. [sarcastic] Sure he was. In my experience husbands don't go missing. They go somewhere. And somebody pays me to find out where.Audio & Speech
Synthesize the speech below as a spacecraft's onboard intelligence.
AUDIO PROFILE: HELIOS, onboard intelligence of the deep-space freighter Ardent
SCENE: Reactor containment is failing. The corridor is empty. HELIOS is addressing a crew that may already be gone.
DIRECTOR'S NOTES
Style: Clinical and unhurried. No panic and no warmth, but a thin thread of something almost like regret running underneath the numbers.
Pace: Even and metered, with identical spacing between each clause, like a countdown.
Accent: Neutral and unplaceable.
TRANSCRIPT:
[serious] Attention. Containment integrity is at nineteen percent and falling. Estimated time to breach: four minutes, ten seconds. All personnel should proceed immediately to the aft escape modules. I have attempted to reach the bridge nine hundred and forty times. There has been no response. I will continue trying. [whispers] Please acknowledge, if you can hear me.Audio & Speech
Synthesize the speech below as a wildlife documentary narrator.
AUDIO PROFILE: Margaret Ainsley, natural history narrator, forty years in the field
SCENE: A hedgerow at dusk in the Yorkshire Dales, filmed from six inches away. Nothing dramatic is happening, and she finds that fascinating.
DIRECTOR'S NOTES
Style: Hushed reverence with genuine affection for the animal. Understated — she trusts the picture and never oversells it.
Pace: Measured and patient, slowing further on the closing line.
Accent: Northern English, Yorkshire.
TRANSCRIPT:
[curious] He has been awake for four minutes, and already the evening is going badly. Forty grams of hedgehog, one slug, and a rival twice his size who arrived first. He will not win this. He knows he will not win this. [amused] And yet he tries — because a hedgehog's ambition has never once been troubled by arithmetic.Audio & Speech
API quick start
Run gemini-3.1-tts with a single API call.
POST https://api.wiro.ai/v1/Run/google/gemini-3-1-tts
{
"prompt": "Synthesize the speech below as a live hor…",
"voice": "fenrir"
}curl
curl -X POST "https://api.wiro.ai/v1/Run/google/gemini-3-1-tts" \
-H "Content-Type: application/json" \
-H "x-api-key: YOUR_WIRO_API_KEY" \
--data-binary @- <<'JSON'
{
"prompt": "Synthesize the speech below as a live hor…",
"voice": "fenrir"
}
JSON