Try MiniMax FastH3 from Fastvideo →
Models
Agents
Workflows
Studio
PricingBlogDocs
ExploreDiscover models by categoryBrowse All ModelsBrowse the complete catalogSee FavoritesSign in to view saved models
Generative Media AgentCreate and edit media by chattingWorkflow AgentBuild visual workflows with Agent
OverviewThe platform at a glanceLearnSkills, knowledge, guardrailsAnatomyWhat makes agents reasonBuild Your AgentPick skills, set tier, deploy
Pre-built AgentsBrowse the catalog
Agent Usecases
Ad Campaign ManagerApp Event ManagerApp Review RepliesBarber BookingCustomer Win-BackEcommerce ListingsRestaurant Reviews
Sign InStart Building

Task History

Click to see output list

No tasks yet

Go to Models
Explore models/
Audio & SpeechActive

google / gemini-3.1-tts

gemini-3.1-tts

bygoogle

Generate natural speech from text with Gemini 3.1 Flash TTS. Use voice options and expressive tags to control tone and pacing for narration, apps, and accessibility.

Text to Speech
Model ID
gemini-3.1-tts
Provider
google
Updated
1788788171
gemini-3.1-tts
5
Comments
Average rating : 5 (2 users)
Providergoogle
Modelgemini-3.1-tts
Text to Speech
wiro playground—google/gemini-3.1-tts
Reset to defaults

Required

Required

Sample outputs
google-gemini-3-1-tts-sample-5.mp3
google-gemini-3-1-tts-sample-6.mp3
google-gemini-3-1-tts-sample-7.mp3
google-gemini-3-1-tts-sample-8.mp3
Updated 1788788171

Overview

Gemini 3.1 Flash TTS is a text-to-speech model from Google. It uses a large language model to decide both the words and the delivery. You steer performance with plain text direction plus inline square-bracket tags. It outputs spoken audio that’s suitable for voiceovers, narration, and spoken UI.

The model is built for controllable speech. You can change pacing, emotion, and delivery mid-sentence. This helps you ship audio that sounds intentional, not flat.

What you can build

  • App voiceovers with a consistent “brand voice”
  • Audiobook-style narration with pacing and pause control
  • Character dialogue for games (single-voice on this Wiro page)
  • Accessibility narration for articles, screens, and UI flows
  • Multilingual announcements and help prompts
  • Podcast-style scripted segments where exact wording matters

Inputs

  • The script you want spoken, provided as plain text. Keep it under roughly 16K tokens of text for the Gemini 3.1 Flash TTS model.
  • A style direction written directly into the script, like “Say cheerfully: …” so the model knows how to perform.
  • Inline audio tags in square brackets, placed inside the script where you want the change. Examples include tags for laughter, whispering, emotion, and pauses. Google documents 200+ supported tags.
  • One prebuilt voice, selected from 30 options: Zephyr, Puck, Charon, Kore, Fenrir, Leda, Orus, Aoede, Callirrhoe, Autonoe, Enceladus, Iapetus, Umbriel, Algieba, Despina, Erinome, Algenib, Rasalgethi, Laomedeia, Achernar, Alnilam, Schedar, Gacrux, Pulcherrima, Achird, Zubenelgenubi, Vindemiatrix, Sadachbia, Sadaltager, Sulafat.

Outputs

  • A single speech audio result that reads your script in the selected voice.
  • The output is delivered as a playable audio file on Wiro.
  • The underlying Gemini TTS examples commonly use mono 24,000 Hz, 16-bit PCM audio wrapped in a WAV container.
  • Audio generated by Gemini 3.1 Flash TTS is watermarked with SynthID.

Limitations

  • This model is in Preview on Google’s side. Behavior and limits can change.
  • It’s text-to-speech. It’s meant to read provided text, not invent a new script.
  • Multi-speaker dialogue exists in Google’s Gemini TTS offering, but it’s limited to 2 speakers in the Gemini API. This Wiro page exposes a single voice choice.
  • If you use Gemini TTS through Google Cloud Text-to-Speech, the “text” and “prompt” fields can each be at most 4,000 bytes, with 8,000 bytes combined. That path can truncate output around 655 seconds.
  • Very long scripts can drift in prosody or volume over time. Split long narration into shorter chunks.
  • Messy input text can sound bad. Expect problems with OCR errors, missing punctuation, or inconsistent casing.

Safety & compliance

  • Gemini 3.1 Flash TTS audio includes SynthID watermarking to help detect AI-generated speech.
  • Don’t use it for deception. Don’t impersonate real people, mislead listeners, or hide that audio is synthetic.
  • Follow Google’s generative AI policies. Avoid disallowed content such as hate, harassment, sexual content, or instructions for wrongdoing.

Example prompts

Great starting points for gemini-3.1-tts.

Synthesize the speech below as a live horse-racing commentator. AUDIO PROFILE: Ray Tulloch, veteran racecourse commentator SCENE: The final furlong at Doncaster. Forty thousand people on their feet, hooves thundering, the two leaders inseparable. DIRECTOR'S NOTES Style: Controlled chaos. Rising urgency that cracks at the peak. He is calling names faster than he can breathe. Pace: Very fast and accelerating. No gaps, no dead air. Compress the words together at the finish. Accent: Northern English, Doncaster. TRANSCRIPT: And they're into the final furlong, [very fast] Kestrel Lane and Marble Arch stride for stride, nothing between them, nothing at all — [shouting] Marble Arch comes again! Marble Arch on the far rail! Kestrel Lane will not lie down, they are locked together at the line and — [gasp] oh, that is too close to call.Audio & Speech
Synthesize the speech below as a hard-boiled detective's voiceover. AUDIO PROFILE: Frank Doyle, private investigator, fifty-one, twenty years too tired SCENE: A one-room office above a laundromat at two in the morning. Rain on the window, a bottle at his elbow, one lamp still burning. DIRECTOR'S NOTES Style: Dry, worn down, faintly amused at his own bad luck. He is talking to himself, not to an audience. Pace: Slow and deliberate, with long pauses between sentences. Let each line land before starting the next. Accent: Mid-century American, working-class Chicago. TRANSCRIPT: She walked in at a quarter past midnight wearing a coat worth more than my car. [sighs] They always come at midnight. That's when the money gets nervous. She told me her husband was missing. [sarcastic] Sure he was. In my experience husbands don't go missing. They go somewhere. And somebody pays me to find out where.Audio & Speech
Synthesize the speech below as a spacecraft's onboard intelligence. AUDIO PROFILE: HELIOS, onboard intelligence of the deep-space freighter Ardent SCENE: Reactor containment is failing. The corridor is empty. HELIOS is addressing a crew that may already be gone. DIRECTOR'S NOTES Style: Clinical and unhurried. No panic and no warmth, but a thin thread of something almost like regret running underneath the numbers. Pace: Even and metered, with identical spacing between each clause, like a countdown. Accent: Neutral and unplaceable. TRANSCRIPT: [serious] Attention. Containment integrity is at nineteen percent and falling. Estimated time to breach: four minutes, ten seconds. All personnel should proceed immediately to the aft escape modules. I have attempted to reach the bridge nine hundred and forty times. There has been no response. I will continue trying. [whispers] Please acknowledge, if you can hear me.Audio & Speech
Synthesize the speech below as a wildlife documentary narrator. AUDIO PROFILE: Margaret Ainsley, natural history narrator, forty years in the field SCENE: A hedgerow at dusk in the Yorkshire Dales, filmed from six inches away. Nothing dramatic is happening, and she finds that fascinating. DIRECTOR'S NOTES Style: Hushed reverence with genuine affection for the animal. Understated — she trusts the picture and never oversells it. Pace: Measured and patient, slowing further on the closing line. Accent: Northern English, Yorkshire. TRANSCRIPT: [curious] He has been awake for four minutes, and already the evening is going badly. Forty grams of hedgehog, one slug, and a rival twice his size who arrived first. He will not win this. He knows he will not win this. [amused] And yet he tries — because a hedgehog's ambition has never once been troubled by arithmetic.Audio & Speech

API quick start

Run gemini-3.1-tts with a single API call.

POST https://api.wiro.ai/v1/Run/google/gemini-3.1-tts
{
  "prompt": "Synthesize the speech below as a live hor…",
  "voice": "fenrir"
}
curl
curl -X POST "https://api.wiro.ai/v1/Run/google/gemini-3.1-tts" \
  -H "Content-Type: application/json" \
  -H "x-api-key: YOUR_WIRO_API_KEY" \
  --data-binary @- <<'JSON'
{
  "prompt": "Synthesize the speech below as a live hor…",
  "voice": "fenrir"
}
JSON

Pricing: Final costs are determined after the task is finished. Input: $1.00 per 1M tokens (text). Output: $20.00 per 1M tokens (audio).

View full API docs

Discover, test, and run AI models, build workflows and agents with one unified API.

All systems operational
WiroAboutBlogCareersContact
ProductModelsAgentsPricingPartnerChangelogStatusFAQ
Getting StartedIntroductionAuthenticationProjectsCode ExamplesWiro MCP ServerSelf-Hosted MCPn8n IntegrationLLMs.txt
API ReferenceModelsRun a ModelModel ParametersTasksLLM & Chat StreamingWebSocketRealtime VoiceFiles
© 2026 Wiro AI. All rights reserved.
PrivacyTermsData Deletion