Try MiniMax FastH3 from Fastvideo →
Models
Agents
Workflows
Studio
PricingBlogDocs
ExploreDiscover models by categoryBrowse All ModelsBrowse the complete catalogSee FavoritesSign in to view saved models
Generative Media AgentCreate and edit media by chattingWorkflow AgentBuild visual workflows with Agent
OverviewThe platform at a glanceLearnSkills, knowledge, guardrailsAnatomyWhat makes agents reasonBuild Your AgentPick skills, set tier, deploy
Pre-built AgentsBrowse the catalog
Agent Usecases
Ad Campaign ManagerApp Event ManagerApp Review RepliesBarber BookingCustomer Win-BackEcommerce ListingsRestaurant Reviews
Sign InStart Building

Task History

Click to see output list

No tasks yet

Go to Models
Explore models/
Video GenerationActive

wiro / video-caption

video-caption

bywiro

Burn captions into a video with TikTok-style timing. Transcribe spoken audio word-by-word or overlay fixed text, with full control over caption style.

Video to VideoUtility
Model ID
video-caption
Provider
wiro
Updated
1788870099
video-caption
5
Comments
Average rating : 5 (2 users)
Providerwiro
Modelvideo-caption
Video to VideoUtility
wiro playground—wiro/video-caption
Reset to defaults
0 / 1
Maximum 1 video allowed

Required: the video to caption. mp4, mov, mkv and m4v keep their own format; anything else comes back as mp4.

Required. Default 'speech': the speech in the video is transcribed and captioned word by word, timed to the voice. 'fixed' burns the caption text you write below instead, and needs no audio at all.

Required. Choose how captions appear: 'Rolling Line' keeps the last three words on screen and drops the oldest as the next one arrives, 'Flash Each Word' shows words briefly, 'Sentence' displays the whole line at once - which in 'fixed' mode leaves it on screen from start to end.

Required with 'fixed' mode: the line burned onto the video. Left empty, the run falls back to 'speech' and captions the video's own speech instead. Optional with 'speech', where it is used only if no speech can be made out at all.

Sample outputs

No samples yet

Run this model to create outputs and build up samples.

Updated 1788870099

Overview

Video Caption by Wiro adds burned-in captions to an existing video. It can transcribe the video’s speech and time each word to the voice, then render those words onto the frames. It can also burn a fixed caption line you provide, even if the video has no audio. The output solves the last-mile problem for social clips: you get a share-ready video with captions already embedded.

What you can build

  • TikTok, Reels, and Shorts clips with word-timed captions
  • Silent promo videos with a fixed headline burned onto the footage
  • Accessibility-friendly exports for platforms that don’t show subtitle tracks well
  • Creator templates with consistent caption font, color, and placement
  • Quick “quote” videos where the same line stays on screen

Inputs

  • One input video you want to caption. Provide an uploaded file or a direct URL. If the input is MP4, MOV, MKV, or M4V, it keeps the same container on export; other formats typically return as MP4.
  • A captioning mode: either transcribe the speech in the video and caption it with timestamps, or burn the fixed caption text you provide.
  • A caption display effect that controls how text appears on screen. Options include a rolling word window, flashing each word, or showing a full sentence at once.
  • Optional caption text you write. This is required when you choose fixed-text captioning. If you leave it blank, the run falls back to speech-based captions.
  • Optional caption placement on the video: top, center, or bottom.
  • An optional font choice from the Poppins font family.
  • An optional font size in pixels from 8 to 400. If you leave it empty, the model scales the text based on the video size.
  • An optional text color as a hex code (for example, #FFFFFF).
  • An optional background box color as a hex code (for example, #000000).
  • An optional background box opacity from 0.0 to 1.0. Set 0.0 for no box.
  • An optional margin from the top or bottom edge in pixels from 0 to 1000. Leave it empty to scale with video size. This setting is ignored when captions are centered.
  • For the rolling-word effect, an optional maximum number of words to keep on screen from 1 to 20.

Outputs

The model returns a single captioned video file.

The video includes captions burned into the frames, so it plays correctly on any platform without needing a separate subtitle track. For MP4, MOV, MKV, and M4V inputs, it preserves the original container; for other input types, it typically exports as MP4.

Recommended settings

  • For TikTok-style captions on spoken clips, use speech-based captioning with the “flash each word” effect.
  • For fast speakers, use the rolling-word effect and keep the on-screen word window small (often 3–5 words).
  • If your video has busy backgrounds, add a caption box and set opacity around 0.3 to 0.6 for readability.
  • If you need consistent visuals across many resolutions, leave font size empty so it scales with the input video.
  • If platform UI covers the bottom area (common on Shorts), move captions to the top or raise the bottom margin.
  • For a title card that stays up the whole time, use fixed-text captioning with the sentence effect.

Limitations

  • Speech-based captions depend on audio quality. Heavy noise, music, or overlapping speakers can reduce accuracy and timing.
  • Fixed-text captioning does not auto-time words to speech. It burns the same line you provide.
  • The sentence effect can keep the full line on screen for the entire clip in fixed-text mode, which may not fit fast-paced edits.
  • Since captions are burned in, you can’t turn them off later or restyle them without re-running the model.
  • Low-quality inputs can hurt results. This includes heavy compression, clipped audio, very quiet speech, and videos with existing hardcoded subtitles.

Safety & compliance

  • Only upload videos you own or have permission to edit and republish.
  • Captions can reveal personal data spoken in the clip. Review outputs before sharing.
  • Wiro file URLs are designed to be accessible by link, so treat output links as public if shared. (wiro.ai)
  • If you need to remove a run’s media, Wiro provides a task-level option to delete both uploaded inputs and generated outputs after completion. (wiro.ai)

Example prompts

Great starting points for video-caption.

Synthesize the speech below as a live horse-racing commentator. AUDIO PROFILE: Ray Tulloch, veteran racecourse commentator SCENE: The final furlong at Doncaster. Forty thousand people on their feet, hooves thundering, the two leaders inseparable. DIRECTOR'S NOTES Style: Controlled chaos. Rising urgency that cracks at the peak. He is calling names faster than he can breathe. Pace: Very fast and accelerating. No gaps, no dead air. Compress the words together at the finish. Accent: Northern English, Doncaster. TRANSCRIPT: And they're into the final furlong, [very fast] Kestrel Lane and Marble Arch stride for stride, nothing between them, nothing at all — [shouting] Marble Arch comes again! Marble Arch on the far rail! Kestrel Lane will not lie down, they are locked together at the line and — [gasp] oh, that is too close to call.Video Generation
Synthesize the speech below as a hard-boiled detective's voiceover. AUDIO PROFILE: Frank Doyle, private investigator, fifty-one, twenty years too tired SCENE: A one-room office above a laundromat at two in the morning. Rain on the window, a bottle at his elbow, one lamp still burning. DIRECTOR'S NOTES Style: Dry, worn down, faintly amused at his own bad luck. He is talking to himself, not to an audience. Pace: Slow and deliberate, with long pauses between sentences. Let each line land before starting the next. Accent: Mid-century American, working-class Chicago. TRANSCRIPT: She walked in at a quarter past midnight wearing a coat worth more than my car. [sighs] They always come at midnight. That's when the money gets nervous. She told me her husband was missing. [sarcastic] Sure he was. In my experience husbands don't go missing. They go somewhere. And somebody pays me to find out where.Video Generation
Synthesize the speech below as a spacecraft's onboard intelligence. AUDIO PROFILE: HELIOS, onboard intelligence of the deep-space freighter Ardent SCENE: Reactor containment is failing. The corridor is empty. HELIOS is addressing a crew that may already be gone. DIRECTOR'S NOTES Style: Clinical and unhurried. No panic and no warmth, but a thin thread of something almost like regret running underneath the numbers. Pace: Even and metered, with identical spacing between each clause, like a countdown. Accent: Neutral and unplaceable. TRANSCRIPT: [serious] Attention. Containment integrity is at nineteen percent and falling. Estimated time to breach: four minutes, ten seconds. All personnel should proceed immediately to the aft escape modules. I have attempted to reach the bridge nine hundred and forty times. There has been no response. I will continue trying. [whispers] Please acknowledge, if you can hear me.Video Generation
Synthesize the speech below as a wildlife documentary narrator. AUDIO PROFILE: Margaret Ainsley, natural history narrator, forty years in the field SCENE: A hedgerow at dusk in the Yorkshire Dales, filmed from six inches away. Nothing dramatic is happening, and she finds that fascinating. DIRECTOR'S NOTES Style: Hushed reverence with genuine affection for the animal. Understated — she trusts the picture and never oversells it. Pace: Measured and patient, slowing further on the closing line. Accent: Northern English, Yorkshire. TRANSCRIPT: [curious] He has been awake for four minutes, and already the evening is going badly. Forty grams of hedgehog, one slug, and a rival twice his size who arrived first. He will not win this. He knows he will not win this. [amused] And yet he tries — because a hedgehog's ambition has never once been troubled by arithmetic.Video Generation

API quick start

Run video-caption with a single API call.

POST https://api.wiro.ai/v1/Run/wiro/video-caption
{
  "inputVideo": "https://your-cdn.com/input.mp4",
  "mode": "speech",
  "captionEffect": "wordbyword",
  "text": "..."
}
curl
curl -X POST "https://api.wiro.ai/v1/Run/wiro/video-caption" \
  -H "Content-Type: application/json" \
  -H "x-api-key: YOUR_WIRO_API_KEY" \
  --data-binary @- <<'JSON'
{
  "inputVideo": "https://your-cdn.com/input.mp4",
  "mode": "speech",
  "captionEffect": "wordbyword",
  "text": "..."
}
JSON

Pricing: 'fixed' caption costs 0.001 per video, flat. 'speech' costs 0.004 per started minute of video, since the speech has to be transcribed: up to 1:00 is 0.004, 1:01 to 2:00 is 0.008, 2:01 to 3:00 is 0.012, and so on. Only successful outputs are charged.

View full API docs

Discover, test, and run AI models, build workflows and agents with one unified API.

All systems operational
WiroAboutBlogCareersContact
ProductModelsAgentsPricingPartnerChangelogStatusFAQ
Getting StartedIntroductionAuthenticationProjectsCode ExamplesWiro MCP ServerSelf-Hosted MCPn8n IntegrationLLMs.txt
API ReferenceModelsRun a ModelModel ParametersTasksLLM & Chat StreamingWebSocketRealtime VoiceFiles
© 2026 Wiro AI. All rights reserved.
PrivacyTermsData Deletion