Try Google Gemini Omni Flash Video Generator from Google →
Models
Agents
WorkflowsStudioPricingBlogDocs
ExploreDiscover models by categoryBrowse All ModelsBrowse the complete catalogSee FavoritesSign in to view saved models
OverviewThe platform at a glanceLearnSkills, knowledge, guardrailsAnatomyWhat makes agents reasonBuild Your AgentPick skills, set tier, deploy
Pre-built AgentsBrowse the catalog
Agent Usecases
Ad Campaign ManagerApp Event ManagerApp Review RepliesBarber BookingCustomer Win-BackEcommerce ListingsRestaurant Reviews
Sign InStart Building

Task History

Click to see output list

No tasks yet

Go to Models
Explore models/
Video GenerationActive

xai / grok-imagine-video

grok-imagine-video

byxai

Grok Imagine Video by xAI generates short MP4 videos with synced audio from a text prompt, with an optional first-frame image. Choose duration, aspect, and 480p or 720p.

Text to VideoImage to Video
Model ID
grok-imagine-video
Provider
xai
Updated
1776333262
8
Comments
Average rating : 3.5 (9 users)
Providerxai
Modelgrok-imagine-video
Text to VideoImage to Video
wiro playground—xai/grok-imagine-video
Reset to defaults
0 / 1
Maximum 1 image allowed
Drop image to upload

OR

Click to browse your device

Supports: JPG, JPEG, PNG, GIF, WEBP, HEIC

Optional. First frame image for image-to-video generation. Max 20MiB. Supported: jpg, jpeg, png. Only 1 image allowed.

Required. Text description of the desired video.

Sample outputs
Updated 1776333262
## Overview Grok Imagine Video is a video generation model made by xAI. It creates short videos with audio from a text prompt. You can also provide a first-frame image to guide the shot. ## What you can build - Social clips for TikTok, Reels, and Shorts in common aspect ratios - Cinematic establishing shots with camera moves and lighting direction - Product spins and simple ad-style hero shots from a product photo - Animated portraits from a headshot or character still - Concept trailers and mood boards for pre-visualization ## Inputs - A text prompt describing the scene, action, camera, lighting, and mood - An optional first-frame image for image-to-video generation - Max size: 20 MiB - Formats: JPG, JPEG, PNG - Limit: 1 image - A fixed output duration - Options: 5 seconds, 10 seconds, 15 seconds - An output aspect ratio selection - Options: auto, 16:9, 9:16, 1:1, 4:3, 3:4, 3:2, 2:3 - Auto matches the image aspect ratio when an image is provided - An output resolution selection - Options: 480p, 720p ## Outputs - One video file in MP4 format - The video includes an audio track generated with the clip - Frame rate: 24 fps - Resolution matches your selection up to 720p - Example: 16:9 at 720p is 1280×720 ## Recommended settings - Fast iteration and fewer artifacts - Duration: 5 seconds - Resolution: 480p - Prompt: one subject, one main action, one camera move - Best overall quality for sharing - Duration: 5 or 10 seconds - Resolution: 720p - Aspect ratio: 16:9 for YouTube, 9:16 for mobile-first - Image-to-video (first frame matters) - Aspect ratio: auto, unless you want a forced crop - Prompt: describe motion, not the static details in the image - Add audio cues in the prompt if you care about sound Prompt writing tips that work well: - Use a director style: subject + action + setting + camera + lighting - Call out camera motion: “slow dolly in”, “tracking shot”, “locked-off tripod” - Mention sound when you want it: footsteps, wind, crowd noise, music style ## Limitations - This Wiro listing supports text-to-video and image-to-video with a single first-frame image. - It doesn’t expose video editing, video extension, or multi-reference workflows. - Output resolution tops out at 720p. - Output is short-form video. Longer durations can show more motion drift. - Generated outputs can be blocked or filtered by content moderation. ## Safety & compliance - Follow all applicable laws and platform rules for synthetic media. - Don’t generate sexual content involving minors. Don’t create CSAM. - Don’t create non-consensual intimate imagery or sexual deepfakes. - Get permission before using a real person’s image. - Avoid impersonation, fraud, or misleading political content. - Clearly label AI-generated video when disclosure is required.

Example prompts

Great starting points for grok-imagine-video.

Cinematic macro shot, 120fps slow-motion. A thick, dark espresso is poured from above into a heavy, intricately cut crystal glass filled with irregular, clear ice cubes. The camera is locked on the side of the glass. The amber liquid swirls dynamically and realistically around the ice. The background is completely dark, but a sharp, warm spotlight hits the glass directly from behind. This creates complex, physically accurate light refractions and glowing caustic reflections through the moving liquid and the thick crystal. The solid geometry of the glass and ice remains perfectly rigid and distinct from the fluid.Video Generation
Extreme macro probe lens shot inside the open casing of an antique, brass skeleton pocket watch. Tiny, intricately engraved brass and silver cogs, mainsprings, and escapement wheels are ticking and rotating with flawless, rigid mechanical precision. No parts blend, morph, or melt into one another. The camera slowly pushes in through the complex layers of moving gears. A single, distinct, dusty sunbeam cuts through the dark metallic environment, illuminating floating dust motes and creating a highly realistic, smooth bokeh effect in the background.Video Generation
Cinematic, slow tracking shot following an older man with white hair in light beige clothing as he walks away down a sunlit Italian sidewalk during golden hour. To the right is a vibrant, weathered yellow building featuring a vintage 'FOTOAUTOMATICA' photo booth with a gently swaying red curtain. A cluster of complex directional street signs, including 'Porta Romana', remains perfectly rigid and legible without morphing or hallucinating. Street art, including a sketched cherub and a butterfly, stays structurally fixed to the textured wall. The camera captures deep, realistic cast shadows. In the softly blurred background, other pedestrians stroll naturally along a narrow street lined with parked cars. Flawless spatial consistency, accurate human gait mechanics, and precise architectural geometryVideo Generation
Cinematic, low-angle tracking shot following a small, fluffy golden-brown dog briskly trotting from right to left across a wet asphalt forest path. The dog's long fur and curled tail bounce dynamically and naturally with every step, showcasing complex hair physics. The dog pants slightly with its mouth open, looking ahead. It is lightly raining, with visible rain streaks falling. The wet road creates soft, realistic reflections of the dog's paws, which kick up tiny, physically accurate droplets of water upon impact. The background remains a deeply blurred, atmospheric autumn forest fading into mist, maintaining consistent depth of field. Flawless anatomical geometry of the dog's legs and gait.Video Generation

API quick start

Run grok-imagine-video with a single API call.

POST https://api.wiro.ai/v1/Run/xai/grok-imagine-video
{
  "prompt": "Cinematic, slow tracking shot following a…",
  "duration": 5,
  "aspectRatio": "auto",
  "resolution": "480p"
}
View full API docs

Discover, test, and run AI models, build workflows and agents with one unified API.

All systems operational
WiroAboutBlogCareersContact
ProductModelsAgentsPricingChangelogStatusFAQ
Getting StartedIntroductionAuthenticationProjectsCode ExamplesWiro MCP ServerSelf-Hosted MCPn8n IntegrationLLMs.txt
API ReferenceModelsRun a ModelModel ParametersTasksLLM & Chat StreamingWebSocketRealtime VoiceFiles
© 2026 Wiro AI. All rights reserved.
PrivacyTermsData Deletion