Try Google Gemini Omni Flash Video Generator from Google →
Models
Agents
WorkflowsStudioPricingBlogDocs
ExploreDiscover models by categoryBrowse All ModelsBrowse the complete catalogSee FavoritesSign in to view saved models
OverviewThe platform at a glanceLearnSkills, knowledge, guardrailsAnatomyWhat makes agents reasonBuild Your AgentPick skills, set tier, deploy
Pre-built AgentsBrowse the catalog
Agent Usecases
Ad Campaign ManagerApp Event ManagerApp Review RepliesBarber BookingCustomer Win-BackEcommerce ListingsRestaurant Reviews
Sign InStart Building

Task History

Click to see output list

No tasks yet

Go to Models
Explore models/
Video GenerationActive

xai / grok-imagine-r2v

grok-imagine-r2v

byxai

Create short cinematic videos from a text prompt, guided by up to 16 reference images for consistent style and subjects. Built by xAI.

Image to Video
Model ID
grok-imagine-r2v
Provider
xai
Updated
1779108217
7
Comments
Average rating : 5 (7 users)
Providerxai
Modelgrok-imagine-r2v
Image to Video
wiro playground—xai/grok-imagine-r2v
Reset to defaults
Delete All
0 / 7
Maximum 7 image allowed
Drop image to upload

OR

Click to browse your device

Supports: JPG, JPEG, PNG, GIF, WEBP, HEIC

Required. Reference images to guide style and content. Max 20MiB each. Supported: jpg, jpeg, png. Max 7 images.

Required. Text description of the desired video. Use <IMAGE_1>, <IMAGE_2>, etc. to reference specific images.

Sample outputs
Updated 1779108217
## Overview Grok Imagine R2V is a reference-to-video model by xAI. It generates short videos using your prompt plus reference images. The images act as creative direction, not a fixed first frame. ## What you can build - Character-consistent clips from multiple angles and outfits - Style-guided shots, like noir, watercolor, or retro film looks - Product shots that keep brand style while changing scenes - Mood boards and pitch visuals with camera moves and transitions - Remix videos that blend several reference subjects into one scene ## Inputs - Reference images (required) - 1 to 16 images - Supported formats: JPG, JPEG, PNG - Max size: 20 MiB per image - Use tokens like , in the prompt to point at specific images - Prompt (required) - Describe what happens in the clip - Include camera direction, motion, and scene beats - Duration (required) - 5, 10, or 15 seconds - Aspect ratio (required) - 16:9, 9:16, 1:1, 4:3, 3:4, 3:2, 2:3 - Resolution (required) - 480p or 720p ## Outputs - A generated video file in MP4 format (video/mp4) - Frame rate: 24 fps - Width and height match the chosen aspect ratio and resolution ## Recommended settings - Fast iteration - Duration: 5 seconds - Resolution: 480p - Prompt: one subject, one action, one camera move - Higher-quality previews - Duration: 10 seconds - Resolution: 720p - Use 2 to 6 reference images for stable style - Social-first output - Aspect ratio: 9:16 - Duration: 5 or 10 seconds - Prompt: specify close framing and readable subject motion - Prompt tips that work well for R2V - Focus on motion, not appearance - Call out camera moves like “slow push-in” or “crane shot” - Describe actions in time order, one beat per sentence ## Limitations - R2V is a separate mode from image-to-video and video editing. - Very large reference images can fail upload or processing limits. - Some upstream R2V deployments cap duration at 10 seconds. - Generated file URLs may be temporary on some providers. ## Safety & compliance - Generated videos may be filtered by content policy review. - Don’t create sexual content involving minors, or content that suggests it. - Don’t create non-consensual intimate imagery or sexual deepfakes. - Don’t upload images you don’t have rights to use. - Avoid using real-person likeness for deception or impersonation.

Example prompts

Great starting points for grok-imagine-r2v.

Cinematic, sweeping crane shot transitioning from extreme macro to a vast wide-angle landscape. The shot begins at ground level on a bed of dry brown autumn leaves and pine needles, focusing tightly on a vibrant green oak leaf covered in perfectly spherical, physically accurate water droplets. The droplets subtly vibrate with a gentle breeze, reflecting the ambient light. The camera smoothly tilts up and tracks forward, seamlessly racking focus to reveal a serene, misty autumn wetland. Small, still pools of deep blue water perfectly reflect the sky, surrounded by vibrant golden marsh grasses and scattered pine trees. In the distance, rolling green and yellow forested hills rise toward a towering, rugged rocky mountain peak set against a dramatic, heavy cloudy sky. Flawless depth-of-field transition, perfectly stable water reflections, and rigid environmental geometry throughout the camera movement.Video Generation
Cinematic, slow-motion tracking shot moving through a dense, sunlit field of vibrant yellow tulips. A young woman with long dark hair, wearing an elegant white strapless gown with sheer, ruffled lace sleeves, walks gracefully through the blooming flowers. The delicate fabric of her dress catches the gentle breeze, displaying flawless cloth physics. She approaches a young man dressed in a dark black tunic and white turban, who is standing calmly on a dirt path intersecting the field, holding the reins of a majestic white horse with a black and gold saddle. The camera smoothly orbits the scene, maintaining a shallow depth of field that beautifully blurs the foreground and background tulips into a bright, warm bokeh. Highly detailed textures on the lace, the horse\'s realistic muscle movements, and perfectly stable floral geometry throughout the shot.Video Generation
Cinematic, slow-motion medium tracking shot on a brightly lit city street. The background is a highly detailed, weathered yellow wall densely covered in vibrant red graffiti, torn layers of paper, and a large, prominent poster featuring a winged man holding a gold trophy with the clearly legible text \'UN SOFFIO DI FIATO\'. Standing confidently in front of the wall are two stylish Black women. The woman on the left leans slightly against the textured wall; she wears a black leather jacket draped over her shoulder, a red crop top over a white collared shirt, and baggy, faded blue jeans. In her hands, she holds a vintage silver and dark leather rangefinder camera, her fingers naturally turning the metallic lens ring. A gold pocket watch with a purple face hangs from a delicate gold chain attached to her jeans\' belt loop, swaying with accurate pendulum physics. Beside her stands the second woman with an afro puff, wearing a maroon leather jacket, a white cutout top, and black cargo pants, casually holding a delicate sprig of tiny purple heather flowers. The camera captures the deep textures of the leather, denim, and peeling street art. Flawless spatial consistency: the background posters, graffiti, and text remain perfectly rigid and do not warp or hallucinate as the camera slowly moves.Video Generation
Cinematic, sweeping crane shot transitioning from ground level to the high canopy within a dense, lush green forest of towering, moss-covered trees. The shot begins at a low angle, sharply focused on the smooth planks of a sturdy wooden footbridge. Standing alert on the edge of the bridge is a small, highly detailed brown and white bird with distinct black neck bands and a bright yellow-ringed eye. The bird darts its head with rapid, flawless avian mechanics, its feathers ruffling slightly in a gentle breeze. The camera smoothly cranes upward and tracks forward, leaving the bridge and pushing seamlessly through the vibrant green foliage. Dappled sunlight dynamically shifts across the leaves with accurate lighting physics. As the camera reaches the mid-canopy, the focus softly racks to reveal a fluffy grey koala clinging securely to a thin, gently swaying branch. The koala’s thick fur is highly textured and moves naturally in the wind. The koala slowly turns its head to look at the camera, chewing on a green leaf with perfectly realistic jaw and facial mechanics. Flawless spatial consistency, rigorous depth-of-field transitions, and perfectly stable anatomical and environmental geometry without hallucinating or morphing.Video Generation

API quick start

Run grok-imagine-r2v with a single API call.

POST https://api.wiro.ai/v1/Run/xai/grok-imagine-r2v
{
  "prompt": "Cinematic, sweeping crane shot transition…",
  "inputImage": "https://your-cdn.com/input.png",
  "duration": 5,
  "aspectRatio": "16:9"
}
View full API docs

Discover, test, and run AI models, build workflows and agents with one unified API.

All systems operational
WiroAboutBlogCareersContact
ProductModelsAgentsPricingChangelogStatusFAQ
Getting StartedIntroductionAuthenticationProjectsCode ExamplesWiro MCP ServerSelf-Hosted MCPn8n IntegrationLLMs.txt
API ReferenceModelsRun a ModelModel ParametersTasksLLM & Chat StreamingWebSocketRealtime VoiceFiles
© 2026 Wiro AI. All rights reserved.
PrivacyTermsData Deletion