Try MiniMax H3 (Text-to-Video) (Image-to-Video) from MiniMax →
Models
Agents
WorkflowsStudioPricingBlogDocs
ExploreDiscover models by categoryBrowse All ModelsBrowse the complete catalogSee FavoritesSign in to view saved models
OverviewThe platform at a glanceLearnSkills, knowledge, guardrailsAnatomyWhat makes agents reasonBuild Your AgentPick skills, set tier, deploy
Pre-built AgentsBrowse the catalog
Agent Usecases
Ad Campaign ManagerApp Event ManagerApp Review RepliesBarber BookingCustomer Win-BackEcommerce ListingsRestaurant Reviews
Sign InStart Building

Task History

Click to see output list

No tasks yet

Go to Models
Explore models/
Video GenerationActive

MiniMax / h3-r2v

h3-r2v

byminimax

MiniMax H3 R2V generates 4–15s videos with synced stereo audio from reference images, clips, and optional audio. Use it for consistent characters, style, and motion.

Image to VideoVideo to Video
Model ID
h3-r2v
Provider
minimax
Updated
1785946082
11
Comments
Average rating : 5 (20 users)
Providerminimax
Modelh3-r2v
Image to VideoVideo to Video
wiro playground—minimax/h3-r2v
Reset to defaults
Delete All
0 / 9
Maximum 9 image allowed
Drop image to upload

OR

Click to browse your device

Supports: JPG, JPEG, PNG, GIF, WEBP, HEIC

Optional. Subject or style references. Up to 9 images. Cite them in the prompt as Image 1, Image 2, and so on. At least one image or video is required when audio is used. Images, videos, and audio together max 12 files.

Delete All
0 / 3
Maximum 3 video allowed
Drop video to upload

OR

Click to browse your device

Supports: MP4, WEBM, MOV

Optional. Motion references. Up to 3 clips, each 2–15 seconds, total at most 15 seconds. Cite them in the prompt as Video 1, Video 2, and so on.

Required. Describe the video and refer to inputs by order: Image 1, Video 1, Audio 1, and so on.

Required. Length of the video in whole seconds, between 4 and 15.

Required. Default 768P. Choose 2K for a higher resolution result.

Sample outputs
Updated 1785946082
## Overview MiniMax H3 R2V is MiniMax’s reference-to-video generation model. It reads your prompt plus optional reference images, video clips, and audio clips in one context. It generates a short MP4 video with native stereo audio, not a separate audio pass. This helps when you need stable identity, style, motion, and voice across a single shot. ## What you can build - Product hero shots that keep the product design identical while the camera moves - Character-consistent clips from multiple reference images (outfit, face, props) - Motion transfer using a short reference video as movement or pacing guidance - Style transfer using a style board image plus a separate subject image - Video edits driven by a source clip, like replacing a subject while keeping motion - Dialogue and sound effects timed to the action inside one generated clip ## Inputs - Reference images you want the model to follow, such as a subject, outfit, or style board (JPG, PNG, WEBP, HEIC, or HEIF). You can upload up to 9 images. - Reference video clips you want to use for motion, pacing, or as a source for editing (MP4 or MOV). You can upload up to 3 clips. Each clip must be 2–15 seconds, with 15 seconds total. - Reference audio clips you want to use for voice or sound guidance (WAV or MP3). You can upload up to 3 clips. Each clip must be 2–15 seconds, with 15 seconds total. Audio can’t be the only reference. - A text prompt that describes the target video. Cite your uploads by their order, like “Image 1”, “Video 1”, and “Audio 1”. - The target clip length in whole seconds. It must be between 4 and 15 seconds. - An aspect ratio choice. You can let the model adapt to your references or force a fixed ratio like 16:9 or 9:16. - An output resolution choice. You can generate at 768P or 2K. ## Outputs The model returns a single MP4 video. The file includes the video stream plus a synced stereo audio track. Video is generated at 24 fps. Audio is 32 kHz stereo. The duration matches your requested 4–15 seconds. ## Limitations - Maximum clip length is 15 seconds. - Reference inputs have hard caps: 9 images, 3 videos, and 3 audio clips, with 12 files total. - Audio references must be paired with at least one image or video reference. - Reference-to-video prompting is sensitive to how you name and assign each reference. Vague prompts can ignore parts of your intent. - Conflicting references can cause identity drift, style blending, or sudden changes between shots. - Low-quality inputs can degrade results. This includes compressed videos, noisy audio, low-resolution images, and scanned or heavily filtered assets. ## Safety & compliance MiniMax applies automated moderation to user inputs and generated prompts. Content suspected of being unlawful, pornographic, or infringing third-party rights may be blocked. Filtering can still produce false positives and false negatives. You are responsible for lawful use, rights clearance, and following the MiniMax H3 Community License terms.

Example prompts

Great starting points for h3-r2v.

Image 1 is the watch. Keep the design identical while the second hand ticks and the camera slowly orbits around it, premium product film, soft reflections.Video Generation
Image 1 is the barista. Keep her face and apron consistent as she pours steamed milk into a cup, making latte art, soft steam drifting up, natural handheld feel.Video Generation
Video 1 is the motion and camera path. Keep the same walk and framing but set it at night in a neon wet street market, rain on the ground, cinematic.Video Generation
Video 1 is the motion and camera path. Keep the same downhill race and tracking shot, but replace the rider outfit with red and blue and make the helmet white, same rocky trail and dust, cinematic.Video Generation

API quick start

Run h3-r2v with a single API call.

POST https://api.wiro.ai/v1/Run/MiniMax/h3-r2v
{
  "prompt": "Image 1 is the watch. Keep the design ide…",
  "inputImage": "https://your-cdn.com/input.png",
  "inputVideo": "https://your-cdn.com/input.mp4",
  "inputAudio": "https://your-cdn.com/input.mp3"
}
View full API docs

Discover, test, and run AI models, build workflows and agents with one unified API.

All systems operational
WiroAboutBlogCareersContact
ProductModelsAgentsPricingPartnerChangelogStatusFAQ
Getting StartedIntroductionAuthenticationProjectsCode ExamplesWiro MCP ServerSelf-Hosted MCPn8n IntegrationLLMs.txt
API ReferenceModelsRun a ModelModel ParametersTasksLLM & Chat StreamingWebSocketRealtime VoiceFiles
© 2026 Wiro AI. All rights reserved.
PrivacyTermsData Deletion