Try MiniMax FastH3 from Fastvideo →
Models
Agents
Workflows
Studio
PricingBlogDocs
ExploreDiscover models by categoryBrowse All ModelsBrowse the complete catalogSee FavoritesSign in to view saved models
Generative Media AgentCreate and edit media by chattingWorkflow AgentBuild visual workflows with Agent
OverviewThe platform at a glanceLearnSkills, knowledge, guardrailsAnatomyWhat makes agents reasonBuild Your AgentPick skills, set tier, deploy
Pre-built AgentsBrowse the catalog
Agent Usecases
Ad Campaign ManagerApp Event ManagerApp Review RepliesBarber BookingCustomer Win-BackEcommerce ListingsRestaurant Reviews
Sign InStart Building

Task History

Click to see output list

No tasks yet

Go to Models
Explore models/
Video GenerationActive

alibaba / wan 3.0 prime r2v

wan 3.0 prime r2v

byalibaba

Wan 3.0 Prime R2V by Alibaba generates 2–30s MP4 videos from prompts plus reference images, clips, and audio for consistent identity and style.

Image to VideoVideo to Video
Model ID
wan 3.0 prime r2v
Provider
alibaba
Updated
1789737310
14
Comments
Average rating : 3.5 (17 users)
Provideralibaba
Modelwan 3.0 prime r2v
Image to VideoVideo to Video
wiro playground—alibaba/wan 3.0 prime r2v
Reset to defaults
0 / 10
Maximum 10 image allowed
Delete All

Optional. Up to 10 subject or style references. JPEG/JPG/PNG (no transparency)/BMP/WEBP, 240-8000px, ratio up to 8:1, max 20MB.

Required. The prompt to generate the video. You can point to the inputs by index, like "Image 1" or "Video 1".

Required. Total duration of the generated video (in seconds)

Required. Quality tier for the generated video

Optional. Aspect ratio of the output. Adaptive derives it from the prompt and the input media.

Sample outputs
Updated 1789737310

Overview

Wan 3.0 Prime R2V is Alibaba’s reference-to-video model in the Wan 3.0 series. You provide a detailed prompt plus optional references. The model blends them into a single MP4 video at 30 fps. It’s useful when you need stable identity, style, and voice across shots.

What you can build

  • Character-consistent clips using a small set of reference images
  • Product videos that match brand style from reference frames
  • Motion-guided scenes using a short reference video as movement inspiration
  • Dialogue-driven scenes that mimic a reference voice from an audio clip
  • Multi-shot sequences written as a timed script inside one prompt

Inputs

  • A text prompt that describes the scene, actions, and camera language. You can refer to references by index, like “Image 1” or “Video 1”.
  • Optional reference images to lock subject identity or style. Provide up to 10 images in JPEG/JPG/PNG (no transparency), BMP, or WEBP. Each image must be 240–8000 px per side, up to 20 MB, and no wider than 8:1.
  • Optional reference videos to guide motion, pacing, or framing. Provide up to 5 clips in MP4 or MOV. Total reference video time must stay within 15 seconds. Each clip must be at least 16 fps and no larger than 100 MB.
  • Optional reference audio to guide voice timbre or soundtrack feel. Provide up to 5 WAV or MP3 files. Total reference audio time must stay within 15 seconds. Each file must be no larger than 15 MB.
  • A target output duration in seconds. The generated video can be 2–30 seconds long. If you include reference video, keep input plus output within 30 seconds total.
  • A resolution tier for the output video: 480p, 720p, or 1080p.
  • An aspect ratio choice. “Adaptive” lets the model choose framing. Fixed ratios include 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16.
  • An option to include or omit an audio track in the output video.
  • An optional prompt expansion setting that rewrites short prompts to add detail.
  • An optional random seed to help repeat a result range.

Outputs

  • A generated MP4 video file.
  • The output includes the final width and height, frame rate (30 fps), and the realized duration.
  • The response also includes the seed used for the generation, which helps you rerun similar attempts.

Recommended settings

  • For character consistency: use 3–8 reference images with varied angles and lighting.
  • For motion control: add 1 short reference video and keep the output at 5–10 seconds.
  • For short prompts: keep prompt expansion on, then edit the expanded prompt if needed.
  • For iteration: start at 480p, then rerun at 1080p after you like composition.
  • For repeatable experiments: set a fixed seed and change one input at a time.

Limitations

  • Reference video and reference audio have a hard total limit of 15 seconds each.
  • If you include reference video, input plus output duration must stay within 30 seconds.
  • PNG transparency is not supported for reference images.
  • A fixed seed can’t guarantee identical output across runs.
  • Blurry, low-resolution, watermarked, or heavily compressed references reduce identity stability.
  • Noisy audio, overlapping speech, or strong background music can degrade voice matching.
  • Some providers return a download link that expires after about 24 hours. Save outputs promptly.
  • Some user interfaces may force an “AI generated” watermark for compliance.

Safety & compliance

  • Many deployments apply automated safety checks to both inputs and outputs.
  • Don’t upload references you don’t have rights to use.
  • Get consent before using a real person’s likeness or voice.
  • Label AI-generated video where your policy or local law requires it.
  • Avoid illegal content and prohibited harm, including impersonation for fraud.

Example prompts

Great starting points for wan 3.0 prime r2v.

A teenager in a loose t-shirt picks up Image 1, drops it on the ground and pushes off across a sunlit skatepark. He races toward a set of stairs, pops the board into a clean kickflip in slow motion and lands it cleanly, rolling away as dust kicks up behind the wheels. Handheld camera tracking low beside him, warm late afternoon light, wheels rumbling on concrete and the crack of the board landing.Video Generation
A young man sitting on a log by a night campfire picks up Image 1, settles it on his knee and begins to play, firelight flickering across the wood grain and his hands. Friends around the fire nod along and sparks drift up into the dark trees. The camera slowly circles the fire. Warm cinematic firelight, crackling flames and soft acoustic strumming.Video Generation
Image 1 perches on a weathered railing overlooking a misty jungle waterfall at sunrise. It ruffles its feathers, turns its head toward the camera and lets out a sharp call, then spreads its wings and launches off the railing into the mist. The camera follows it into the open air. Vivid color, golden morning light, rushing water and bird calls.Video Generation
Image 1 rests on a slab of dark wet slate as the camera slowly orbits it in extreme close-up. Thin beams of light sweep across the crystal and the steel case, water droplets bead on the surface, and the second hand ticks forward while faint mist rolls past. The camera pulls back to reveal the watch alone in a pool of light. Luxury commercial look, dramatic lighting, soft ticking.Video Generation

API quick start

Run wan 3.0 prime r2v with a single API call.

POST https://api.wiro.ai/v1/Run/alibaba/wan-3-0-prime-r2v
{
  "prompt": "A teenager in a loose t-shirt picks up Im…",
  "duration": 5,
  "resolution": "480P",
  "inputImage": "https://your-cdn.com/input.png"
}
curl
curl -X POST "https://api.wiro.ai/v1/Run/alibaba/wan-3-0-prime-r2v" \
  -H "Content-Type: application/json" \
  -H "x-api-key: YOUR_WIRO_API_KEY" \
  --data-binary @- <<'JSON'
{
  "prompt": "A teenager in a loose t-shirt picks up Im…",
  "duration": 5,
  "resolution": "480P",
  "inputImage": "https://your-cdn.com/input.png"
}
JSON
View full API docs

Discover, test, and run AI models, build workflows and agents with one unified API.

All systems operational
WiroAboutBlogCareersContact
ProductModelsAgentsPricingPartnerChangelogStatusFAQ
Getting StartedIntroductionAuthenticationProjectsCode ExamplesWiro MCP ServerSelf-Hosted MCPn8n IntegrationLLMs.txt
API ReferenceModelsRun a ModelModel ParametersTasksLLM & Chat StreamingWebSocketRealtime VoiceFiles
© 2026 Wiro AI. All rights reserved.
PrivacyTermsData Deletion