Kling V3 Motion Control was tested here as a motion-transfer tool, not as a general text-to-video generator. The question was simple: can one reference photo keep a recognizable person while a separate driving clip supplies the body movement? Three existing outputs put that idea under different visual pressure: a clean studio move, a relit club scene, and a fashion shot with moving hair and confetti. The clips below are the original model outputs, published without manual retouching.
Model
Kling V3 Motion Control on Wiro accepts a reference image, a driving video, and an optional text prompt. The image supplies the person, setting, and other visual anchors. The video supplies the action. The prompt can add or alter details such as lighting, atmosphere, and styling. That division matters: a prompt can guide the result, but it cannot reliably replace a clear source image or readable driving motion.
What the test set out to check
The three tests were designed to separate three common failure points in image-driven motion transfer. First, identity retention: does the face, silhouette, and wardrobe still read as the person in the input image? Second, motion adherence: do the arms, torso, and steps follow the driving clip instead of becoming a generic dance? Third, scene edits: can the prompt change light, atmosphere, or small environmental details without causing the subject to fall apart?
This is a useful distinction from a prompt-first video test. Kling V3 vs Veo 3.1 Fast looks at text-led generation. Here, the source video is the motion plan. It gives the model far less room to invent choreography, and it makes errors at hands, hair, and object boundaries easier to spot.
Inputs and settings
- Mode:
std, the standard option used for these runs. - Reference image: one portrait image per test, used for the subject identity and base scene.
- Driving video: one motion reference per test, used for the character action.
- Original sound:
no. - Character orientation:
video, so the result follows the orientation in the driving video. The model documentation says this setting supports reference videos up to 30 seconds; theimagesetting instead preserves the image orientation and caps the video at 10 seconds. - Post-production: none. The embedded clips are the outputs as generated.
The model documentation also lists image support for JPG, JPEG, and PNG files up to 10 MB, with source dimensions from 340 to 3850 pixels and aspect ratios from 1:2.5 to 2.5:1. For video, it accepts MP4 or MOV. Those limits are practical guardrails: a clean, full-body image and a driving clip with visible limbs give the model more usable information than a heavily cropped portrait or an action hidden by motion blur.
What each Kling V3 Motion Control output shows
Test 1: studio portrait with a confident move
This was the baseline test. A neutral background and mostly fixed camera leave very little visual noise, so the result shows whether the core transfer works before stylization is added. The output follows the raised arms and forward step in a convincing sequence. The face, clothing, and studio setting remain coherent through the move. The weak point appears where the hands travel fastest: arm edges soften briefly and show slight blending. That is still a usable result for a short social cut or a proof of motion, but it is a reason to avoid fast hand flourishes when a close inspection matters.
Test 2: neon club stage lighting
The second test asks the prompt to do more work. It keeps the same identity and dance intent while replacing the clean studio look with colored rim light, haze, and mild camera sway. The output shows that the lighting edit reads clearly on the body and background instead of sitting on top as a flat color wash. The trade-off is visible around hair and shoulders, where neon spill can become a little too broad during movement. The clip works when the goal is a quick mood change from a single subject image. It is less suited to a shot where hair edges or exact fabric texture must stay perfect frame by frame.
Test 3: fashion campaign wind and confetti
The third result stresses foreground interaction. It asks Kling V3 Motion Control to preserve a studio identity while adding subtle wind and a layer of floating confetti. Hair and clothing gain motion that makes the shot feel less static, and the subject remains recognizable. The confetti reveals the limit: a few pieces clip through the hair rather than passing cleanly in front of or behind it. That makes this a good choice for atmospheric campaign material, where individual particles are not the focal point. It is a poor fit for product footage or compositing work that needs reliable occlusion on every frame.
Run time and cost on Wiro
| Output | Elapsed task time | Cost shown for this output | Read on the result |
|---|---|---|---|
| Test 1 | 788 seconds | Not recorded | Best clean baseline for checking identity and body motion. |
| Test 2 | 676 seconds | Not recorded | Best for a fast relight and atmosphere change. |
| Test 3 | 673 seconds | Not recorded | Best for subtle secondary motion with a simple environment. |
These are task elapsed times from the three runs, not a promise of future queue time. They range from 673 to 788 seconds. The documentation exposes task-cost fields in its example response, but it does not provide a price for these three completed outputs. No cost is claimed here rather than borrowing the example value or inventing a rate. Queue load, source duration, and processing conditions can change the wall-clock result.
How to get stronger motion transfer
- Use a source photo that shows the body sections needed by the driving motion. A waist-up crop cannot give the model reliable information for a full-body step.
- Choose a driving clip with clear limb separation. Hands crossing the face or body are where artifacts are most likely.
- Keep the prompt focused on one visual change. Test 2 worked because lighting and atmosphere were the point; adding a new outfit, location, and camera move at once would make diagnosis harder.
- Use the video-orientation setting when the driving clip’s pose direction is more important than matching the still image’s pose.
When to pick Kling V3 Motion Control
Pick Kling V3 Motion Control when a specific reference person or character needs to perform an existing motion. It is a practical fit for turning a campaign portrait into a short motion asset, adapting a character image to a dance or gesture reference, or making several visual treatments from the same performance clip. Choose a text-to-video workflow instead when there is no driving performance to preserve. For a wider view of that category, see Top 5 Image-to-Video APIs in 2026 and Kling V3 Omni: 3 Sound-On Text-to-Video Tests.
The practical verdict is narrow but useful. This model transfers broad body action and holds identity well enough for short clips, even when the prompt changes the lighting or adds restrained atmosphere. It needs cleaner motion when hands move quickly, and it should not be trusted for precise particle occlusion. Start with a stable source image, a readable driving clip, and one focused visual instruction. For a product that needs a person to move like a reference, that is exactly the job this model is built to do.
Official source
For the maker’s current product information, see Kling AI.
Try it
Run Kling V3 Motion Control on Wiro with a clear source image and a driving clip that shows the motion you want to keep.