{"id":1298,"date":"2026-02-26T20:18:02","date_gmt":"2026-02-26T20:18:02","guid":{"rendered":"https:\/\/wiro.ai\/blog\/?p=1298"},"modified":"2026-09-27T20:10:35","modified_gmt":"2026-09-27T20:10:35","slug":"p-video-text-to-video-and-image-to-video-in-6-tests","status":"publish","type":"post","link":"https:\/\/wiro.ai\/blog\/p-video-text-to-video-and-image-to-video-in-6-tests\/","title":{"rendered":"P-Video: Text-to-Video and Image-to-Video in 6 Tests"},"content":{"rendered":"<h2>P-Video text-to-video and image-to-video in 6 tests<\/h2>\n<p><strong>P-Video text-to-video<\/strong> was tested here as a short-form production tool, not as a demo reel generator. The question was simple: can one model make usable motion for a product shot, rainy city b-roll, vertical food content, two image-led action scenes, and a dense fantasy shot without asking the prompt to solve everything at once?<\/p>\n<p>All six clips remain embedded below. They are the actual outputs from the original test set. Three began from text alone. Tests 04 and 05 used a supplied image, so the opening composition came from that image rather than the ratio selector. The published test record also says one run used seven seconds; it does not identify which clip, so this article does not assign that setting to a specific output.<\/p>\n<h2>Contents<\/h2>\n<ul>\n<li><a href=\"#setup\">Setup and controls<\/a><\/li>\n<li><a href=\"#results\">What each P-Video output shows<\/a><\/li>\n<li><a href=\"#speed\">Run time and cost record<\/a><\/li>\n<li><a href=\"#pick\">When to choose P-Video<\/a><\/li>\n<\/ul>\n<h2 id=\"setup\">P-Video test setup<\/h2>\n<p>P-Video accepts a prompt and can optionally take an image or an audio file. Its current Wiro controls expose 1-10 second duration, 720p or 1080p resolution, 24 or 48 fps, a seed, audio saving, draft mode, and prompt upsampling. Text-only work can select 1:1, 3:2, 2:3, 4:3, 3:4, 16:9, or 9:16. Once an input image is attached, the image determines the frame shape instead.<\/p>\n<p>The six published runs used 720p and 24 fps. The three text-to-video clips use 16:9, 16:9, and 9:16. The two image-to-video runs inherit their source-image composition. Prompt upsampling was disabled for one run, although the old task record does not say which one. That missing detail matters: it would be dishonest to credit a particular result to the setting.<\/p>\n<p>The prompts deliberately ask for a single dominant camera move. That is the useful constraint in this set. A turntable, tracking move, top-down food shot, or slow push-in gives the model a clear job. Combining several unrelated transitions would make it much harder to judge whether a failure came from the scene, the subject, or the camera instruction.<\/p>\n<p>For background on the provider, see the <a href=\"https:\/\/github.com\/PrunaAI\/pruna\" target=\"_blank\" rel=\"noopener\">Pruna open-source repository<\/a>. P-Video is available on <a href=\"https:\/\/wiro.ai\/models\/pruna\/p-video\">Wiro&#8217;s P-Video model page<\/a>.<\/p>\n<h2 id=\"speed\">Run time and cost record<\/h2>\n<table>\n<thead>\n<tr>\n<th>Test<\/th>\n<th>Mode<\/th>\n<th>Frame shape<\/th>\n<th>Recorded elapsed time<\/th>\n<th>Cost shown for this output<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>01<\/td>\n<td>Text-to-video<\/td>\n<td>16:9<\/td>\n<td>18 seconds<\/td>\n<td>Not recorded<\/td>\n<\/tr>\n<tr>\n<td>02<\/td>\n<td>Text-to-video<\/td>\n<td>16:9<\/td>\n<td>16 seconds<\/td>\n<td>Not recorded<\/td>\n<\/tr>\n<tr>\n<td>03<\/td>\n<td>Text-to-video<\/td>\n<td>9:16<\/td>\n<td>16 seconds<\/td>\n<td>Not recorded<\/td>\n<\/tr>\n<tr>\n<td>04<\/td>\n<td>Image-to-video<\/td>\n<td>Input image<\/td>\n<td>25 seconds<\/td>\n<td>Not recorded<\/td>\n<\/tr>\n<tr>\n<td>05<\/td>\n<td>Image-to-video<\/td>\n<td>Input image<\/td>\n<td>18 seconds<\/td>\n<td>Not recorded<\/td>\n<\/tr>\n<tr>\n<td>06<\/td>\n<td>Text-to-video<\/td>\n<td>16:9<\/td>\n<td>21 seconds<\/td>\n<td>Not recorded<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>The observed range was 16-25 seconds per output. Image-to-video was not automatically slower: one image-led clip took 25 seconds and the other took 18. The original six task records do not expose a cost value, so no per-clip price is claimed here. Wiro&#8217;s P-Video documentation includes a separate cancelled example task with a displayed total cost of $0.003510 for six elapsed seconds, but that is an example task, not a price quote for these six runs. Duration, resolution, queue state, and product pricing can change, so check the model page before budgeting a batch.<\/p>\n<h2 id=\"results\">What the six outputs actually show<\/h2>\n<h3>Test 01: product watch turntable<\/h3>\n<p><strong>Prompt:<\/strong> Ultra realistic product ad. A stainless steel wristwatch on matte black marble. Slow rotating turntable. Macro lens. Crisp reflections. Subtle dust motes. Camera does a gentle push in. High end studio lighting.<\/p>\n<figure><video controls preload=\"metadata\" src=\"https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/02\/p-video-test-01-sound.mp4\"><\/video><figcaption>Test 01: the original P-Video product watch output.<\/figcaption><\/figure>\n<p>This is the cleanest production-shaped request in the set. The clip shows why a simple product scene works: dark marble, one object, studio light, rotation, then a restrained push-in. Reflections sell the material and the motion has a clear direction. The weak point appears when inspecting fine geometry. Watch hands and bezel edges can shift slightly between frames, so this is a good fit for a mood-led ad cut or a fast social insert, not proof of exact product accuracy.<\/p>\n<h3>Test 02: rainy neon street tracking shot<\/h3>\n<p><strong>Prompt:<\/strong> Cinematic neon street at night in heavy rain. A dark sedan drives past. Wet asphalt reflections. Neon signs blur in the background. Smooth tracking shot from left to right. Realistic raindrops on the lens.<\/p>\n<figure><video controls preload=\"metadata\" src=\"https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/02\/p-video-test-02-sound.mp4\"><\/video><figcaption>Test 02: original rainy neon street output.<\/figcaption><\/figure>\n<p>The output succeeds first as atmosphere. The wet road, blur, and light reflections give the shot depth while the sedan supplies an easy motion anchor. Heavy rain adds useful texture but also raises the failure risk. Sign shapes can change and the car body can wobble when rain, lens droplets, and vehicle motion compete. Pick this kind of P-Video text-to-video prompt for establishing shots where the viewer reads light and movement before tiny structural details.<\/p>\n<h3>Test 03: vertical ramen motion<\/h3>\n<p><strong>Prompt:<\/strong> Top down food video. A steaming ramen bowl. Chopsticks lift noodles slowly. Steam swirls. Warm soft lighting. Shallow depth of field. Natural motion. Realistic textures.<\/p>\n<figure><video controls preload=\"metadata\" src=\"https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/02\/p-video-test-03-sound.mp4\"><\/video><figcaption>Test 03: original vertical ramen output.<\/figcaption><\/figure>\n<p>The 9:16 framing makes this the clearest short-form test. Steam and food texture read well at this length, and the overhead composition avoids a busy background. The hard part is the chopsticks. Hands, utensils, noodles, and liquid all need to stay coherent while moving together. Here, the controlled lift keeps the clip readable, but a brand that needs a precise hand action should still review every frame. Use a short duration and a single food action rather than asking for cutting, pouring, plating, and a camera move in one prompt.<\/p>\n<h3>Test 04: image-to-video snowboard tracking<\/h3>\n<p><strong>Prompt:<\/strong> Action sports. A snowboard instructor carves down a bright slope. Smooth dynamic tracking shot behind the rider. Crisp daylight. Snow spray in slow motion.<\/p>\n<figure><video controls preload=\"metadata\" src=\"https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/02\/p-video-test-04-sound.mp4\"><\/video><figcaption>Test 04: original image-to-video snowboard output.<\/figcaption><\/figure>\n<p>Supplying an image changes the job. Instead of inventing the opening subject placement, P-Video starts with a chosen rider, slope, and lighting direction. That gives this clip a stronger opening composition than a text-only action prompt usually gets. It does not remove motion problems. Boards, boots, and spray move quickly, so local warping remains possible. Choose image-to-video when the first frame, wardrobe, product, or brand art must look a certain way, then ask for one believable movement.<\/p>\n<h3>Test 05: image-to-video city pan<\/h3>\n<p><strong>Prompt:<\/strong> Cinematic urban street at dusk. A sleek dark sedan moves across frame. Background traffic approaches camera. Warm street lamps. Slow steady pan tracking the sedan. Photorealistic.<\/p>\n<figure><video controls preload=\"metadata\" src=\"https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/02\/p-video-test-05-sound.mp4\"><\/video><figcaption>Test 05: original image-to-video sedan pan.<\/figcaption><\/figure>\n<p>This output tests a safer version of vehicle motion. The pan follows one sedan, while background traffic adds depth without demanding a scene change. The image input anchors the street layout; the slow pan gives the model time to preserve it. The review points are practical: wheel rotation, headlight flicker, and lane geometry. For lifestyle b-roll, that trade-off works well. For a dealership spot where the exact vehicle must stay mechanically correct, use real footage or treat this as a concept cut.<\/p>\n<h3>Test 06: fantasy wildlife stress test<\/h3>\n<p><strong>Prompt:<\/strong> High fantasy documentary. A majestic stag with mossy antlers in a foggy redwood forest. Bioluminescent spores floating. Slow camera push in from wide to medium close up. Golden hour god rays. Shallow depth of field.<\/p>\n<figure><video controls preload=\"metadata\" src=\"https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/02\/p-video-test-06-sound.mp4\"><\/video><figcaption>Test 06: original fantasy wildlife output.<\/figcaption><\/figure>\n<p>This is the deliberate stress test. It asks for a subject with fine antlers, fog, particles, sun rays, a depth change, and a camera push-in. The output can still make an appealing fantasy beat, but it exposes the cost of stacking detail: antlers and drifting spores need temporal consistency from frame to frame. This is where P-Video works best as a creative concept tool. Keep the strongest seconds, cut around unstable frames, and do not promise continuity that the output does not show.<\/p>\n<h2 id=\"pick\">When to pick P-Video<\/h2>\n<p>Choose P-Video when the brief needs a fast, short clip with one central action. It fits product atmosphere, city b-roll, food loops, fantasy inserts, and image-led motion where the first frame matters. Start at 720p and 24 fps when validating an idea. Use prompt upsampling when a rough request needs help expanding; turn it off when exact prompt wording matters more than extra descriptive interpretation. Use draft mode for lower-quality previews before committing to a final render.<\/p>\n<p>For text-to-video, say what is in frame, name one camera move, and keep the event sequence short. For image-to-video, spend time on the source image because it controls the opening composition and frame shape. Avoid close-up hands, branded text, exact mechanical motion, or several independent moving subjects when the final use needs frame-level reliability.<\/p>\n<p>For more tests in this category, see <a href=\"https:\/\/wiro.ai\/blog\/lance-text-to-video-review\/\">Lance Text to Video Review<\/a>, <a href=\"https:\/\/wiro.ai\/blog\/cinematic-image-to-video-models\/\">Best Cinematic Image-to-Video Models<\/a>, and <a href=\"https:\/\/wiro.ai\/blog\/wan-2-7-video-5-video-prompt-tests\/\">WAN 2.7 Video: 5 Video Prompt Tests<\/a>.<\/p>\n<h2>Run P-Video<\/h2>\n<p><a href=\"https:\/\/wiro.ai\/models\/pruna\/p-video\">Run P-Video on Wiro<\/a> with a focused prompt first, then use the result to decide whether the shot needs an image anchor, a simpler action, or a shorter cut.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>P-Video text-to-video and image-to-video in 6 tests P-Video text-to-video was tested here as a short-form production tool, not as a demo reel&hellip;<\/p>\n","protected":false},"author":4,"featured_media":1299,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[52],"tags":[58,114,93,57],"class_list":["post-1298","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-model-reviews","tag-image-to-video","tag-p-video","tag-pruna","tag-text-to-video"],"_links":{"self":[{"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/posts\/1298","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/users\/4"}],"replies":[{"embeddable":true,"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/comments?post=1298"}],"version-history":[{"count":4,"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/posts\/1298\/revisions"}],"predecessor-version":[{"id":4252,"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/posts\/1298\/revisions\/4252"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/media\/1299"}],"wp:attachment":[{"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/media?parent=1298"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/categories?post=1298"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/tags?post=1298"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}