Skip to content
Before / After

MMAudio: 4 Video-to-Audio Before/After Tests

MMAudio video-to-audio tests ask a simple question: can a generated soundtrack make a short silent clip feel like one event rather than a stack of guessed noises? The four examples below use the same input clips already published here and pair each one with a short sound description. The useful test is not whether every sound is perfect. It is whether the main sound, ambience, and timing agree with what the viewer sees.

MMAudio on Wiro accepts a video URL or uploaded video plus an optional prompt. For video input, the service matches output duration to the input clip. The public model documentation lists 30 denoising steps, strength 4.5, seed 123456, and a negative prompt of bad, blurry as defaults. The documentation does not publish a fixed Wiro price or expected turnaround per output, so no price or timing has been inferred for these four clips. The research paper reports 1.23 seconds for an 8-second clip in its own setup; that is not a Wiro runtime guarantee.

What the MMAudio video-to-audio tests check

Each test gives MMAudio one primary event and a few supporting sounds. That makes the result easier to judge. A beach clip needs surf first, then wind and birds. A horse clip needs hoof rhythm before breath or distant ambience. A storm needs rain and thunder, but the timing matters: thunder should not feel like a loop pasted onto the picture. The city prompt is deliberately crowded, so it also checks whether the model can keep a scene readable when several sound sources compete.

The underlying project describes MMAudio as a video-and-text conditioned audio generator with a synchronization module. Its authors say the model was trained with audio-visual and audio-text data so that it can use the video for timing and the prompt for semantic direction. The official MMAudio repository also flags real limits: speech-like artifacts can occur, background music is not a strong use case, and unfamiliar concepts may fail. Those caveats matter more than a clean demo reel.

Four before-and-after outputs

Test 1: beach surf, wind, and seabirds

Prompt: ocean waves crashing on the beach, seagulls, wind

This is the cleanest test because the visual scene points to an obvious sound hierarchy. Listen first for the broad wash of surf, then for wind and occasional bird detail. The generated result should feel anchored by wave movement rather than turn into a generic ocean ambience track. It is a good fit for MMAudio because the visual actions and prompt agree.

Before: the beach input video.
After: MMAudio output for the beach prompt.

Test 2: hoof impacts and horse movement

Prompt: horse galloping, hooves on dirt, snorting, breath

The key signal here is cadence. Hooves should suggest repeated ground contact instead of a single static animal sound. Breath and snorts are secondary details, so they should not overwhelm the movement. This is a tougher synchronization check than the beach: viewers can spot a misplaced impact quickly when the horse is moving on screen.

Before: the galloping-horse input video.
After: MMAudio output for the horse prompt.

Test 3: rain, wind, and thunder

Prompt: lightning storm, thunder, heavy rain, strong wind

This output tests layers rather than a single action. Rain and wind can supply a continuous bed, while thunder needs enough separation to read as an event. The clip is useful for checking restraint: a soundtrack that puts every element at the same level sounds busy even if each ingredient matches the prompt. For synthetic weather footage, that balance often matters more than extra detail.

Before: the storm input video.
After: MMAudio output for the storm prompt.

Test 4: a dense rainy street

Prompt: busy city street, cars passing, distant siren, light rain

The fourth case is the stress test. Cars, rain, and a distant siren all make sense, but the prompt does not tell the model exactly when each should arrive. This makes it a better test of scene-level plausibility than frame-perfect Foley. Listen for whether traffic remains the base layer and whether the siren stays distant. If the prompt feels too busy, remove one secondary sound before increasing strength.

Before: the input video retained from the original fourth test.
After: MMAudio output for the rainy-city prompt.

Parameters used and what to change

The documented defaults are 30 steps, strength 4.5, seed 123456, and the negative prompt bad, blurry. In a video-to-audio run, duration follows the video; the 8-second duration field applies when generating audio from text without a video. More steps trade speed for quality. Strength controls prompt influence: the docs recommend 4.0 to 5.0, so 4.5 is a sensible starting point. Keep the seed fixed when comparing prompt edits, then change it only when seeking a different variation.

For a practical pass, name the sound that must land with the picture first. Add one or two background elements, not a shopping list. Use the negative prompt for sounds that would break the scene, such as music or crowd noise. The open-source project notes that higher-resolution input can add encoding and decoding time without improving quality, so a shorter, clean source clip is usually the better test asset.

When to pick MMAudio

Pick MMAudio when a silent clip needs Foley-style ambience or event sound and the video already tells the model what is happening. It suits short product shots, generated B-roll, nature clips, simple action, and prototypes where a sound designer will later polish the mix. It is less suitable for intelligible dialogue, a finished music bed, or highly specific unfamiliar effects. Those limits are consistent with the authors’ published failure notes and with the model’s framing as synchronized audio generation rather than a full audio post-production suite.

For adjacent workflow ideas, see Kling V3 Omni sound-on video tests, Video Background Music v2 tests, and cinematic image-to-video model tests. For the research behind MMAudio, read the paper page on Hugging Face and the official GitHub repository.

Run MMAudio on Wiro with one clear sound priority, then compare variations with the same seed before committing to a final mix.