An AI video generator can nail the motion and still hand back a clip with nothing on the audio track, no footsteps, no engine hum, no room tone, just silence under moving pixels. That gap is common enough that a small set of models exists purely to close it, watching the footage, working out what should be making noise, and generating a synced track for it. Here is how that actually works on Picasso IA, and where it still needs a human pass.
These models take a video file as input and return the same clip with a generated audio track synced to it. Under the hood they read the frames, work out what is moving and when, and produce a waveform timed to those moments rather than a generic loop dropped underneath. That is the difference between this and stock sound libraries: the audio is built for your specific footage, not searched for and trimmed to fit it.
Picasso IA runs three models built for this in the video-editing part of the full catalog. Mirelo video-to-sfx-v1.5 is built for foley, footsteps, taps, impacts, the kind of sound that needs to land on a specific frame. MMAudio takes a short text prompt plus an optional negative prompt, which makes it good at ambient beds like traffic, wind or crowd murmur while excluding sounds you do not want. Thinksound accepts a caption plus a longer chain-of-thought description, so you can spell out several sound sources in one scene and have the model reason through how they layer. None of the three generates music by design, all three are built for sound effects and ambience.
The three models overlap in what they can attempt but differ in what they are actually good at, so picking the right one first saves a round of regenerating. This is the honest breakdown, not a marketing comparison.
| Model | Give It | Best At | Skip It For |
|---|
| video-to-sfx-v1.5 (mirelo) | Just the clip, optional style words | Foley hits like footsteps, taps, impacts | Long ambient beds |
| mmaudio | A short prompt plus a negative prompt | Steerable ambience, city noise, weather | Precise timing on fast cuts |
| thinksound | A caption and a chain-of-thought description | Layered scenes with several sound sources | Quick one-word requests |
| video-audio-merge | A finished audio file and the original clip | Mixing or replacing a track after generation | Generating any sound itself |
If the clip is one clear action, a door closing, a ball bouncing, start with video-to-sfx-v1.5. If it is a wide environmental shot, a street, a forest, a beach, start with mmaudio and lean on the negative prompt to keep music out. If the scene has several things happening at once and you can describe them in order, thinksound is worth the extra typing.
The video itself does most of the timing work, but the words you add still steer what the model hears in the frame. These are the habits that consistently produce a usable first take instead of a third regeneration.
- Name the source, not the mood: write footsteps on gravel rather than tense atmosphere, since the models respond to concrete sound sources far better than adjectives about feeling.
- Match duration to the clip: set the requested length to the actual clip length instead of the default, so the track does not cut off mid-action or trail on past the last frame.
- Run more than one variation: generate two or three takes per clip and pick the one that lands, since the same prompt can produce a tighter or a looser interpretation each run.
- Say what to leave out: use the negative prompt field on mmaudio to exclude music or voices when you only want environmental sound under the scene.
- Keep the reference short: one clear sentence beats a paragraph, and a chain-of-thought description on thinksound works best as an ordered list of sounds, not a mood board.
This is the full loop from a silent clip to an exported file with synced audio, using the models above and the merge tool that follows them. It assumes only a video file with no usable sound.
- Upload the silent clip. Open the Toolkit and pick a video file, an AI-generated clip or a phone recording works the same way here.
- Choose a sound model. Pick video-to-sfx-v1.5 for a specific action, mmaudio for ambience, or thinksound for a layered scene with several sound sources.
- Write a short, concrete prompt. Name the sounds you expect to hear and set the negative prompt if the model supports one, then set the duration to match the clip.
- Generate a few variations and compare. Play each take against the video with the sound on, not just as a waveform, since timing errors are easy to miss without watching the picture.
- Merge, replace or mix the track. Run the winning take through video-audio-merge to combine it with the original clip, replacing the audio entirely or blending it under any existing sound.
A generated sound effects track rarely sounds finished the moment it comes out of the model, it usually needs to sit at the right level against whatever else is in the clip. Video-audio-merge exists for exactly this step. It takes the original video and a separate audio file, lets you choose between replacing the original sound entirely or mixing the new track on top of it, and applies a volume multiplier before it renders, so a generated effect that reads too loud or too quiet gets fixed here rather than by regenerating the sound itself.
Keep a trace of the original ambience under the generated effects instead of muting it completely, real footage rarely has true silence, and a hint of room tone under the new sound is what stops the mix from sounding pasted in.
Once the levels sit right, pick an output format and codec, MP4 with H264 covers most destinations, and export. The whole pass, generate, compare, mix, export, usually takes a few minutes per clip once you know which model to reach for first.
An honest page says where the tool stops. For video-to-sound-effects generation there are real limits worth knowing before you rely on it for something that matters.
Fast cuts and quick multi-action sequences can outrun the timing these models produce, a hit that should land on frame twelve sometimes lands close but not exact, and you notice on a loop even if a first watch does not. None of the three models generates music, by design, so a track that needs a score still needs a separate music tool or a licensed cue. Mirelo video-to-sfx-v1.5 trims longer source videos to ten seconds per run, so a longer scene needs the start offset stepped across it in multiple passes rather than one single generation. Dialogue and lip sync are outside the scope entirely, these models add effects and ambience around speech, they do not generate or align a voice. And because the model is inferring sound from pixels, an unusual or ambiguous action can produce a plausible-sounding effect that is simply wrong for what actually happened, which is why comparing the audio against the picture before publishing matters more than trusting the first take.
Can AI really add sound effects that match what is happening in a video?
Yes, within limits. Video-to-sfx-v1.5, mmaudio and thinksound all read the frames of an uploaded clip and generate an audio track timed to the visible action rather than a generic loop. For a clear single action like a door closing or footsteps on gravel, the timing is usually close. For fast cuts or several things happening at once, review the result against the picture before you trust it, since timing can drift on complex scenes.
Which model should I use for footstep or impact sounds specifically?
Start with video-to-sfx-v1.5 from mirelo. It is built for foley-style sounds that need to land on a specific frame, footsteps, taps, a door latch, an impact, and it lets you generate several variations per run so you can pick the take with the tightest sync. Mmaudio and thinksound can attempt the same job but tend to work better on ambient beds than on a single precise hit.
Does this generate music too, or only sound effects?
Only sound effects and ambience, by design. None of the three video-to-audio models on Picasso IA generates music, and mmaudio actually defaults to a negative prompt that excludes music from its output so it does not accidentally add a score you did not ask for. If a clip needs a soundtrack, that is a separate step with a dedicated music model, not something these tools attempt.
How long can the video be?
It depends on the model. Video-to-sfx-v1.5 trims a source clip to ten seconds per generation, with a start offset you can move to cover a longer video in multiple passes. Mmaudio and thinksound let you set duration directly, typically up to the length of the clip you upload. For anything longer than a single shot, plan on generating in segments and merging the results rather than expecting one pass to cover a full scene.
Can I control the style of the sound with a text prompt?
Yes, with different depth per model. Mmaudio takes a short prompt plus an optional negative prompt, so you can ask for wind and distant traffic while excluding music or voices. Thinksound goes further with a caption and a chain-of-thought field where you describe several sound sources in order. Video-to-sfx-v1.5 accepts an optional prompt too, alongside a creativity setting that pushes the result toward tighter realism or looser, more stylized audio.
What if the generated sound does not match the video well?
Regenerate before you rewrite the prompt. All three models are stochastic, so the same input can produce a noticeably different take on the next run, and video-to-sfx-v1.5 explicitly supports generating several variations in one go for this reason. If two or three attempts all miss the same detail, that is when adjusting the prompt or switching models, mmaudio for ambience versus thinksound for a layered scene, usually fixes it.
Can I mix the generated sound with the video's original audio instead of replacing it?
Yes. Video-audio-merge on Picasso IA has a mix mode that layers the new track on top of the original sound rather than replacing it outright, along with a volume multiplier to balance the two. This matters when a clip already has usable ambient noise and you only want to add specific effects on top, rather than wiping out everything that was already there.
Does this work for video clips generated on Picasso IA itself, or only uploaded footage?
Both. A clip generated by a text-to-video or image-to-video model on Picasso IA downloads the same way any video does, and you can feed it straight into video-to-sfx-v1.5, mmaudio or thinksound the same as footage shot on a phone or a camera. This is actually the most common use, since AI-generated video clips frequently come out with no audio track at all.
How much does adding sound effects to a video cost?
Each generation costs credits, and the amount depends on the model and settings you choose, so a specific number here would go stale fast. New accounts start with free credits, enough to test video-to-sfx-v1.5, mmaudio and thinksound against the same clip and compare which one fits your footage before spending anything further. Current plans and credit allowances live on the pricing page, which is the only place worth trusting for numbers.
Is this a replacement for a sound library or a sound designer?
For a lot of short-form content, close enough to replace the search step: instead of scrubbing through a stock library for a footstep or a door creak that almost fits, you generate one timed to your actual footage. For a film or a project where sound design carries real weight, treat this as a fast first pass rather than the final mix, Picasso IA vs Artlist walks through where a stock library still wins on hand-curated, human-mixed audio.
A silent clip is one upload away from having sound. Open the Toolkit, pick video-to-sfx-v1.5, mmaudio or thinksound, and give your next AI video something to sound like.