A short film is not one long generation, and treating it like one is the fastest way to get eight seconds of nothing usable. Every current text-to-video model generates a single shot of a few seconds at a time. A film or a trailer is what happens after: several of those shots, chosen, trimmed and merged in an order that reads as one story. This page walks through that assembly line end to end, shot list to exported cut, with the real models involved at each stage.
The instinct is to ask for the whole scene in one prompt and hope the model holds it together for thirty seconds. It will not, not yet. What text-to-video models are genuinely good at is a single beat: a character turning to camera, a car pulling into frame, a door opening on a lit room. Once you accept the shot as the unit of work instead of the film, the workflow stops fighting the tool.
This is also why a catalog approach helps more here than with a single style of image. A trailer usually needs a wide shot, a character beat, something fast, and a final punch, and no single video model wins at all four. Picasso IA runs 488 models behind one interface, including several text-to-video engines and a separate video-editing category for everything that happens after generation.
The single habit that separates a coherent thirty-second cut from a folder of unrelated clips is writing the shot list first, on paper or in a notes app, before opening the Toolkit. List every beat the film needs in order: the establishing shot, the character entrance, the turn, the reveal, the final frame. Two or three sentences per shot, written like a photo description, is enough detail for a video prompt.
Keep each shot's duration honest to what the model actually returns. Most text-to-video generations land between four and ten seconds, so a thirty-second trailer is realistically six to eight separate generations, not one. Six short generations that each nail their beat cut together far better than one long generation that drifts by the second half.
Write the dialogue or narration line, if there is one, into the shot list too, even if you plan to record it separately. A shot generated to match a specific spoken line sits together with its audio far more naturally than a shot generated first and narrated over afterward as an afterthought.
Not every beat in a trailer wants the same engine, and this is where a shot list pays off again: it tells you, before you generate anything, which kind of shot you are about to ask for. The table below is a starting point, not a rule, and every model listed sits in the text-to-video category of the catalog alongside dozens of others worth testing on a specific shot.
| Shot Type | A Model to Try | Why It Fits |
|---|
| Wide establishing shot | google/veo-3.1 | Strong scene coherence and native ambient audio for atmosphere |
| Character close-up or turn | kwaivgi/kling-v2.6 | Handles a single subject's motion smoothly across the shot |
| Fast action or a chase beat | bytedance/seedance-1-pro | Keeps quick camera or subject motion from breaking apart |
| Product or object reveal | lightricks/ltx-2-pro | Clean, high-resolution motion for a trailer's final punch |
| Narrated montage frame | kwaivgi/kling-v2.6 or google/veo-3.1 | Simple, readable motion that a voiceover can sit on top of |
Generate each shot two or three times before moving on. Video generation is less predictable than a still image, and the difference between a usable take and a stiff one is often the third attempt. Picking the strongest clip per beat costs minutes and saves a re-edit later.
Generation is the visible half of the job; assembly is where a shot list actually becomes a film. This is the order that turns six or eight separate clips into one exported cut, using the editing models in the video-editing category alongside the generators above.
- Generate every shot on your list. Run each beat through the text-to-video model that fits it, keep the two or three best takes per shot, and discard the rest before moving on.
- Trim each clip to its usable range. Cut out the first and last frames where a generation warms up or drifts, using a trimming tool such as trim-video, so only the clean motion survives into the edit.
- Merge the trimmed clips into one file. Feed the ordered clips into video-merge, which concatenates them in sequence and can match frame rate and resolution across shots that came from different models.
- Add music, sound effects or narration. Layer a soundtrack or narration onto the merged file with video-audio-merge, or generate ambient sound directly from the picture with a model such as mmaudio when a shot needs texture the source clip did not carry.
- Caption and export the final cut. Run the finished sequence through autocaption if the trailer needs burned-in text, then export at the resolution your platform expects, reframing with reframe-video first if the destination needs a different aspect ratio.
None of these five steps requires software outside the browser. Each one is its own model with its own settings, which means a bad merge or a mistimed cut is a five-minute re-run, not a lost afternoon in a timeline editor.
The clearest tell that a cut was assembled from separate generations is a character or location that subtly changes between shots: a different jacket, a different face, a grade that shifts halfway through. Fixing this is a habit, not a model choice: generate one strong reference image for your protagonist and your key location, then feed that same reference into every shot that needs it instead of describing the character from scratch each time.
The technique is the same one covered on the consistent AI characters page: lock a reference image, vary only the pose, the angle and the action per shot. For a trailer this matters even more than for a single portrait, because the eye catches inconsistency fastest across a cut, in the half second between two shots of the same person.
Pro tip: generate your protagonist's reference image first, alone, in neutral lighting, before writing a single shot prompt. Every later shot references that one image instead of a fresh text description, and that single decision does more for visual continuity than any prompt wording ever will.
Style words help too, the same way they do for a still image: pick two or three consistent adjectives, such as muted teal grade or handheld documentary look, and repeat them in every shot's prompt. The effects page collects preset looks that are a faster way to find that vocabulary than guessing at it shot by shot.
An honest workflow page says where it stops working, and for AI-assembled film this is not a short list. Budget for these five limits before you start, and none of them will surprise you mid-project.
- Shot length: current text-to-video models return a handful of seconds per generation, so anything longer than a trailer means dozens of separate shots, and the editing time scales with that count, not with the runtime.
- Lip-synced dialogue: models can produce a character talking, but tight, script-accurate lip sync across a full line of dialogue is still unreliable, and a trailer with a voiceover over the picture avoids the problem entirely.
- Cross-shot continuity: even with a locked reference image, small details, a prop's exact position, the fold of a jacket, drift between separate generations, in a way a single continuous camera take never would.
- Sound design cohesion: music, ambient sound and dialogue arrive from different tools and need a manual mixing pass; nothing on the platform auto-balances levels across a full cut for you.
- Full-length narrative: this workflow is built for a trailer or a short, a few minutes of assembled shots at most, not a feature-length film with continuous scenes and a full cast held consistent across hours.
None of this argues against the workflow. It argues for building trailers, teasers and short-form pieces, the formats where a handful of strong shots genuinely outweighs one continuous take.
Can AI actually generate a full short film in one go?
No, and any tool that claims otherwise is describing a demo, not a deliverable. Every current text-to-video model generates a shot of a few seconds at a time. A short film or trailer is always an assembly of several of those shots, chosen and merged afterward, the same way a real film is shot in setups and edited together rather than filmed in one take.
Which Picasso IA models should I use for the shots themselves?
There is no single best model, which is the argument for testing a few on your specific shot. google/veo-3.1 and kwaivgi/kling-v2.6 handle character beats with strong coherence, bytedance/seedance-1-pro holds together fast motion, and lightricks/ltx-2-pro produces high-resolution shots for a final reveal. The text-to-video category has dozens more worth testing.
How do I combine the separate clips into one video?
Once every shot is generated and trimmed, video-merge in the video-editing category concatenates the clips into a single file in the order you upload them, and can normalize frame rate and resolution across shots that came from different source models. It works the same whether you join two clips or twenty.
Can I add music or a voiceover to the finished cut?
Yes. Once the shots are merged into one file, video-audio-merge lays a music track, a recorded voiceover or a sound effect on top of the picture. If a shot needs ambient texture the original generation lacks, such as footsteps or wind, a model like mmaudio can generate audio directly from the video before that clip goes into the merge.
How do I keep the same character looking right across every shot?
Generate one strong reference image of the character first, in neutral lighting, before writing any shot prompts, then feed that same image into every generation that needs the character rather than describing them fresh each time. This is the same approach used for consistent AI characters, and it is the single change that does the most for visual continuity across a multi-shot cut.
Does dialogue actually sync to the character's mouth?
Sometimes, not reliably enough to build a trailer around it. Current video models can produce a character speaking, but frame-accurate lip sync to a specific scripted line is inconsistent across a full sentence. The more dependable pattern is to write the trailer around a voiceover or narration that sits over the picture, saving generated lip sync for short, forgiving lines rather than a full dialogue scene.
How long should each generated shot actually be?
Plan around four to ten seconds per generation, since that is what most current text-to-video models return in a single run. A thirty-second trailer built from six five-second shots edits together cleanly; one long generation stretched toward thirty seconds tends to drift well before it gets there.
What if the aspect ratio needs to change for a different platform?
Reframe the merged cut with reframe-video rather than regenerating every shot at a new ratio. It adjusts the frame to fit a new aspect ratio, vertical for a phone screen or square for a feed post, without touching the underlying generations, which is faster and keeps every shot's content intact.
Can I add captions or on-screen text to the trailer?
Yes, once the cut is fully merged. autocaption reads the audio track, whether that is dialogue, narration or a voiceover you added during the merge, and burns in styled subtitles with control over font, color and position. Running it as the last step, after picture and sound are locked, avoids having to redo the captions every time an earlier shot changes.
How much does a full short film or trailer cost to generate?
Each shot generation and each editing pass costs credits, and the total depends on how many shots you generate, how many takes you keep, and which models you choose, so a number here would be wrong within a month. New accounts start with free credits, enough to test the workflow before committing further, and current plans live on the pricing page.
Your shot list is the only thing standing between an idea and a cut you can actually watch. Open the Toolkit and generate the first shot, or see how the Picasso IA vs Runway comparison frames the same workflow against a single-model editor.