The "hug my younger self" trend spread because it does something a filter never could: it takes two moments that never touched, a childhood photo and an adult one, or two people who never shared a room, and lets them hold each other for a few seconds. Picasso IA does not fake this with a canned template. It runs the same image-to-video models built for product shots and film previz, pointed at a photo instead, so the embrace moves the way the model understands motion, not a preset.
The appeal is not really about video quality, it is about the specific gap the clip closes. A childhood photo of yourself, a portrait of a parent who has passed, a picture of a friend who moved away, all of these are moments frozen at arm's length. Turning one into three seconds of motion where the two figures step together and embrace gives the still image a kind of closure a static photo cannot offer, and that emotional weight is why the format keeps resurfacing across every social platform instead of fading like a typical filter trend.
It also happens to sit exactly where current AI video models are strong. Image-to-video generation is good at plausible human motion over a short clip, and a hug is a simple, well-defined motion: two figures close a gap and their arms wrap around each other. That is a much easier ask than a dancing routine or a complex action sequence, which is part of why the results tend to look convincing on the first or second try rather than needing dozens of attempts.
There is more than one path to the same result, and which one fits depends on what you are starting with: one photo where both people already appear together, two separate photos that need to become one scene first, or a photo plus a specific motion you want copied exactly.
| Starting point | Model to use | What you upload | Best for |
|---|
| One photo, both people already in frame | wan-2.7-i2v or hailuo-02 | The single photo, plus a prompt describing the hug | The simplest case, fastest to try |
| Two separate photos, different times or places | qwen-image-edit-plus or nano-banana-2-lite, then a video model | Both photos, merged into one composite first | "Hug my younger self", reunions, pets and owners |
| A specific hug motion you want copied | kling-v2.6-motion-control | Your photo plus a short reference video of the embrace | Matching a particular gesture or timing |
| Quick test with no reference video | kling-v2.1 or kling-v2.6 | The photo and a plain text description | Fast iteration before committing to a longer clip |
Start from the top row if your inputs allow it. The two-step path in row two takes longer because it runs two separate generations, but it is the only route when the two people in the video never appeared in the same photograph to begin with.
Most "hug my younger self" videos are not built from one photo, they are built from two, merged into a single scene before anything moves. This is the step people skip when they are disappointed with a result: feeding two separate images straight into a video model usually produces something confused, because the model has no instruction for how the two photos relate to each other in space.
The fix is to merge first with an image editing model built for exactly this. qwen-image-edit-plus and nano-banana-2-lite both accept multiple reference images and a plain-language instruction, and both can place a person from one photo into the scene of another while keeping their face and proportions recognizable. A prompt like place the person from image one next to the person from image two, both standing, facing each other, warm indoor lighting produces a single composite photo where the two subjects finally share a frame.
That composite is what you then hand to the video model. Because the two people already share consistent lighting, scale and a believable pose, the video model has an easier job: closing the small remaining gap between them rather than inventing a shared scene from nothing. This order, merge first, animate second, is the single biggest factor in whether the clip looks staged or looks real.
This is the full path from two photos on your phone to a short video, covering both the one-photo and two-photo cases so you can skip the step that does not apply to you.
- Pick your photos. One clear, well-lit photo per person, faces visible and unobstructed, is enough for either the single-photo or two-photo path.
- Merge if you are starting from two photos. Open a multi-image editor like qwen-image-edit-plus or nano-banana-2-lite in the Toolkit, upload both photos, and describe how the two people should be positioned together.
- Pick a video model and describe the motion. For a straightforward embrace, wan-2.7-i2v or hailuo-02 work well; for a specific gesture, kling-v2.6-motion-control can copy it from a short reference clip.
- Generate at a short duration first. Run a five- or six-second clip before spending credits on a longer one, and check that the arms, faces and proportions hold up across the motion.
- Regenerate with a tighter prompt if needed. Add detail about pacing, camera framing or where the hands should land, and run it again; most people land on a version they like within two or three attempts.
The difference between a hug video that reads as touching and one that reads as unsettling almost always comes down to timing and framing, not the underlying model. A prompt that just says two people hug tends to produce something abrupt: arms snap into place, the embrace starts and ends within the same second, and the motion feels mechanical rather than warm.
Describing the motion in stages fixes most of this. Break the hug into its beats: they walk toward each other, pause, open their arms, and embrace slowly, holding for a moment. Naming the pause matters more than any other word, because it is the part a rushed prompt skips first and the part that reads as emotion rather than choreography. Keeping the camera mostly still, rather than a sweeping cinematic move, also helps, since a static shot lets the viewer focus on the embrace.
If the generated hands or arms look wrong where the two people connect, do not keep regenerating the whole clip blindly. Try a shorter duration first, since contact points are usually where a video model is least confident, and a five-second clip gives it less time to drift from a mistake it made early on.
Lighting consistency between the two source photos matters more than people expect. A merged photo where one person is lit warmly indoors and the other by cold daylight tends to carry that mismatch into the video, since the model treats it as part of the scene rather than something to correct.
This format is genuinely moving when it works, which is exactly why it deserves an honest account of when it does not.
- Hands at the point of contact: the moment where fingers press into a shoulder or back is the hardest part of any hug to render, and it is where extra fingers or a blurred grip show up most often.
- Photos from very different eras: a faded 1980s print merged with a crisp modern phone photo can produce a visible mismatch in grain, color and sharpness that the merge step cannot fully hide.
- Faces at odd angles: a hug naturally turns both faces partly away from the camera, and profile or three-quarter views are harder for these models to hold steady than a straight-on portrait.
- Long clips: pushing past six to ten seconds increases the chance of drift, where a face or proportion slowly shifts partway through the embrace.
- People who never resemble each other in the source photos: if the two images differ wildly in lighting, resolution or angle, the merge step has to guess at more than it can reliably fill in, and the seam sometimes shows.
None of this means the format does not work. It means budgeting for two or three attempts rather than expecting the first generation to be the keeper, which is true of most AI video generation.
Can I make a hug video from just one photo?
Yes, if both people already appear together in that one photo. Upload it to an image-to-video model such as wan-2.7-i2v or hailuo-02 and describe the embrace you want; the model animates the existing scene rather than needing to combine anything. The two-photo merge step only matters when your subjects come from separate pictures that were never taken together.
How do I combine two separate photos of different people into one hug scene?
Use a multi-image editing model, qwen-image-edit-plus or nano-banana-2-lite, both available in the Toolkit. Upload both photos and write a plain instruction describing how the two people should be placed together, standing close, facing each other, in consistent lighting. The result is one composite photo, which you then animate with a video model as a separate step.
Which model gives the most realistic hug motion?
There is no single best option, since the right pick depends on what you are starting from. wan-2.7-i2v and hailuo-02 both produce natural physics-aware motion from a single photo and a text prompt. kling-v2.6-motion-control is the strongest choice when you have a reference video of the exact embrace you want copied, because it transfers that specific motion rather than generating one from a description.
Why do the hands or arms look wrong in my generated video?
The point where hands press into a shoulder or back is the hardest part of a hug for any current video model to render cleanly, since fingers involve fine detail packed into a small, fast-moving area of the frame. Shorter clips tend to hold up better than long ones, and a tighter prompt describing exactly where the hands should land often helps more than simply regenerating the same request.
Can I animate a photo of myself hugging a deceased relative?
Yes, and this is one of the most common uses of the format. Upload a clear photo of yourself and a clear photo of the person you want to include, merge them into one scene with an editing model, and animate the result. Treat the output as a gentle visual tribute, since the model is generating a plausible embrace, not reconstructing a real moment.
How long can an AI hug video be?
Current image-to-video models on Picasso IA generate short clips, typically five to ten seconds depending on the model. That length suits the format well, since a hug is a brief motion to begin with, and pushing past ten seconds increases the chance that a face or proportion drifts partway through the clip rather than adding anything the trend needs.
Do the two people in the video need to be facing the camera?
Not necessarily, but a mostly forward-facing or three-quarter angle in the source photos gives the model the clearest information to work with. A hug naturally turns both faces partly away from the camera as the embrace happens, which is already a harder angle for these models than a straight-on portrait, so starting from a clear, well-lit frontal photo gives the generation the best foundation.
Can I use a photo of a pet instead of a person?
Yes, the same two-step approach works for a person hugging a pet, or even two pets, as long as the source photos are clear. Merge the two subjects into one scene first if they come from separate photos, then animate the embrace with the same video models used for people. Fur and paws behave differently from hands, and tend to render a little more forgivingly at the point of contact.
What happens if the two source photos have very different lighting?
The mismatch usually carries through to the final video, since the merge step treats both photos as part of one intended scene rather than correcting for how differently they were lit. Where possible, mention the lighting you want in the merge prompt, warm indoor light or soft daylight, so the editing model evens out the difference before the video model ever sees the composite.
How much does making a hug video cost?
Both the merge step and the video generation step cost credits, and the total depends on which models and settings you choose, since a longer or higher-resolution video costs more than a short, lower-resolution one. New accounts start with free credits, which is enough to try the two-step workflow before spending anything further, and current plan pricing lives on the pricing page, which is the only place worth trusting for exact numbers.
Two photos and a few minutes are all this takes to try. Open the Toolkit, merge your images if you need to, and see what the embrace looks like in motion.