Everyone has a photo they wish could move. A portrait, a mascot drawing, a product character standing still on a white background. Motion transfer models close that gap without a single frame of manual animation: give them a still image and a reference video of someone dancing, and they generate a clip where your subject performs the same moves, at the same tempo. This page covers how that works on Picasso IA, how to pick inputs that produce clean motion, and where the technique falls apart.
Dance is one of the hardest things to describe in a text prompt and one of the easiest things to hand over as a reference. A sentence like energetic hip hop with fast footwork tells a text-to-video model almost nothing about timing, weight or rhythm. A ten-second clip of an actual dancer tells a motion transfer model everything it needs, because it reads real joint positions frame by frame instead of guessing from words.
That is the shift these models make. Instead of generating movement from a description, they copy movement from footage and apply it to a different subject. The result reads as choreographed because it is: the choreography already happened, on camera, before your photo entered the pipeline. All the model has to solve is mapping that motion onto a new character convincingly, which is a narrower problem than inventing dance from scratch.
Picasso IA runs two models built specifically for this: kling-v2.6-motion-control from kwaivgi and wan-2.2-animate-animation from wan-video, both living in the text-to-video side of the catalog alongside dozens of other video models. You give either one a reference image and a reference video. The model reads the actions, gestures and posture changes in the video and reproduces them on the subject in your image, frame by frame, without you touching a timeline or a rig.
The two behave a little differently. Kling v2.6 Motion Control runs in a Standard mode for fast, cheaper drafts and a Professional mode for sharper output, and it can keep the original audio from your reference clip. It also lets you choose whether the final orientation follows your photo or your reference video, which sets how long a clip you can generate: up to ten seconds one way, up to thirty the other. Wan 2.2 Animate Animation outputs at 720p or 480p, runs at up to 24 frames per second, and includes an experimental audio-merge option. Running the same pair of inputs through both is a reasonable way to see which reads your footage better.
The reference video does almost all the work, so the quality of your result tracks the quality of that clip more than anything else in the process. A dancer filmed head to toe, against a plain background, moving at a normal human pace, gives the model a clean signal to copy. A blurry phone clip of a crowd, or a routine with rapid spins and camera cuts, gives it far less to work with.
| Reference Video Type | What Transfers Well | Good Starting Point For |
|---|
| Solo choreography, full body in frame | Footwork, arm sweeps, turns | Dance-style social clips and reels |
| Simple gesture or upper-body clip | Head turns, hand waves, subtle sway | Animating a portrait or mascot image |
| Short loop, three to ten seconds | One clean movement repeated | Fast tests, sticker-style loops |
| Full routine up to thirty seconds | An entire sequence start to finish | Finished, shareable dance edits |
| Slow, deliberate movement | Precise joint tracking, few glitches | A first attempt on a new photo |
Framing matters as much as the footage itself. Both models want the character in your reference image to stand in roughly the same orientation as the person in the video, so a portrait cropped at the shoulders paired with a full-body clip tends to confuse the mapping. Start with a photo that shows as much of the body as your reference video does.
Reuse a reference clip you like across several character photos before you go hunting for new footage. Once you know a particular video transfers cleanly, running your next three subjects through that same clip is faster and far more predictable than gambling on an unfamiliar one each time.
The workflow fits in one sitting once you have your two inputs ready: a clear photo of the subject and a reference video of the moves you want it to perform.
- Pick a clean subject photo. A well-lit image with the subject facing the camera and the relevant part of the body visible, cropped close to how your reference video is framed.
- Choose a reference dance clip. Full body for a full-body result, upper body for a subtler animated portrait, filmed steadily with the dancer clearly separated from the background.
- Open Kling v2.6 Motion Control or Wan 2.2 Animate Animation. In the Picasso IA toolkit, upload the photo as the character image and the clip as the motion reference.
- Set the mode and orientation. Standard mode for a quick check, Professional for a cleaner final pass, and match the orientation setting to whichever input shows the fuller body.
- Generate, review and refine. Watch the clip at full size for warped hands or a drifting face, then rerun with a steadier reference clip or a different photo crop if something looks off, and upscale the export through the video upscaling workflow if you need a sharper final file.
A handful of habits separate a convincing dance clip from a wobbly one, none requiring technical skill, just attention to what you feed the model before you hit generate.
- Match the framing: keep the crop of your photo close to the crop of your reference video, full body to full body or upper body to upper body, so the model is not stretching motion across a mismatched pose.
- Light the subject evenly: hard shadows across a face or torso get read as extra detail during motion, which shows up as flicker once the subject starts moving.
- Start with Standard mode: run a cheap draft first to confirm the pairing works before spending a Professional-mode generation on a clip that might need a different reference anyway.
- Keep the background simple: a plain or blurred background in your photo gives the model one less thing to reconcile with the motion coming from the reference clip.
- Test one variable at a time: if a result looks off, change either the photo or the reference clip between runs, not both, so you can tell which input caused the problem.
This technique is genuinely useful and genuinely limited, and both are worth saying plainly. Fast, complex choreography with rapid spins or floor work confuses the motion mapping more often than a walk or a wave does, showing up as blurred limbs or a pose that briefly loses the shape of the body. Hands stay the hardest part of any figure to animate correctly, and a dance clip with finger detail close to the camera is where an extra joint or a warped grip is most likely. Long clips near the thirty-second ceiling drift more than short ones, with small errors in face shape compounding as the sequence runs on. Two dancers moving together, a duet or group routine, sits outside what either model is built for, since each generation maps one reference performance onto one character image. Audio merging, where offered, is explicitly experimental: expect timing that is close but not frame-accurate.
There is also a rights question worth taking seriously before you post anything. The model copies motion, not identity, so it does not reproduce the original dancer's face or likeness. It does reproduce their specific choreography, which some creators consider their creative work regardless of who performs it in the final clip. Use footage you filmed yourself, have permission to reference, or that is clearly licensed, the same standard you would apply to any other creative source.
Can any photo be turned into a dancing AI video?
Most photos work, but the cleanest results come from a well-lit subject facing the camera with the relevant part of the body visible and unobstructed. Photos that are heavily cropped, poorly lit, or show the subject at an odd angle tend to produce warped or unstable motion. If a first attempt looks rough, try a different crop of the same photo or a reference video with a similar framing before assuming the technique will not work for that subject.
Do I need to record my own dance video to use as a reference?
No, though it helps if you can. Any video that clearly shows the moves you want transferred works as a reference, whether you filmed it yourself or sourced footage you have rights to use. What matters technically is that the dancer is clearly visible against the background and moving at a normal, trackable pace, not where the footage originally came from.
Will the AI dance video keep my face and clothing accurate?
Largely yes, since the reference video only supplies motion and the subject image supplies everything about how the character looks, including face, outfit and background. Expect some softening of fine detail as the pose changes quickly, and expect longer clips to drift a little more than short ones as small errors accumulate across more frames.
How long can a dance video generated from a photo be?
It depends on the model and the orientation setting you choose. Kling v2.6 Motion Control generates up to ten seconds when the output follows your photo's orientation and up to thirty seconds when it follows the reference video's orientation instead. Wan 2.2 Animate Animation follows the length of the reference clip you provide. For a first test, a short clip is faster to review and cheaper to iterate on than a long one.
Which is better for dance videos, Kling Motion Control or Wan 2.2 Animate Animation?
Neither wins outright, which is why Picasso IA offers both. Kling v2.6 Motion Control gives a Standard and Professional quality split plus flexible clip length, while Wan 2.2 Animate Animation runs at a steady 24 frames per second with a built-in audio-merge option. Running the same photo and reference clip through both takes minutes and is the fastest way to see which handles your footage better.
Can I keep the music from my reference video in the finished clip?
Both models offer an audio option. Kling v2.6 Motion Control has a straightforward toggle to keep the original sound from the reference video. Wan 2.2 Animate Animation offers an audio-merge setting that the model itself describes as experimental, meaning the timing may not land perfectly. If precise sync matters, treat the merged audio as a draft and be ready to add music separately afterward.
Is there a free way to try this before paying for anything?
New Picasso IA accounts start with free credits, which is enough to run a first motion-transfer test and see whether the technique works well for your photo and reference clip before spending anything further. Beyond the free credits, current plans and their limits are listed on the pricing page, which is the only place worth checking for up-to-date numbers.
Can the same photo be animated with more than one dance style?
Yes, and this is one of the more useful parts of the workflow. The subject photo stays fixed while the reference clip changes, so you can run the same character image against several different dance videos and end up with a small library of distinct clips from a single source photo, without regenerating the character itself each time.
Why does the motion look wrong on some attempts?
The two most common causes are a mismatch between how the photo and the reference video are framed, and reference footage that moves too fast or cuts between camera angles for the model to track cleanly. Recrop the photo to match the reference video's framing, or pick a steadier, more continuous reference clip, and rerun the generation before concluding the subject simply does not work.
Is it okay to animate a photo using someone else's choreography?
The model does not copy a person's face or identity, only the pattern of movement in the footage you provide, so using your own photos raises no likeness issue. The choreography itself is a separate question: some dancers and choreographers consider a specific routine their creative work, so use footage you filmed, have explicit permission to reference, or that is clearly licensed for this kind of use, rather than any clip you happen to find online.
A photo sitting still on your camera roll is one upload away from becoming a moving clip. Open the Picasso IA toolkit, pick a reference video, and see what your own photo looks like dancing.