How to Animate a Photo With AI
The whole process takes about five minutes. Most of the quality is decided in the first thirty seconds — which image you pick.
18+ only · fictional characters only · no public gallery
How it works
Three steps, about five minutes end to end.
Choose the right source image
Sharp, well-lit, one clear subject, subject not cut off at the edge of frame. If your image fails any of these, fix it before generating — nothing later compensates for a weak source.
Match the motion to the subject
Portraits want subtle subject motion. Products and landscapes want camera motion. Applying a strong camera move to a face is the most common first-attempt mistake.
Run short, judge honestly, re-roll
Generate, watch it twice, and decide. Video generation is sampling, not rendering — a second attempt at the same settings is a genuinely different result, and it is cheaper than trying to prompt your way out of a bad one.
What makes it different
Upscale before, not after
Fixing resolution after generation is fighting the compression. If your source is small, upscale the still first — the model then has real detail to animate.
Crop to the subject
A busy frame splits the model's attention and every subject gets a little worse. One clear subject, filling a reasonable part of the frame, is the strongest input you can give.
Keep hands out of it
Hands are the weakest area of every video model available today. A composition that does not feature them is not a compromise, it is a good decision.
Expect to keep half
A 50% keep rate is normal, including for people who do this daily. Budget runs accordingly instead of treating a failed generation as a mistake.
Troubleshooting table
| Symptom | Cause | Fix |
|---|---|---|
| Face changes partway through | Clip too long, motion too strong | Shorter clip, gentler template |
| Almost no movement | Flat or very busy source | Crop tighter, choose a camera-move template |
| Warped hands | Model limitation | Crop them out or keep them still |
| Blurry, mushy output | Low-resolution source | Upscale the still first |
| Text in frame turns to nonsense | Model limitation | Remove or crop out the text |
| Background boils | Cluttered background | Blur or simplify it before generating |
Step one: pick a photo that can survive it
Animation amplifies whatever is already in the frame. A sharp source gets sharper motion; a soft, noisy source gets soft, noisy motion with new artefacts on top. Before uploading anything, ask three questions: is the subject in focus, is there exactly one thing the viewer is meant to look at, and is the subject fully inside the frame? A subject cropped at the edge tends to smear as the model tries to invent what is outside it.
If the answer to any of those is no, the cheapest fix is upstream. Crop, upscale, or pick a different photo. Every one of those costs nothing; a failed generation costs credits.
Step two: pick motion that suits the subject
There are two broad families. Subject motion moves what is in the frame — hair, fabric, breathing, a head turn. Camera motion moves the frame itself — a push-in, an orbit, a slow drift. Portraits and characters generally want the first. Products, landscapes and architecture generally want the second.
The failure to avoid is combining strong subject motion with strong camera motion. Both ask the model to change a lot of pixels at once, and that is exactly when identity and structure break down. If you want a dramatic result, get it from one of the two, not both.
Step three: read the result properly
Watch the clip twice at full size before judging it. The first pass tells you whether the motion is right; the second tells you whether the subject stayed itself. Most people conflate the two and end up changing the wrong thing.
If the motion is right and the subject drifted, shorten the clip. If the subject held and the motion was wrong, change the template. If both were wrong, the source image is the problem — go back to step one. This ordering saves more credits than any prompt trick.
When not to use AI animation at all
If you need an exact, repeatable performance — a specific gesture at a specific moment, matched to audio — this is the wrong tool. Rigged animation does that; generation does not. If the shot must be identical across ten variations, generation will fight you, because each run is a fresh sample.
And if the source photo shows a real, identifiable person, the rules are narrow: ordinary non-intimate content with their consent only. Intimate or sexual content involving real people is prohibited outright here and illegal in a growing number of jurisdictions.
Frequently asked questions
How long does it take to animate a photo?
Minutes per run in practice, and roughly five minutes end to end for a first usable result including one re-roll.
What is the best photo to start with?
Sharp, well-lit, one subject, fully inside the frame. That covers most of the quality difference between good and bad results.
Why did my result look melted?
Usually too much motion for the source. Choose a gentler template rather than adjusting the prompt.
Can I animate an old family photo?
Yes — restoration-style gentle parallax works well on archive photos. Upscale first if the scan is small.
Do I need editing software?
No. You get an MP4 back. An editor is only needed if you want to add sound or join several clips.
How many attempts should I expect?
Two or three for a first usable clip. Keeping about half of your generations is normal.
Try it on your own image
New accounts get starter credits. No card needed to run a first generation.
Try it on your own photo