KitanaAI Photo & VideoStudio
video maker · 5 min read

Choosing the Right Motion for an AI Photo-to-Video Clip

The four camera and subject motions, which kind of photo each one suits, and why the wrong motion is the most common reason a generated clip looks fake.

By the Kitana team
CTVIDEO MAKER

A photo-to-video model does one thing well: it continues the moment in a still image for a few seconds, with motion that is subtle and physically plausible. Camera drifts, hair moves, light shifts, the subject breathes. What it does badly is invent things the photo never showed. Almost every fake-looking clip comes from asking it to do the second thing while thinking you asked for the first, and the motion preset is where that choice is made.

What the four motions do

Kitana's Photo to Video tool offers four presets. Two move the camera, one moves the subject, and one does both.

Slow Push In. The camera creeps closer along a straight line. Nothing new is revealed; the scene simply grows in the frame. The model has to keep the image coherent as it enlarges, which is the easiest task of the four. Works on portraits, products, landscapes, almost anything.

Gentle Drift. A soft sideways float. The camera slides a little to one side, which reveals a sliver of the scene beyond the frame and shifts the parallax between foreground and background. Suits landscapes, interiors and any photo with clear depth layers; suits flat, tightly cropped portraits less well.

Wind & Hair. The camera stays put and the subject moves: air through hair, a scarf, a curtain, leaves, the surface of water. The model animates whatever loose material it can see. Wonderful on a portrait with hair down; pointless on a bald head against a brick wall.

Cinematic Orbit. The camera arcs around the subject, revealing the side that the photo did not show. This is the most striking preset and the most demanding, because it asks the model to invent what was hidden. It needs a subject with real depth, a clear silhouette and enough background for the parallax to read.

Matching the motion to the photo

The reliable way to choose is to ask what the photo already implies.

  • A head-and-shoulders portrait, front-lit, plain background. Push In. There is nothing to drift past and nothing to orbit around. Wind & Hair if the hair is down.
  • A portrait with hair down, or a scarf, or an outdoor breeze. Wind & Hair first. Push In as the safe alternative.
  • A landscape or cityscape with foreground and distance. Gentle Drift. The parallax between near and far is what makes it feel filmed.
  • A three-quarter portrait with visible depth, a shoulder turned, a background with structure. Cinematic Orbit is possible here, because there is a side to reveal and depth to arc through.
  • A product on a surface. Push In or a very gentle Orbit. Products have simple silhouettes and the model handles the hidden side well.
  • A group photo. Push In only. Every other motion multiplies the chance that one face drifts.

If two motions seem plausible, choose the one that reveals less. A clip that could have been more dramatic is a clip; a clip that warps the subject is a discard.

What the model keeps, and what it will not do

The tool is instructed to keep the subject's identity, clothing and setting exactly as they appear in the photo: the same moment continuing, not a new scene. Motion is subtle and physically plausible. That is a deliberate boundary, and knowing it prevents most disappointment.

It will not make the person speak, wave, walk, or turn to face the camera. It will not add people or objects. It will not change the weather or the time of day. Each of those is a new scene, and asking for one produces either nothing or a warping mess. How the underlying model works, and why these limits exist, is covered in AI photo to video: how the technology works.

What it will do, and does well: a breath, a blink, a shift of weight, hair and fabric in air, water and leaves, the drift of light, and the camera moves above. That is enough to make a still feel like a frame from footage, which is the effect the format is for.

The source photo for video

Video is less forgiving than a still restyle, because any flaw in the photo is animated.

  1. Sharp. A soft photo becomes a soft clip, and the motion makes the softness obvious.
  2. Vertical. The output is a vertical clip sized for Reels, TikTok and Shorts. A landscape photo will be cropped, and the crop decides the composition, so shoot or choose vertical.
  3. Subject clear of the edges. Every motion moves the frame a little. A subject touching the edge gets cut off or stretched.
  4. Uncluttered background. Fewer objects means fewer things for the model to keep consistent while the camera moves.
  5. For Orbit, a silhouette. A subject that separates cleanly from the background, with some sense of a side to reveal.

The general checklist for video sources is in preparing for AI video: tips for content creators.

Judging a clip

Watch it twice. The first time, look only at the subject's face and ask whether it stayed the same person throughout. The second time, look only at the background and ask whether anything popped in, vanished or slid unnaturally. A clip that passes both is usable. A clip that fails the first is a discard regardless of how good the motion looks.

A clip takes a few minutes rather than seconds. You do not have to wait on the screen; tap Notify Me and the app sends a notification when it is ready. The wait is the same for every motion, so there is no speed reason to pick one over another.

Photo to Video is the one video tool among the fourteen in the Kitana apps. The free plan includes one short clip; Kitana Pro includes ten a month at full length and resolution.

Frequently asked questions

Why does my clip look fake even though the photo was good?
Usually the motion asked for something the photo could not support. An orbit around a flat, front-lit selfie has no depth to reveal, so the model invents it and the face warps. Match the motion to the photo and most of the fakeness disappears.
Which motion is safest?
Slow Push In. It moves the camera along one axis toward the subject, which reveals nothing new and asks the model only to keep the scene coherent while it enlarges. It works on almost any photo.
Can the person in the clip talk or wave?
No. The clip continues the moment in the photo with subtle, physically plausible motion: camera movement, breathing, hair and fabric, light. Large gestures and speech would require inventing a new scene, which is not what this tool does.
Why does the subject sometimes turn into a different person mid-clip?
Because the motion asked the model to show a view it had never seen. An orbit reveals the side of a face that was hidden, and the model has to invent it. Push In and Drift keep the visible view, which is why they drift far less.
Does the free plan include video?
The free plan includes one short clip, and Pro includes ten a month at full length and resolution. Both use the same four motions.

Ready to put this into practice?

Create with Kitana using the tool that fits this guide.

Turn a photo into video

Ready to try it yourself?

Download Kitana and create your first AI photo in under a minute.