HOME/BLOG/GUIDES
Guides

How to Turn an Image Into a Video With AI

Step-by-step image-to-video AI workflow: prep your image, choose a model, write a motion-only prompt, and export a polished clip. No guesswork, no wasted credits.

Egemen KüpçüJune 19, 202615 min read
How to Turn an Image Into a Video With AI
Share

You have an image you love — a product shot, a portrait, an illustration, a landscape — and you want it to move. Knowing how to turn an image into a video with AI is now one of the most useful skills a creator can have: you hand the model a still frame, describe the motion you want, and it animates the scene. But the gap between a clip that looks magical and one that looks like a melting nightmare comes down to workflow. This guide walks you through the complete image-to-video AI workflow from start to finish — image prep, model choice, prompting, settings, and export — so you get a usable clip without burning through credits on guesswork.

This is the practical, end-to-end companion to our deeper dive on how to prompt AI video generators, which focuses on the cinematography language itself. Here we zoom out to the whole process.

What Image-to-Video AI Actually Does (and What It Doesn't)

The single most useful thing to understand before you start: in image-to-video, your image is the first frame. The model doesn't redraw your scene from scratch — it takes the pixels you give it and predicts how they should move over the next few seconds. Everything that exists in the output grows out of what was already in the input.

That has two big consequences:

  • Input quality is a hard ceiling. A blurry, low-resolution, or heavily compressed image will produce a blurry, low-resolution video — often worse, because motion amplifies every flaw. The model can't invent detail that isn't there.
  • It animates what it sees, not what you wish were there. If your subject is half cut off at the edge of the frame, asking for a "full body turn" will fail. The model can only work with the composition in front of it.

This is the core difference from text-to-video (t2v), where you describe a scene in words and the model generates everything — composition included. Roughly, the rule of thumb is: use i2v when you already have the exact visual you want (a specific product, a real face, a finished illustration) and need to preserve it precisely; use t2v when you're still exploring and want the model to invent the scene. For a breakdown of which models excel at each, see our 2026 AI video model comparison.

On Oxava, image-to-video runs in the studio on two models: Kling v3 and Seedance 2.0. Both take a still image plus a motion prompt and return a short clip. We'll cover when to reach for each below.

Step 1: Input Image Preparation

Because the input is your ceiling, this is the step most people skip and most regret. Spend two minutes here and the rest of the workflow gets easier.

Resolution. Aim for at least 1024px on the long edge; somewhere in the 1K–2K range is the sweet spot. Going to 4K rarely improves the result — most models downscale internally for motion generation, so you spend extra detail for nothing. If your source is below ~1024px, upscale it first (more on that in Step 5).

Format. Use a PNG or a high-quality JPEG. Avoid screenshots saved as small, heavily compressed JPEGs — compression artifacts that are barely visible in a still become crawling, shimmering noise once the frame starts moving.

Composition. Keep your subject roughly centered with a little breathing room at the edges. Models animate the center of the frame most reliably, and edge crops limit what kind of motion is even physically possible.

Lighting. Soft, even light animates cleanly. Hard shadows with sharp edges tend to flicker or smear as the model interpolates motion, because it has to guess what's hidden in the shadow as things shift.

Background. Simpler is safer. A clean or shallow-depth background gives the model fewer competing elements to get wrong; a busy, cluttered scene invites warping and drifting artifacts.

Aspect ratio. Lock this now, before you generate. If your final video needs to be 9:16 for Reels, prepare a 9:16 image — cropping a 16:9 clip afterward throws away resolution and often ruins the composition.

Faces and small text. If a face matters, make sure it's fully visible and in sharp focus. And know that small text — logos, labels, captions — tends to warp during animation; if your image contains it, a constraint in your negative prompt ("no text distortion, keep label stable") helps, but the most reliable fix is keeping critical text out of the moving area.

Step 2: Choosing Model and Mode

The two models on Oxava have different temperaments, and matching the model to your image type saves you re-rolls.

Kling v3 outputs up to 1080p and is known for strong character consistency and the ability to hold a subject steady across multi-shot motion. It's a dependable choice for people, products, and anything where you need the subject to stay recognizably itself through the clip.

Seedance 2.0 leans into multi-reference workflows — it can take several reference inputs and use an object-reference lock to keep a specific element consistent. That makes it strong when you're combining or anchoring multiple visual elements.

Here's a quick map from image type to a sensible starting model:

Image type Suggested model Why
Portrait / person Kling v3 Strong character + face consistency
Product (single object) Kling v3 or Seedance 2.0 Both hold the object; Seedance adds reference lock
Illustration / AI art Kling v3 Holds stylized look without forcing realism
Landscape / environment Either Forgiving subject; both handle ambient motion well
Multi-reference scene Seedance 2.0 Built for multiple references + object lock

Quality mode. Most models offer a Standard and a Professional (or higher) mode. Professional mode produces cleaner detail and steadier motion but costs more credits. The efficient habit: prototype in Standard to nail your prompt and settings, then re-run the winning version in Professional. There's no point paying premium-mode rates for a take you're going to discard.

If you want a side-by-side on capabilities and cost before you commit, our model comparison goes deeper.

Step 3: Writing a Motion-Only Prompt

This is where image-to-video prompting diverges sharply from image prompting, and it's the mistake we see most often. You do not need to re-describe the contents of the image. The model already has the picture — it's the first frame. Re-describing the subject ("a red sneaker on a white background") wastes the prompt and can even confuse the model into trying to regenerate the scene. Describe only the motion.

For Seedance 2.0, the documentation around the model describes a useful six-part structure — reported by guides like apiyi's Seedance 2.0 prompt reference — that's worth following:

  1. Subject — who/what moves (briefly)
  2. Action — the primary motion
  3. Environment — how the surroundings respond
  4. Camera — one primary camera command
  5. Style — mood/pacing cues
  6. Constraints — what to avoid

The same source suggests keeping it to roughly 60–100 words and issuing a single primary camera move — stacking three camera commands into one prompt is the fastest way to get chaotic, drifting footage.

For Kling v3, a similar logic applies, often framed (as in VicSee's Kling 3.0 prompt guide) as: a dynamic verb for the action + a clear camera technique + how elements interact. Lead with the strongest verb, then specify the camera.

A practical camera-move vocabulary, sorted by how reliably AI handles it:

  • Easy (high success rate): slow push-in, gentle pan left/right, static locked shot, slow zoom-out.
  • Medium: tracking shot, dolly forward, slow orbit around the subject.
  • Hard (expect re-rolls): handheld shake, walking POV, complex multi-axis moves combining several motions at once.

When in doubt, choose an easy move. A clean slow push-in beats a botched orbit every time.

Motion intensity matters. Asking for too much movement — a person doing three things while the camera orbits and the background swirls — overloads the model. Keep one primary action plus one camera move per clip.

A reusable negative prompt template covers the usual failure modes:

distortion, warping, melting, extra limbs, deformed hands, flickering, text artifacts, sudden cuts, morphing

And for anything more than a single beat, a timestamped beat structure helps the model pace itself — e.g. "0–2s: subject begins to turn; 2–4s: camera pushes in; 4–5s: subject settles, holds gaze." Our cinematic prompting guide covers this temporal language in much more depth.

Step 4: Duration, Resolution, and Aspect Ratio Settings

The settings panel quietly decides whether your clip lands. Match the duration to the complexity of what you're asking for:

Duration Best for
3–5 sec A single action or a clean loop (push-in, product rotate, hair in the wind)
6–8 sec A two-beat sequence (turn, then react; reveal, then settle)
9–12 sec Complex multi-step motion — higher risk, more re-rolls

A counterintuitive truth: longer isn't better. The more seconds you ask for in a single generation, the more chances the model has to drift, warp, or lose the subject. For a longer final video, generate several short, clean clips and stitch them together in editing rather than fighting for one perfect twelve-second take.

Aspect ratio follows your destination:

Platform Aspect
Website / YouTube 16:9
Reels / TikTok / Shorts / Stories 9:16
Instagram feed 4:5 or 1:1
LinkedIn 1:1 or 16:9

Set this to match the image you prepared in Step 1.

Resolution and credit efficiency. Test in Standard resolution first to confirm the motion and framing are right, then re-generate the keeper in Professional / higher resolution. This is the single biggest credit-saver in the whole workflow.

Step 5: What Changes by Image Type

This is where a generic workflow breaks down — a portrait and a landscape need genuinely different handling. Here's how to adapt each step by what's actually in your image.

Product photos. Keep the background clean and lean on motion-led prompts (slow rotation, parallax, a soft camera push) rather than asking the product itself to deform. Seedance 2.0's object-reference lock helps hold the product's shape and label steady. If you want eye-catching social content, the floating product video effect is one of the highest-performing formats — suspended in mid-air with a slow orbit. Product video is a deep topic of its own, so we've put the full process — including hero-shot framing and ad-ready output — in a dedicated guide on turning product photos into videos.

Illustrations and AI art. Lower your expectations for realistic physics — stylized art doesn't move like the real world, and pushing for photoreal motion fights the source. Instead, lock the style and ask for atmospheric motion: drifting clouds, flickering light, a gentle parallax, hair or fabric in a breeze. These keep the illustration alive without breaking its look.

Portraits. Use Professional mode here — faces are unforgiving. If your model supports element binding or a face/identity lock, turn it on. Be cautious with hand animation: hands are the highest-risk element in any AI video, so favor head turns, expression shifts, and subtle camera moves over having the subject gesture or manipulate objects. If you want full talking-head, lip-synced output, that's a different channel — see our guide on UGC-style product videos with AI avatars.

Landscapes and environments. The most forgiving category. There's no face or hand to deform, so you can be ambitious: parallax through layers, moving light and shadow, water ripples, wind through foliage, drifting fog. Ambient motion looks great and rarely warps.

Low-resolution inputs. Don't animate them — upscale first. Motion magnifies the softness and artifacts in a small image, so a quick upscale before i2v dramatically improves the result. Our image upscaling guide covers how to get a clean, high-resolution input ready for video.

Step 6: QA Checklist and Common Errors

Most "the AI is bad" frustration is actually a fixable input or settings problem. Here are the failures we see most, and what causes them:

Symptom Likely cause Fix
Warping / drifting background Cluttered input image Use a cleaner, simpler background
Subject does the wrong thing Vague motion prompt Use one clear action verb + one camera move
Cropping / black bars Mismatched aspect ratio Match image and output ratio
Chaotic, jittery motion Too many simultaneous motions One primary action per clip
Soft, mushy output Low-resolution input Upscale before generating
Subject regenerates / morphs Re-describing image contents Prompt motion only
Drift in long takes Single clip too long Split into shorter clips

A short pre-flight QA before you generate:

  • Input is ≥1024px, clean format, no heavy compression
  • Subject centered with edge breathing room
  • Aspect ratio matches the target platform
  • Prompt describes motion only — one action, one camera move
  • Negative prompt covers warping/distortion

And after you generate, review for:

  • Face stability — does the face stay consistent and recognizable?
  • Hand integrity — any extra fingers, melting, or morphing?
  • Physics — does the motion look plausible for the subject?
  • Logo / text — do any labels or words stay readable?
  • Loop — if it's meant to loop, does the last frame return cleanly to the first?

If a take fails one of these, change one variable and re-roll rather than rewriting everything — it's how you learn what the model responds to.

Step 7: Export and Post-Production

A raw AI clip is rarely the finished product — treat it as strong footage, not a final cut.

  • Trim weak frames. AI clips often have a soft first or last half-second as the motion ramps in or out. Trim them; your clip gets punchier.
  • Add the layer the model can't. Captions, voiceover, and music do enormous work for engagement, especially on social platforms where most people watch muted.
  • Match the platform's pacing. A clip that feels slow on a website can feel perfect as a fast, looping Reel. Cut to the platform.
  • Brand it. Drop in your logo, colors, or end card in editing rather than hoping the model bakes them in.
  • Pick the right codec. MP4 / H.264 is the universal, upload-anywhere choice. If you're doing heavier editing, ProRes preserves more quality through the process.

On Oxava, your generated clips download directly from the studio, so you can take them straight into your editor of choice and finish the cut.

Frequently Asked Questions

Do I need a long, detailed prompt for image-to-video? No — and a long prompt can actually hurt. Because your image is already the first frame, you only need to describe the motion: one clear action plus one camera move. A focused 60–100 word motion prompt usually outperforms a sprawling one that re-describes everything in the scene.

My faces and hands keep distorting. What should I do? Faces and hands are the hardest elements for any AI video model. Use a higher quality mode (Professional), enable any face/identity lock your model offers, and add "deformed hands, extra fingers, face warping" to your negative prompt. The most reliable fix, though, is to avoid hand-heavy motion entirely — favor head turns, expression changes, and camera moves over gestures.

How long should my clip be? For a single action or a loop, 3–5 seconds is ideal. For a two-beat sequence, 6–8 seconds. Going beyond that in one generation raises the odds of drift and warping — for longer videos, stitch several short, clean clips together in editing.

Can I animate illustrations or AI art, not just photos? Yes. Illustrations and AI art animate well as long as you don't expect real-world physics. Lock the style and ask for atmospheric motion — drifting light, parallax, a gentle breeze — rather than complex realistic movement. The stylized look is preserved and the scene comes to life.

Kling v3 or Seedance 2.0 — which should I pick? For portraits and character-consistent motion, start with Kling v3. For multi-reference scenes or when you need to lock a specific object's appearance, Seedance 2.0 is built for it. For most single-subject clips either works, so prototype with the one that fits your image type. Our model comparison breaks down the trade-offs in detail.

Bring Your Image to Life

Learning how to turn an image into a video with AI isn't about finding a magic prompt — it's a repeatable image-to-video AI workflow: prepare a clean, high-resolution input; pick the model that matches your image type; write a motion-only prompt with one action and one camera move; dial in duration and aspect ratio for your platform; and finish the clip in editing. Get those steps right and the model does the rest.

To go further, our cinematic prompting guide covers the temporal and camera language in depth, and our AI video model comparison helps you choose the right model for every project.

Ready to animate your image? Open the Oxava studio, upload your image, choose Kling v3 or Seedance 2.0, and generate your first clip. Start in Standard mode to find the motion you love, then re-run it in Professional for the final export — and you'll have a polished video from a single still.

FOUNDER & AUTHOR

Egemen Küpçü

Egemen Küpçü is the founder of Oxava, with 10+ years of hands-on experience in 3D and visual production. He writes about the craft of generating product, brand and campaign visuals with AI.

Subscribe to our newsletter

Be the first to hear about new techniques, model updates and ideas on AI generation.