HOME/BLOG/GUIDES
Guides

How to Make AI Video for Instagram Reels & TikTok (9:16)

Make AI video natively in 9:16 for Instagram Reels and TikTok — native vertical vs. cropping, safe zones, the 3-second hook, and which Oxava model to pick.

Egemen KüpçüJuly 7, 202616 min read
How to Make AI Video for Instagram Reels & TikTok (9:16)
Share

Open TikTok or Instagram Reels and watch how you hold your phone: upright, one thumb, full screen. That posture is the whole reason short-form video is vertical, and it's why a clip built for a horizontal frame arrives on these platforms looking like a guest who wore the wrong outfit — pillarboxed, shrunk to a strip in the middle, or auto-cropped in a way that lops the head off your subject. If you're learning how to make AI video for Instagram Reels and TikTok in the 9:16 format, the single most important decision happens before you write a word of your prompt: you generate vertical from the start, instead of trying to rescue a horizontal clip later.

This guide is about that format decision and everything that follows from it — why 9:16 is the default and not a preference, why native generation beats cropping, exactly which Oxava model and mode let you lock 9:16 directly, where each platform's interface hides your content, and how to structure the first three seconds and the overall rhythm so the clip survives the scroll. It's a companion to our deeper cinematic prompting guide, which covers the camera and motion language itself; here we focus on the vertical, platform-specific application of it.

Why vertical (9:16) is the default, not a preference

For years, "proper" video meant 16:9 — the shape of a TV, a cinema screen, a YouTube player. Short-form flipped that. TikTok, Reels, and Shorts all assume a 9:16 vertical canvas at 1080×1920 pixels, and they build their playback experience around it. A vertical clip fills the entire screen edge to edge. A horizontal one gets letterboxed into a small central band, or the platform auto-crops it to fit — and auto-crop rarely lands where you'd have chosen.

The stakes are highest in the opening moment. On feeds that autoplay one clip after another, viewers decide almost instantly whether to stay: studies of short-form behavior find that roughly 60% of viewers make the keep-or-scroll call within the first one to three seconds. If your video shows up as a shrunken strip while everything around it is full-bleed and immersive, you've handed those viewers an easy reason to swipe before your content ever gets a chance.

So vertical isn't a stylistic choice you make for polish. It's the native shape of the surface. Building anything else means fighting the platform's own layout — and the platform always wins.

Native 9:16 vs. cropping a horizontal video later

Here's the temptation: generate a normal 16:9 clip, then crop it to 9:16 in an editor. It feels efficient. It almost never works well, and the reason is specifically about how AI video generators think.

An AI model composes for the aspect ratio you give it. When you ask for a 16:9 clip, the model arranges the subject, the negative space, the camera move, and the action to look right in a wide frame — the subject sits in the middle third, with room breathing out to the left and right. Crop that down to a tall 9:16 slice and you throw away the sides the model deliberately used, while the top and bottom of the vertical frame stay empty because the model never planned for them. You end up with a subject that's too small, dead space above and below, and a camera move that was choreographed for a width you no longer have.

The principle, borrowed from how experienced creators treat aspect ratio: treat 9:16 as a starting constraint, not a post-production fix. When the model knows from the first token that it's filling a tall frame, it composes vertically — the subject stands upright in the space, the camera tilts along the vertical axis, and the important action sits where the tall canvas can hold it. This is "generate first, not crop later," and it's the difference between a clip that feels made for the phone and one that feels squeezed onto it.

Vertical composition also rewards different instincts than horizontal. Think close-ups and single upright subjects rather than sprawling wide shots; think vertical camera moves — a tilt up from a subject's shoes to their face, a push-in on a standing figure — rather than long horizontal pans that run out of frame almost immediately. A wide landscape that sings in 16:9 usually falls apart in 9:16; a person, a product held in the hand, or a single hero object stands tall and fills the frame naturally.

Choosing the right model and mode for vertical output in Oxava

This is where the practical decisions live, and Oxava's video models behave differently enough that picking the right one saves you a re-roll. The core question is simple: can I select 9:16 directly, or do I need to start from a vertical source image? The answer depends on the model and the mode.

Two modes matter here. In text-to-video (t2v) you describe a scene in words and the model generates everything, framing included. In image-to-video (i2v) you hand the model a still as its first frame and it animates from there. Where a model lets you set the aspect ratio changes between these two modes.

Text-to-video: pick 9:16 directly on Kling v3 and Seedance 2.0

In text-to-video, both Kling v3 (Standard, Pro, and Turbo) and Seedance 2.0 accept an aspect-ratio setting, so you simply choose 9:16 in the panel and the model composes vertically from scratch. This is the cleanest path to a native Reels/TikTok clip when you don't already have a specific image to animate — you write your motion and scene, set the frame to 9:16, and generate.

A vertical-aware t2v prompt looks like this:

"Medium vertical shot of a woman in a mustard raincoat walking toward camera down a rain-slicked city street at night, neon signs glowing on both sides. Slow push-in; she fills the tall frame from waist up. Handheld micro-movement, shallow depth of field, cinematic, 9:16 vertical."

Notice how the framing cues ("vertical shot," "fills the tall frame," "9:16 vertical") reinforce the setting rather than fight it.

Image-to-video: who lets you set 9:16, and who needs a vertical source

Image-to-video is where the models split, and it's the nuance most people miss.

  • Seedance 2.0 accepts an aspect-ratio setting even in i2v, so you can select 9:16 directly regardless of your source image's shape. This makes it the most flexible choice for vertical control.
  • Gemini Omni Flash (image-to-video only) offers a 16:9 / 9:16 choice in its panel, so you set 9:16 there too.
  • Kling v3 (all variants) does not take an aspect-ratio setting in i2v — the output ratio follows your source image. To get a vertical clip, you must start from a vertical (9:16) image.
  • Grok Imagine 1.5 (image-to-video only) behaves the same way: no aspect setting in i2v, so the output inherits the source image's ratio. Again, feed it a vertical image to get vertical out.

Here's the whole picture at a glance:

Model Text-to-video Image-to-video How to get 9:16
Seedance 2.0 Pick 9:16 directly Pick 9:16 directly Choose 9:16 in either mode — most flexible
Kling v3 (Std / Pro / Turbo) Pick 9:16 directly Follows the source image t2v: pick 9:16 · i2v: start from a vertical image
Gemini Omni Flash i2v only Offers 16:9 / 9:16 Choose 9:16 in the i2v panel
Grok Imagine 1.5 i2v only Follows the source image Start from a vertical image

The takeaway is a two-step rule. If you're generating from text, Kling v3 or Seedance 2.0 with 9:16 selected gets you native vertical in one move. If you're animating an existing image and it isn't already vertical, either use Seedance 2.0 or Gemini Omni Flash (which let you force 9:16), or first make a vertical version of the image and hand that to Kling v3 or Grok Imagine 1.5. Because Oxava generates stills and video in the same studio, producing a 9:16 source image to feed an i2v model is a quick first step, not a separate project.

One small reassurance on cost: generating vertical doesn't cost you more. At a given resolution, a 9:16 clip carries the same pixel budget as a 16:9 one, so choosing tall is a purely creative decision — never a budget penalty.

For the complete, tool-agnostic image-to-video process — input prep, motion-only prompting, duration and export — see our dedicated image-to-video workflow guide; we won't repeat those steps here, since this guide's job is the format layer on top of them.

The safe-zone map: where the interface hides your content

Getting 9:16 right is only half the battle. Every short-form platform layers its own interface on top of your video — buttons, captions, usernames, progress bars — and those elements sit in predictable zones. Put something important where a "Follow" button or a caption lands, and a chunk of your audience never sees it.

The obstructed regions differ slightly by platform, but the pattern is consistent: the top holds navigation, the right rail holds the engagement buttons (like, comment, share, profile), and the bottom holds the caption, handle, and audio label — plus, in ads, your call-to-action.

Zone TikTok (approx.) Instagram Reels (approx.) What sits there
Top ~10% ~8% Search, tabs, status bar
Right rail ~10% ~12% Like, comment, share, profile buttons
Bottom ~20% ~25% Caption, username, audio label, CTA

YouTube Shorts follows the same logic — title and channel along the bottom, action buttons down the right — so treating it like a slightly more forgiving Reels layout keeps you safe.

The practical rule that falls out of this: keep your subject and any critical text or logo inside the center third of the frame, comfortably clear of the top strip, the right rail, and — most importantly — the bottom quarter, which Reels eats more aggressively than TikTok. If you burn captions into the video yourself, raise them above the platform's own caption zone so the two don't stack into an unreadable pile. And if a clip is destined for both TikTok and Reels, compose for the stricter of the two — Reels' 25% bottom bite — so the same export is safe everywhere.

The first three seconds: the opening that survives the scroll

Because most viewers decide within one to three seconds, the opening of a short-form clip does more work than the entire rest of it. A vertical AI video gives you a specific advantage here — a full-screen, immersive frame — but only if the first beat earns attention immediately.

A few things reliably help the opening survive:

  • Lead with motion or a face, not a slow fade. A static hold or a gentle fade-in wastes your most valuable second. Start the clip already in motion — a push-in already underway, a subject already turning toward camera, the product already catching the light.
  • Put the payoff, or a promise of it, up front. Don't save the interesting moment for the end where most people won't reach it. Tease it in frame one.
  • Fill the frame. In vertical, an upright subject that occupies most of the tall canvas reads instantly. A small subject swimming in empty space reads as "skip."
  • Match the audio to the first beat. On models that generate sound, prompt an opening that lands on the action — a crisp sound tied to the first movement gives the eye and ear something to lock onto together.

In prompt terms, that means front-loading the action. Instead of "a bottle sits on a table, then the camera moves in," write the motion as already happening:

"Extreme close-up, vertical: a matte-black serum bottle catches a sweep of light as the camera pushes in from the very first frame; a single droplet runs down the glass. Premium, cinematic, 9:16."

Short-form rhythm: shot length, cuts, duration, and export

Short-form has its own metabolism, and it's faster than film. A single AI generation is usually only a few seconds long, which fits perfectly — you build a Reel or TikTok from several short, punchy clips rather than one long take.

Cut roughly every three seconds. Short-form thrives on pace; a static shot that lingers past a few seconds invites the scroll. Rather than asking one generation to carry a whole scene, produce several tight clips — each with one clean action or camera move — and cut between them. This also plays to how AI video behaves: shorter clips drift and warp less than long ones, so the format's rhythm and the model's strengths point the same direction.

Keep each clip to one idea. One primary action plus one camera move per generation. Trying to stage three things in a five-second vertical clip rushes all of them; the cinematic guide goes deep on why one clean beat beats a crowded timeline.

Mind the duration ceilings. Reels, TikTok, and Shorts all now support longer runtimes, but short-form engagement rewards brevity — many of the best-performing clips land in the 15-to-30-second range. Build to the content, not the maximum.

Export for the phone. Render at 1080×1920 in MP4 / H.264 — the universal, upload-anywhere combination every platform ingests cleanly. Oxava clips download straight from the studio, so you can take them into your editor, stack your cuts, add captions above the safe zone, and export the final vertical cut without a format conversion in the middle.

Common mistakes to avoid

A handful of errors account for most disappointing vertical clips:

  • Forcing a horizontal composition into a tall frame. Generating 16:9 and cropping leaves a small subject stranded in a sea of dead space. Generate 9:16 from the start.
  • Choosing the wrong model/mode for vertical. Feeding a horizontal image to Kling v3 or Grok Imagine 1.5 in i2v gives you a horizontal clip — those models follow the source. Use Seedance 2.0 or Gemini Omni Flash if you need to force 9:16 from a non-vertical image, or prepare a vertical source first.
  • Placing the subject in the thumb zone. Anything in the bottom quarter or the right rail gets covered by captions and buttons. Keep it centered.
  • Stacking your captions on the platform's captions. Burned-in text that lands in the platform's own caption band becomes an unreadable pile. Lift yours above it.
  • One long take instead of cuts. A single slow clip reads as sluggish in a fast feed and gives the model more room to drift. Cut every few seconds.
  • Expecting a talking presenter from a standard clip. Lip-synced, speaking presenters are a different tool, not something a general motion prompt reliably produces — see below.

Where vertical AI video fits your other content

If you sell products, the vertical format is how you get catalog shots moving on social. The full process for turning a product photo into ad-ready motion — hero framing, clean backgrounds, and export — lives in our guide on turning product photos into videos; bring those clips into 9:16 with the model rules above and you've got Reels-ready product content.

For a single high-impact effect rather than a full format workflow — like driving a photo with motion control to make a subject dance — see our walkthrough on Kling motion control for Instagram Reels. And if you want a talking, lip-synced presenter delivering a line to camera, that's a dedicated channel of its own; our guide to AI talking-avatar video for creators covers it, since a standard motion prompt won't produce reliable lip-sync.

Frequently Asked Questions

Can I convert a horizontal AI video to 9:16 afterward? You can, but you'll almost always lose quality and composition. Cropping a 16:9 clip to 9:16 throws away the sides the model deliberately composed with and leaves the tall frame's top and bottom empty, so the subject ends up small and awkwardly placed. The far better approach is to generate 9:16 natively from the start, so the model composes for the vertical frame — treat 9:16 as a constraint you set up front, not a fix you apply later.

Which Oxava model gives me full control over vertical output? Seedance 2.0 is the most flexible: it lets you pick 9:16 directly in both text-to-video and image-to-video. For text-to-video, Kling v3 also lets you choose 9:16 directly. In image-to-video, Kling v3 and Grok Imagine 1.5 instead follow your source image's ratio, so you feed them a vertical image; Gemini Omni Flash and Seedance 2.0 let you set 9:16 in the panel.

Can I use the same video for both Reels and TikTok? Yes — both use the same 9:16 / 1080×1920 canvas, so one export works for both. The one thing to plan for is the safe zone: Instagram Reels covers more of the bottom of the frame (~25%) than TikTok (~20%). Compose for the stricter Reels layout — keep your subject and text clear of that bottom quarter — and the same clip stays safe on both platforms.

Does generating vertical video cost more than horizontal? No. At the same resolution, a 9:16 clip carries the same pixel budget as a 16:9 one on Oxava, so choosing vertical is purely a creative decision with no extra cost attached.

How long should a Reels or TikTok clip be? Build from several short clips — each a few seconds with one clean action or camera move — and cut roughly every three seconds to keep the pace up. Total runtime is best kept tight; many top-performing short-form videos land in the 15-to-30-second range. Shorter, well-cut clips also drift and warp less than one long generation.

Start generating vertical in the studio

Making AI video for Reels and TikTok comes down to one habit and a few informed choices: generate 9:16 natively instead of cropping later, pick the model and mode that give you the vertical control you need, keep your subject and text inside the safe center of the frame, and open with motion strong enough to beat the scroll. Get the format right and the rest of your cinematic craft — camera moves, pacing, consistency — finally lands where people will actually see it.

The fastest way to internalize it is to make one. Open the Oxava studio, choose Seedance 2.0 or Kling v3, set the aspect ratio to 9:16 (or, for image-to-video on Kling and Grok, start from a vertical image), and generate a clip built for the phone from the very first frame. Prototype your framing, cut a couple of short clips together, and you'll have a native vertical video ready for Reels and TikTok — no cropping required. When you want to sharpen the motion and camera language inside that vertical frame, the cinematic prompting guide is the next step, and the 2026 video model comparison helps you weigh the models beyond just their vertical behavior.

FOUNDER & AUTHOR

Egemen Küpçü

Egemen Küpçü is the founder of Oxava, with 10+ years of hands-on experience in 3D and visual production. He writes about the craft of generating product, brand and campaign visuals with AI.

Subscribe to our newsletter

Be the first to hear about new techniques, model updates and ideas on AI generation.