
Open X, Reels, or YouTube Shorts this week and you've almost certainly stopped on one: a product ad where the lighting changes mid-shot, a weirdly satisfying physics clip where water pours exactly right, an avatar delivering a line that lands. A lot of what's flooding your feed right now traces back to one tool — Gemini Omni, the video model Google unveiled at I/O on May 19, 2026 and the subject of this review. Part of why it went everywhere so fast: it's surfacing free inside YouTube Shorts / Create, in reach of billions of users, and it edits video in a way nobody's really seen at this scale — you change a finished clip by talking to it. "Make it night." "Have her turn and walk left." "Add rain." No regenerating from scratch, no timeline scrubbing. This Gemini Omni review is the tool behind those videos: what it actually is, what people are making with it right now, where it's genuinely strong, where it falls short — and whether it's worth your time. Short version up top: the headline isn't raw frame quality, it's conversational, scene-memory editing — and Gemini Omni is now live in Oxava's studio, so you can stop scrolling and make your own.
Google positions Gemini Omni as an "any-to-any" multimodal model: it accepts a mix of inputs — text, images, a voice/audio reference, and video — and, right now, produces one kind of output: video. That "right now" matters. Google has said image and audio outputs are on the roadmap, but they are not available today. So despite the "omni" name and the multi-input flexibility, this is a video generation model at launch, not an image or music model. If you're comparing it to a text-to-image tool, you're comparing the wrong things.
The first model in the family is Gemini Omni Flash — the fast, cost-efficient tier meant for high-volume creative work. When people say "Gemini Omni" today, Flash is almost always what they're using.
A framing you'll see in a lot of coverage is that Omni "combines Veo, Genie, and Nano Banana into one model." Treat that as analyst shorthand, not an official spec — reviewers describe it that way to explain the breadth, but Google's own framing is narrower: Omni brings Gemini's world knowledge together with a sharper grasp of physics and motion. It's a useful mental model for what it can do, but don't quote it as Google's architecture claim.
Every current video model can generate a clip. What sets Gemini Omni apart is what happens after the first clip exists.
Omni keeps a memory of the scene — the characters, the objects, the layout, the lighting — and lets you revise it through plain-language instructions, one at a time, without rebuilding from zero. You say "give him a red jacket," and the same character in the same scene gets a red jacket; the face, the framing, and the background stay put. Then "now zoom in slowly," then "add a second person walking past." Each instruction lands as a targeted edit on a persistent scene rather than a fresh roll of the dice.
This is a genuinely different loop from the generate-and-pray cycle most creators know. Instead of tweaking a giant prompt and regenerating the whole thing — hoping the one detail you wanted to change doesn't drag five other things with it — you converge on the shot conversationally. That's exactly why the clips travel so well: creators can chase a specific idea to completion instead of settling for whatever the model happened to produce.
Underneath it is Omni's improved physics intuition. Google emphasizes a better sense of gravity, kinetic energy, and fluid dynamics, so edited motion holds together more believably: a dropped object falls with weight, water moves like water, a character's momentum carries through a turn. That physical coherence is what makes the edits feel like adjustments to a real scene rather than a new hallucination each time — and it's the secret behind a whole genre of "how is this AI?" clips.
The reason Gemini Omni is all over your feed isn't the spec sheet — it's what the talk-to-edit loop lets people ship. Four formats are doing most of the numbers, and you can build a version of each yourself.
Conversational product ads. This is the breakout use. Start from a product shot, generate a clip, then direct it like a tiny film set: "put it on a marble counter," "warm sunset light," "slow push-in," "now show it in a hand." Marketers are cutting polished 10-second spots by talking a scene into shape instead of re-rendering a prompt twenty times. Have a product? Try your own version in Oxava — start from the photo, converse your way to the ad.
Oddly satisfying physics clips. Water pouring, sand collapsing, dominoes, slow-motion splashes — Omni's physics intuition makes these read as real, and "satisfying" content is pure retention fuel on Shorts and TikTok. The move is to generate the base motion, then refine the timing and lighting by instruction until it hits that hypnotic loop.
Avatars and talking-style shorts. Presenter clips, character bits, spokesperson lines — controllable, editable characters suit the format (with the consent caveat below). Creators reshape delivery, wardrobe, and setting conversationally instead of re-rolling a whole generation to fix one detail.
Remixes and reframes. Take a scene and restage it — swap the season, the time of day, the character's action — to spin one idea into a series. This is how a single concept becomes a posting streak, and it's where scene memory pays off most: the "world" stays consistent while you vary one thing at a time.
The honest note: these look effortless in a highlight reel, but the good ones come from precise direction (more on that below), not one lucky prompt. The upside is that the skill transfers — get good at describing one change cleanly and every format above opens up. Instead of envying the feed, make the first one in the studio.
Because this is a fresh launch, treat the exact ceilings as a snapshot — but here's what's known and how to read it.
One number to keep straight: the $0.10/second figure is API pricing, while $7.99/month is the consumer subscription entry point. They're two different things — don't blend them into one "price."
The competitive question isn't "does Omni have the best frames?" — it's "is conversational editing worth building your workflow around?" Here's the honest landscape. On raw per-frame fidelity, independent testers report placing Seedance 2.0 a notch ahead. Omni's differentiator is the editing model, not out-of-the-box image quality.
| Model | Conversational editing | Raw frame quality | Clip length | Editing model |
|---|---|---|---|---|
| Gemini Omni Flash | Yes — scene-memory, talk-to-edit | Strong (testers place it just behind Seedance) | Up to 10s (deployment choice) | Iterative, plain-language edits on a persistent scene |
| Seedance 2.0 | No native talk-to-edit | Top-rated by independent testers | Short-form range | Regenerate / re-prompt |
| Kling 3.0 | No native talk-to-edit | Strong motion & physical realism | Short-form range | Regenerate / re-prompt |
| Veo 3.1 | Limited (prompt-driven) | Benchmark cinematic fidelity + native audio | Short-form range | Prompt-driven, mostly regenerate |
The fair reading: if you want the single sharpest frame, Seedance 2.0 is the one testers point to; if you want cinematic control and native audio, Veo 3.1 remains a top pick; if motion realism is the priority, Kling 3.0 is strong. Gemini Omni's lane is refinement — getting a specific result through conversation instead of re-rolling. For a fuller side-by-side of the current frontier, see our text-to-video AI model comparison for 2026. And if audio is your priority, the single-pass sound-on approach in our Grok Imagine Video 1.5 review is a useful contrast — different bet, different strength.
The editing-first design points cleanly at some jobs and away from others.
Great fits:
Poor fits:
A responsible-use note that applies to avatars especially: because you can direct a synthetic person precisely, use it for your own brand, licensed material, and — if a real person's likeness or voice is involved — only with recorded consent. Don't put words or actions onto real people who didn't agree to it, and disclose AI-generated content where your platform or audience expects it. (The built-in SynthID and C2PA markers help here, but they don't replace getting permission.)
Gemini Omni is impressive, but it's early, and a few sharp edges are worth naming — this is the part the highlight reels leave out.
Alongside Omni, Google also shipped Nano Banana 2 Lite — the fastest, cheapest model in the Nano Banana image line, with roughly 4-second text-to-image generation at about $0.034 per 1,000 images. It's a different tool for a different job (images, not video), but worth a mention: Oxava already uses Nano Banana image models, so the family that powers a lot of on-platform image work just got a faster, lower-cost tier.
Here's the practical part: you don't have to wait on a staggered API rollout — or keep watching other people's clips go viral — to make your own. Gemini Omni is live now in Oxava's studio, including image-to-video and the talk-to-edit workflow. Start from a still or generate a clip, then refine it conversationally: change the lighting, restage a character, add an element, tighten the motion — one plain-language instruction at a time, on a scene the model remembers. That's the exact loop behind the product ads, physics clips, and remixes filling your feed.
It fits naturally into an image-first pipeline. Shape your hero visual, animate it, then converse your way to the final shot — the image-to-video workflow guide walks through the front half of exactly that flow. And because Oxava is multi-model, you're not betting your whole pipeline on one release: pick Omni when editability is the point, reach for a different model when raw fidelity or audio matters more, and keep it all in one place.
To be clear about what Omni does and doesn't do: today it generates video and edits it by conversation. It does not yet output images or audio — those are on Google's roadmap, not in the product now. What's live and genuinely useful right now is the editing loop, and it's the fastest way to feel why "talk to your video" is more than a demo trick.
The takeaway from this Gemini Omni review: stop scrolling past the clips you wish you'd made. Open the studio and edit your first one by talking to it.
Gemini Omni is Google's "any-to-any" multimodal AI model, announced at Google I/O on May 19, 2026. It takes text, images, a voice/audio reference, and video as input and currently produces video output. Its standout feature is conversational, scene-memory editing — you refine a clip by talking to it. The first model in the family is Gemini Omni Flash.
Not yet. At launch, Gemini Omni outputs video only. Google has said image and audio outputs are on the roadmap, but they are not available in the product today. Despite the multi-input "omni" name, treat it as a video model for now.
Two separate prices apply. The Gemini Omni Flash API runs about $0.10 per second of generated video (roughly $1 for a 10-second clip). Separately, consumer access starts around $7.99/month via Google's entry AI plan, with free access surfacing in YouTube Shorts/Create. Don't conflate the per-second API rate with the monthly subscription.
Currently up to 10 seconds. Google describes this as a deployment choice rather than a model limitation and expects the ceiling to increase over time. For now, it's best treated as a shot-by-shot tool rather than a long-form generator.
It depends on the job. On raw frame quality, independent testers place Seedance 2.0 slightly ahead, and Veo 3.1 remains the benchmark for cinematic fidelity and native audio. Gemini Omni's advantage is conversational, scene-memory editing — converging on a specific result without regenerating. Match the model to the shot rather than looking for one overall winner.
Be the first to hear about new techniques, model updates and ideas on AI generation.