VUSION Magazine · Tools

Gemini Omni Flash Review: Google's Conversational Video AI

Google's Gemini Omni Flash makes 720p video with synced audio and multi-turn conversational editing — for iterate-and-refine drafts, not 4K finals.

By Roy Chucri · 2026-07-11

Gemini Omni Flash Review: Google's Conversational Video AI
Released: June 30, 2026 · Model: Gemini Omni Flash (gemini-omni-flash-preview) · Reviewed: July 1, 2026

Verdict

Gemini Omni Flash is worth reaching for when video is a conversation, generate a 720p clip with synchronized audio, then refine it across turns ("make it rain," "add warm light," "drop in the logo") without losing character, audio, or camera continuity; for a 1080p/4K final, hand the winner to Veo 3.1.

Best for: iterate-and-refine drafting, storyboards, social clips, conversational multi-turn edits with audio. Skip it if: you need 1080p or 4K, clips longer than 10 seconds, or a single-shot hero render. Released: June 30, 2026 (developer preview), reviewed July 1, 2026.

By the numbers

10s, maximum clip length at launch (longer durations promised later). 720p, output resolution cap, fine for social and iteration, not for large-screen finals. Batch runs, about half the cost of interactive generation for the same output. 1-pass audio, every clip is generated with synchronized native audio, not bolted on after. Multi-turn, conversational edits build on each other, preserving character, audio, and camera continuity.

What it is

Gemini Omni Flash is Google's multimodal video model: it generates and edits video from combinations of text, image, and video, and refines the result through natural-language conversation. The distinguishing move is conversational editing, after a first clip, you issue instructions that build on each other ("change the background to a rain-soaked Tokyo street," then "add warm amber lighting," then "add a glowing logo"), each applied to the current state without losing continuity. It is aimed at iterative production, not one-shot novelty clips.

Where it fits

Omni Flash is positioned as a complement to Veo 3.1, not a replacement. Veo handles single-shot, high-resolution generation up to 4K; Omni handles the messy middle, five or six rounds of edits before a clip is locked. The intended workflow is explicit: draft and iterate in Omni Flash, then pass the best candidate to Veo 3.1 for the final high-resolution render. Route by the real fork, resolution and workflow. If your process is iterate-and-refine and 720p is enough, Omni; if you need native 1080p/4K for a large screen, Veo.

Conversational editing: what's new

Most video models make you re-roll the whole clip to change one thing. Omni Flash keeps state: each instruction is applied to the current video, so continuity of character, audio, and camera survives across turns. That turns video generation into a directed conversation rather than a slot-machine pull, the reason it slots so cleanly into a storyboard-then-refine pipeline. Native synchronized audio in the same pass means you are not re-timing sound after every visual change.

Gemini Omni Flash vs Veo 3.1

MetricGemini Omni FlashVeo 3.1
Primary roleIterate-and-refine draftingSingle-shot final render
Max resolution720pUp to 4K
Clip lengthUp to 10s~8s
Native audioYes (1-pass, synchronized)Yes
Conversational multi-turn editingYesNo
InputsText, image, videoText, image
Cost (video output)Same per-second rate as Veo 3.1; batch about halfSame per-second rate as Omni Flash
Best useStoryboards, drafts, socialLarge-screen / premium finals

Limits, do NOT

FAQ

What is Gemini Omni Flash?

Google's multimodal video model that generates 720p clips with audio and refines them through conversational, multi-turn editing.

How long and what resolution are the clips?

Up to 10 seconds at 720p at launch.

Does it generate audio?

Yes, every clip is generated with synchronized native audio in one pass.

Omni Flash or Veo 3.1?

Omni for iterate-and-refine 720p drafting; Veo 3.1 for single-shot 1080p/4K finals.

Where can I run it?

In Google AI Studio and through the Gemini API, where it has been in public preview since 30 June 2026. Google publishes the current per-second rate on its own pricing page.

Disclosure: VUSION Magazine is published by Vusion.art, which also builds an AI film production engine of the same name, currently in private beta.