Released: June 30, 2026 · Model: Gemini Omni Flash (gemini-omni-flash-preview) · Reviewed: July 1, 2026
Verdict
Gemini Omni Flash is worth reaching for when video is a conversation, generate a 720p clip with synchronized audio, then refine it across turns ("make it rain," "add warm light," "drop in the logo") without losing character, audio, or camera continuity; for a 1080p/4K final, hand the winner to Veo 3.1.
Best for: iterate-and-refine drafting, storyboards, social clips, conversational multi-turn edits with audio. Skip it if: you need 1080p or 4K, clips longer than 10 seconds, or a single-shot hero render. Released: June 30, 2026 (developer preview), reviewed July 1, 2026.
By the numbers
10s, maximum clip length at launch (longer durations promised later). 720p, output resolution cap, fine for social and iteration, not for large-screen finals. Batch runs, about half the cost of interactive generation for the same output. 1-pass audio, every clip is generated with synchronized native audio, not bolted on after. Multi-turn, conversational edits build on each other, preserving character, audio, and camera continuity.
What it is
Gemini Omni Flash is Google's multimodal video model: it generates and edits video from combinations of text, image, and video, and refines the result through natural-language conversation. The distinguishing move is conversational editing, after a first clip, you issue instructions that build on each other ("change the background to a rain-soaked Tokyo street," then "add warm amber lighting," then "add a glowing logo"), each applied to the current state without losing continuity. It is aimed at iterative production, not one-shot novelty clips.
Where it fits
Omni Flash is positioned as a complement to Veo 3.1, not a replacement. Veo handles single-shot, high-resolution generation up to 4K; Omni handles the messy middle, five or six rounds of edits before a clip is locked. The intended workflow is explicit: draft and iterate in Omni Flash, then pass the best candidate to Veo 3.1 for the final high-resolution render. Route by the real fork, resolution and workflow. If your process is iterate-and-refine and 720p is enough, Omni; if you need native 1080p/4K for a large screen, Veo.
Conversational editing: what's new
Most video models make you re-roll the whole clip to change one thing. Omni Flash keeps state: each instruction is applied to the current video, so continuity of character, audio, and camera survives across turns. That turns video generation into a directed conversation rather than a slot-machine pull, the reason it slots so cleanly into a storyboard-then-refine pipeline. Native synchronized audio in the same pass means you are not re-timing sound after every visual change.
Gemini Omni Flash vs Veo 3.1
| Metric | Gemini Omni Flash | Veo 3.1 |
|---|---|---|
| Primary role | Iterate-and-refine drafting | Single-shot final render |
| Max resolution | 720p | Up to 4K |
| Clip length | Up to 10s | ~8s |
| Native audio | Yes (1-pass, synchronized) | Yes |
| Conversational multi-turn editing | Yes | No |
| Inputs | Text, image, video | Text, image |
| Cost (video output) | Same per-second rate as Veo 3.1; batch about half | Same per-second rate as Omni Flash |
| Best use | Storyboards, drafts, social | Large-screen / premium finals |
Limits, do NOT
- Do not deliver finals from it. 720p is a hard cap; for large-screen or premium brand work, finish in Veo 3.1 or upscale externally.
- Do not plan long takes. Clips top out at 10 seconds at launch; storyboard in beats, not scenes.
- Do not expect one-shot perfection. Its strength is the conversation, budget several refinement turns rather than a single hero prompt.
- Do not skip a review pass on audio. Synchronized audio is generated, not curated, check it before you ship.
FAQ
What is Gemini Omni Flash?
Google's multimodal video model that generates 720p clips with audio and refines them through conversational, multi-turn editing.
How long and what resolution are the clips?
Up to 10 seconds at 720p at launch.
Does it generate audio?
Yes, every clip is generated with synchronized native audio in one pass.
Omni Flash or Veo 3.1?
Omni for iterate-and-refine 720p drafting; Veo 3.1 for single-shot 1080p/4K finals.
Where can I run it?
In Google AI Studio and through the Gemini API, where it has been in public preview since 30 June 2026. Google publishes the current per-second rate on its own pricing page.
Disclosure: VUSION Magazine is published by Vusion.art, which also builds an AI film production engine of the same name, currently in private beta.