The technology behind Gemini Omni 1.1 Flash
How the model works, based on public documentation and sourced evidence.
Gemini Omni 1.1 Flash is Google's video generation and editing model, announced on August 27, 2026. One model covers text to video, animating a still image, reference driven scenes, first and last frame control, natural language editing and scene extension. Google's model page lists the id as gemini-omni-1.1-flash, with clips of 3 to 10 seconds at 24 FPS, returned as mp4 with audio. Every clip carries a generated soundtrack of speech, music and sound effects rather than a silent picture you dub afterwards, and 720p is the standard size, with 1080p and 4K listed as upscaled.
Working with it reads more like a conversation than a render queue. Google's docs describe a previous interaction id that carries the conversation history and the generated video state without re-uploading the previous video, so a change is a short follow-up instruction rather than a fresh job. The role of each supplied image or video is named inside the prompt with simple tags for the first frame, the last frame and each reference, not in separate fields. Audio has no on or off switch, because all generated videos include it. The frame is 16:9 or 9:16 only.
Against Google's own Veo 3.1 the split is control against length. Veo 3.1 fixes clips at 4, 6 or 8 seconds, while Omni takes any length from 3 to 10 seconds, accepts video references of up to 3 seconds each, and extends footage in 10 second increments to a cumulative 40 seconds. Billing differs too. Veo 3.1 carries an official API list price of $0.40 for each second of 720p or 1080p video, while Omni bills video output at $17.50 per million tokens, which Google's pricing page puts at 5,792 tokens for one second of 720p, or roughly $0.10.
On Vidney the model runs three of those modes: text to video, image to video from a first frame, and reference to video with up to seven reference images. Clips are 4, 6, 8 or 10 seconds, at 720p or 1080p, in 16:9 or 9:16. The soundtrack is always generated and cannot be switched off here, so plan the shot with sound in mind. Last frame control, video references, 4K, conversational editing and scene extension stay closed on Vidney for now, which is our integration's current boundary rather than a limit of the model.