Google AI Releases Gemini Omni 1.1 Flash: 40-Second Scene Extension, First/Last Frame Control, and 4K Upscaling Google released Gemini Omni 1.1 Flash, a production update to its native multimodal video generation and editing model, enabling scene extension up to 40 seconds, first/last frame control, and 4K upscaling. The model, available via the Gemini API, Google AI Studio, and the Gemini Enterprise Agent Platform, is already used by Adobe, Figma Weave, GMI Cloud, and Runway. Pricing is $1.50 per 1M input tokens, $9.00 per 1M text output tokens, and $17.50 per 1M video tokens, with SynthID watermarking on all outputs. Google has released Gemini Omni 1.1 Flash gemini-omni-1.1-flash , a production update to its native multimodal video generation and editing model. The release moves Omni from a capable generator to a directable one: scene extension now reads up to 10 seconds of prior context instead of a single final frame, first and last frames can be pinned to control camera movement, drafts render in 360p at a third of 720p cost, finals upscale to 4K, and video clips can be passed as references for character consistency. Gemini Omni Flash is built on three properties Google distinguishes from prior video models: native multimodality text, image, audio, and video processed together , conversational editing through the Interactions API https://ai.google.dev/gemini-api/docs/interactions-overview , and world knowledge inherited from Gemini. Editing is stateful — you pass previous interaction id and the model applies your change while preserving what you did not mention, without re-uploading the prior video. Is it deployable? It is available through the Gemini API https://ai.google.dev/gemini-api/docs/omni in Google AI Studio https://aistudio.google.com/prompts/new chat?model=gemini-omni-1.1-flash and the Gemini Enterprise Agent Platform https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/omni-1-1-flash , with Adobe, Figma Weave, GMI Cloud, and Runway already named as production users. Scene Extension is the Main Change Omni 1.1 analyzes up to 10 seconds of prior context when continuing a clip. Google states that previous models referenced only the final second. Extensions run in 10-second increments to a cumulative 40 seconds , and the model generates a 3–10 second continuation per call. Some final frames of the input are edited to make the seam continuous. The constraints are specific. Extension appends to the end of a clip only — no prepending, no mid-clip insertion. Uploaded input videos must be 10 seconds or shorter, unless you are extending a model-generated video in multi-turn. You cannot add new dialogue when extending an uploaded video where someone is speaking; spoken dialogue is supported in multi-turn extension via previous interaction id . Keyframes and Video References You can now supply a first and last frame and have the model generate the continuous video between them, which is the mechanism behind orbits, dolly-zooms, and seamless loops. Prompts bind media to roles with tags: