cd /news/generative-ai/google-ai-releases-gemini-omni-1-1-f… · home topics generative-ai article
[ARTICLE · art-115158] src=marktechpost.com ↗ pub= topic=generative-ai verified=true sentiment=· neutral

Google AI Releases Gemini Omni 1.1 Flash: 40-Second Scene Extension, First/Last Frame Control, and 4K Upscaling

Google released Gemini Omni 1.1 Flash, a production update to its native multimodal video generation and editing model, enabling scene extension up to 40 seconds, first/last frame control, and 4K upscaling. The model, available via the Gemini API, Google AI Studio, and the Gemini Enterprise Agent Platform, is already used by Adobe, Figma Weave, GMI Cloud, and Runway. Pricing is $1.50 per 1M input tokens, $9.00 per 1M text output tokens, and $17.50 per 1M video tokens, with SynthID watermarking on all outputs.

read4 min views1 publishedAug 29, 2026
Google AI Releases Gemini Omni 1.1 Flash: 40-Second Scene Extension, First/Last Frame Control, and 4K Upscaling
Image: MarkTechPost

Google has released Gemini Omni 1.1 Flash (gemini-omni-1.1-flash ), a production update to its native multimodal video generation and editing model. The release moves Omni from a capable generator to a directable one: scene extension now reads up to 10 seconds of prior context instead of a single final frame, first and last frames can be pinned to control camera movement, drafts render in 360p at a third of 720p cost, finals upscale to 4K, and video clips can be passed as references for character consistency.

Gemini Omni Flash is built on three properties Google distinguishes from prior video models: native multimodality (text, image, audio, and video processed together), conversational editing through the Interactions API, and world knowledge inherited from Gemini. Editing is stateful — you pass previous_interaction_id

and the model applies your change while preserving what you did not mention, without re-up the prior video.

Is it deployable?

It is available through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform, with Adobe, Figma Weave, GMI Cloud, and Runway already named as production users.

Scene Extension is the Main Change

Omni 1.1 analyzes up to 10 seconds of prior context when continuing a clip. Google states that previous models referenced only the final second. Extensions run in 10-second increments to a cumulative 40 seconds, and the model generates a 3–10 second continuation per call. Some final frames of the input are edited to make the seam continuous.

The constraints are specific. Extension appends to the end of a clip only — no prepending, no mid-clip insertion. Uploaded input videos must be 10 seconds or shorter, unless you are extending a model-generated video in multi-turn. You cannot add new dialogue when extending an uploaded video where someone is speaking; spoken dialogue is supported in multi-turn extension via previous_interaction_id

.

Keyframes and Video References

You can now supply a first and last frame and have the model generate the continuous video between them, which is the mechanism behind orbits, dolly-zooms, and seamless loops. Prompts bind media to roles with tags: <FIRST_FRAME>

, <LAST_FRAME>

, <IMAGE_REF_N>

, and <VIDEO_REF_N>

.

Video references accept a maximum of three clips, up to three seconds each, and work best for likenesses. Audio inside a video reference is ignored. Reasoning across multiple videos is not supported and may degrade output.

Cost Control: Draft in 360p, Ship in 4K

The resolution

parameter in response_format

takes 360p

, 720p

(default), 1080p , and 4k

, with the top two upscaled. Google reports 360p previews generate up to 60% faster and at a third of the cost of 720p, based on system throughput of 360p versus 720p. That makes the draft-then-upscale loop the intended production pattern: iterate cheaply, render once.

Pricing, provenance, and limits

Input is $1.50 per 1M tokens (text, image, video, audio). Output is $9.00 per 1M text tokens and $17.50 per 1M video tokens. Video billing runs at 5,792 tokens per second of 720p, an effective ~$0.10 per second under standard pricing.

Every generated video carries SynthID watermarking — invisible to viewers, programmatically detectable for provenance. Notable gaps: no system instructions, temperature, top_p

, stop sequences, or negative prompts (negatives go in the prompt text); voice editing is unsupported; audio references are unsupported; YouTube URLs cannot be used as a source. English is fully supported; other languages are unevaluated. For outputs above 4MB, use delivery="uri"

and poll the Files API until the file is ACTIVE .

Google names Adobe (Firefly), Figma Weave, GMI Cloud, and Runway as customers already running Omni Flash in production. The model is also live in Google Flow for AI Plus, Pro, and Ultra subscribers, with scene extension in the Gemini app.

Comparison

Key Takeaways

  • Scene extension now reads 10s of prior context, up from one final second, and stacks to 40s total.
  • First/last frame interpolation plus <VIDEO_REF_N>

tags give shot-level camera and character control. - 360p drafts run up to 60% faster at a third of 720p cost; 1080p and 4K are upscaled outputs.

  • Paid tier only — ~$0.10 per second of 720p video, no free tier, no provisioned throughput.
Check out the [ Google blog announcement](https://blog.google/innovation-and-ai/technology/developers-tools/build-with-gemini-omni-1-1-flash/),

[,](https://ai.google.dev/gemini-api/docs/omni)

**Gemini API Omni documentation**[, and](https://ai.google.dev/gemini-api/docs/pricing)

Gemini API pricing. Also, feel free to follow us on

Omni quickstart cookbook****and don’t forget to join ourTwitter

and Subscribe to

[150k+ML SubReddit](https://www.reddit.com/r/machinelearningnews/)**. Wait! are you on telegram?**

[our Newsletter](https://magic.beehiiv.com/v1/f5e63dd4-5653-4f09-83e2-321a8b1ba526?email={{email}})

now you can join us on telegram as well.Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us

Michal Sutter is a data science professional with a Master of Science in Data Science from the University of Padova. With a solid foundation in statistical analysis, machine learning, and data engineering, Michal excels at transforming complex datasets into actionable insights.

── more in #generative-ai 4 stories · sorted by recency
── more on @google 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/google-ai-releases-g…] indexed:0 read:4min 2026-08-29 ·