Google released Gemini 3.7 Flash on August 13, 2026, a multimodal model for coding and agent workflows. Google reports higher results than Gemini 3.6 Flash on its cited software engineering, web development, document-processing, and workflow benchmarks. The model supports a 1 million-token context window and is available through Google's developer and enterprise channels, with introductory API pricing through December 31, 2026.
Google released Gemini 3.7 Flash on August 13, 2026, an update to its Flash model line aimed at coding and agent workflows. Google's launch post and Ars Technica place the release three weeks after Gemini 3.6 Flash.
Google describes 3.7 Flash as a model for coding, agents, knowledge work, and web development. The company positions it as a high-efficiency model for multi-step orchestration and software work.
Model interface and deployment
Google's model card lists text, image, audio, and video inputs, text output, a 1,048,576-token context window, and a 65,536-token maximum output. The developer documentation identifies the stable model ID as gemini-3.7-flash.
The model supports configurable reasoning through LOW, MEDIUM, and HIGH thinking levels, with MEDIUM as the default. Google's enterprise documentation notes that MINIMAL is unsupported and returns an API validation error when explicitly set.
Google distributes the model through the Gemini API, AI Studio, Android Studio, Gemini Enterprise Agent Platform, and Google Antigravity. Ars Technica reports that it also powers the Gemini Spark agent for Google AI Pro and Ultra subscribers, while the regular chatbot interface remained on 3.6 Flash at launch.
Google's reported benchmark gains
Google reported improvements over Gemini 3.6 Flash across several evaluations:
- • FrontierCode 1.1 Main: 43.6% for 3.7 Flash, versus 34.4% for 3.6 Flash. - • DeepSWE v1.1: 65.3% for 3.7 Flash. Google's launch post and Ars Technica list 49.0% for 3.6 Flash, while Google DeepMind's model-card table lists 48.6%. - • WebDev Arena: Elo score of 1,588, versus 1,538. - • GDP.pdf: 34.0%, versus 22.0%. - • AutomationBench: 30.4%, versus 17.0%.
These are vendor-reported results rather than independent comparisons. Ars Technica confirmed the figures published in Google's launch post while questioning whether the differences warranted another model release after only three weeks. The 0.4-point DeepSWE baseline difference between Google's launch post and model card remains a source-level inconsistency, not an independently resolved measurement.
For ML engineers, the combination of long context, multimodal inputs, terminal-oriented coding claims, and adjustable thinking levels is relevant to agent evaluation and inference-cost design. Benchmark gains do not by themselves establish reliability in tool use, permission handling, recovery from failed steps, or task-specific code review. Teams evaluating the model will need workload-level tests that measure those properties alongside token spend and latency.
Pricing and release context
Google is offering introductory pricing of $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026. Google states that pricing changes on January 1, 2027, to $1.50 per million input tokens and $7.50 per million output tokens.
Ars Technica noted that 3.7 Flash arrived before Google's long-awaited Gemini 3.5 Pro. Google did not announce a release date for that Pro model in the 3.7 Flash launch materials.
Key Points #
- 1Gemini 3.7 Flash combines multimodal input, a 1 million-token context window, and configurable thinking levels for agent and coding experiments.
- 2Benchmark gains do not by themselves establish reliability in tool use, permissions handling, recovery from failed steps, or task-specific code review.
- 3Introductory token pricing is half the original Gemini 3.6 Flash rate, making cost-per-task comparisons important for teams evaluating agent workloads.
Scoring Rationale #
Gemini 3.7 Flash is a significant new model release from a frontier AI provider, with direct relevance to coding agents, multimodal processing, and long-context applications. Its pricing and reported benchmark gains may inform model selection for production teams, though the available performance evidence is primarily vendor-reported.
Sources #
Primary source and supporting public references used for this report.
Practice with real Ad Tech data
90 SQL & Python problems · 15 industry datasets
[Active Search Campaigns by BudgetEasy](/problems/sql/active-search-campaigns-by-budget)
[High CPC Clicks & Poor Landing PagesMedium](/problems/sql/high-cpc-clicks-poor-landing-page)
[Campaign ROAS by Attribution ModelHard](/problems/sql/campaign-roas-by-attribution-model)
250 free problems · No credit card