cd /news/artificial-intelligence/gemini-3-7-flash-beats-claude-sonnet… · home topics artificial-intelligence article
[ARTICLE · art-98880] src=byteiota.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Gemini 3.7 Flash Beats Claude Sonnet 5 in Coding at Half the Price

Google shipped Gemini 3.7 Flash on August 13, topping FrontierCode 1.1 and DeepSWE v1.1 coding benchmarks with scores of 43.6% and 65.3%, respectively, and priced at $0.75 per million input tokens—four times cheaper than Claude Sonnet 5 and nearly three times cheaper than GPT-5.6 Terra. The model also powers Google Antigravity's default agent platform, with AutomationBench scores jumping from 17.0% to 30.4%, but introductory pricing expires December 31, 2026, doubling to $1.50 per million tokens.

read3 min views1 publishedAug 16, 2026
Gemini 3.7 Flash Beats Claude Sonnet 5 in Coding at Half the Price
Image: Byteiota (auto-discovered)

Google shipped Gemini 3.7 Flash on August 13, and it immediately topped the two coding benchmarks that developers actually argue about — FrontierCode 1.1 and DeepSWE v1.1 — edging out Claude Sonnet 5 and GPT-5.6 Terra in both. The launch price is $0.75 per million input tokens. That is four times cheaper than Sonnet 5, nearly three times cheaper than GPT-5.6 Terra, and the biggest pricing disruption in the AI coding tool space since GPT-5.6 dropped six weeks ago.

The Benchmark Numbers #

Here is what Google published alongside the competition:

Model FrontierCode 1.1 DeepSWE v1.1 Code Arena Elo Input $/1M
Gemini 3.7 Flash 43.6% 65.3% 1588 $0.75*
Claude Sonnet 5 42.7% 49.7% 1541 $3.00
GPT-5.6 Terra 41.3% ~52% 1523 $2.00

The FrontierCode gap is thin — 43.6% vs 42.7% is not something to bet a migration on. But the DeepSWE v1.1 lead is harder to ignore: 65.3% versus 49.7% for Sonnet 5 is a 15-point gap on long-horizon software engineering tasks, the kind where an agent has to write code, run tests, read output, and fix failures across multiple steps. That is not a rounding error. For context, Gemini 3.6 Flash scored 48.6% on the same benchmark — 3.7 Flash gained 16.7 percentage points in a single generation.

What the Price Gap Means at Scale #

Teams paying attention to inference costs are already doing the math. At 1 billion tokens of input per month — a reasonable figure for a team running continuous coding agents — Gemini 3.7 Flash costs $750. Sonnet 5 costs $3,000. GPT-5.6 Terra costs $2,000. If the performance difference between these models for your specific workload is marginal, those numbers make the decision for you.

The catch: introductory pricing expires December 31, 2026. From January 1, 2027, the rate doubles to $1.50/$7.50 per million tokens. That is still cheaper than both competitors at standard pricing, but teams that plan budgets quarterly need to account for the change now, not in December.

Antigravity Users Get It Automatically #

Google Antigravity — the company’s agent-first development platform — updated 3.7 Flash as its default model the day of launch. If you are already running Antigravity on desktop, CLI, or through the managed agents SDK, you are already on 3.7 Flash. No config change required.

The practical difference: Antigravity agents built on 3.7 Flash think more carefully through multi-step plans before executing tool calls, which translates to fewer mid-task derailments on complex coding tasks. The AutomationBench score jumped from 17.0% to 30.4% — a near-doubling that suggests something real is happening there, not just benchmark theater.

Where to Be Careful #

A few things to know before you reroute production traffic.

  • These are Google’s benchmarks on Google’s evals. The DeepSWE lead is not unanimous — other sources show GPT-5.6 Terra ahead on the same benchmark depending on methodology. Independent evaluation is still catching up.
  • API documentation for 3.7 Flash is not fully complete. Google flagged this explicitly. It is not a safe blind replacement for 3.6 Flash in a production pipeline where you depend on specific API behavior.
  • GPT-5.6 Terra still leads on Terminal-bench 2.1 and OSWorld agentic evals. Gemini 3.7 Flash is not the top model everywhere.

The honest read: this is a strong release for coding and agent workloads at a price that makes evaluation essentially free. The smart move is a feature-flagged experiment, not a production flip.

How to Try It #

Gemini 3.7 Flash is available now through the Gemini API, AI Studio, Google Antigravity, OpenRouter, and Cursor. The model string is gemini-3.7-flash

. Context window is 1M tokens, max output is 64k. Thinking levels are tunable (low, medium, high) — worth experimenting with on long-horizon agent tasks.

Google’s official announcement has the full benchmark breakdown. The Decoder’s analysis is worth reading for the more skeptical take on whether the benchmark lead holds.

The benchmark lead may narrow once independent testing catches up. The pricing advantage is real right now. That combination is worth 30 minutes of your time to evaluate.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @google 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/gemini-3-7-flash-bea…] indexed:0 read:3min 2026-08-16 ·