# Google's Gemini 3.7 Flash Beats Rivals on Agent Benchmarks at Half the Price

> Source: <https://startupfortune.com/googles-gemini-37-flash-beats-rivals-on-agent-benchmarks-at-half-the-price/>
> Published: 2026-08-24 09:30:20+00:00

*Google's new Gemini 3.7 Flash tops key agent benchmarks and costs half what its predecessor charged, landing right as Anthropic fights outages and reverses a planned price hike.*

Google shipped Gemini 3.7 Flash on August 13, 2026, just three weeks after Gemini 3.6 Flash. The pitch: "our most intelligent workhorse model yet for coding and agents." Bold words. This time, though, they're backed up. On FrontierCode 1.1 the model hits 43.6%, edging out Claude Sonnet 5 at 42.7% and GPT-5.6 Terra at 41.3%. On AutomationBench, which measures how well a model handles real enterprise workflow automation, it scores 30.4% - nearly triple Claude Sonnet 5's 10.7%.

Price is the real story here. Google is charging $0.75 per million input tokens and $3.75 per million output tokens through the end of 2026: half of what 3.6 Flash cost at launch. That intro rate doubles to $1.50 and $7.50 on January 1, 2027. But for now, if you're running an agent that fires off thousands of tool calls a day, that gap adds up fast.

## Anthropic is on the back foot

You don't have to squint to see why the timing matters. Anthropic made its $2 input, $10 output pricing for Sonnet 5 permanent on August 11, according to reporting tracked by explainx.ai, scrapping a planned September 1 hike to $3 and $15. That's not a company raising prices from strength. It's a company backing off one, days before a competitor undercut it anyway.

Anthropic has bigger problems than pricing, too. Claude logged ten separate outage incidents between August 12 and 19 alone, according to Tech Times, part of a run of 164 documented incidents since January. One evening outage in early August knocked out OAuth authentication along with model access. Paying customers couldn't even log in. Most enterprise SaaS contracts demand 99.9% uptime. Claude's real-world numbers work out to something closer to 14 to 23 hours of downtime per service each quarter.

[How One Judge's Split Ruling on Anthropic Became AI's Copyright Rulebook](https://startupfortune.com/how-one-judges-split-ruling-on-anthropic-became-ais-copyright-rulebook/)

Judge Alsup ruled that training AI on books is fair use, but that Anthropic's piracy of those books from shadow libraries was not, a split that cost the company $1.5 billion. Publishers are now using the same legal seam to sue Google over how Gemini was trained. - [how judges ruled on AI training copyrighted books](https://startupfortune.com/how-one-judges-split-ruling-on-anthropic-became-ais-copyright-rulebook/) - [why anthropic paid 1.5 billion for copyright infringement](https://startupfortune.com/how-one-judges-split-ruling-on-anthropic-became-ais-copyright-rulebook/)

## Not a clean sweep

Frankly, the picture isn't a clean sweep for Google. On DeepSWE v1.1, a benchmark for autonomous software engineering tasks, Gemini 3.7 Flash scores 65.3%, up sharply from 3.6 Flash's 49.0%. GPT-5.6 Terra still leads there at 69.6%. On Agent's Last Exam, which tests multimodal desktop and operating-system tasks, Claude Sonnet 5 comes out on top at 33.3% against Gemini 3.7 Flash's 26.3%. So this isn't a model that beats every rival on every measure. It's a model that's good enough on most of them, and cheap enough that the gap stops mattering for a lot of buyers.

That's the calculation enterprises are actually making right now. Nobody's asking which model wins gold across the board. They're asking which model gets the agent pipeline built without blowing the inference budget or going down mid-shift. A 30.4% AutomationBench score paired with sub-dollar input pricing solves a real procurement problem, even if Claude still wins the occasional headline benchmark.

Google isn't just cutting prices to look generous, either. Gemini 3.6 Flash arrived only three weeks before 3.7 Flash. That tells you Google is iterating fast enough to keep resetting the price-performance bar before competitors catch up. Gemini 3.5 Pro, meanwhile, is still stuck in partner testing with no release date. Google is clearly leaning on Flash as its workhorse tier while the bigger model stays in the oven.

## What this means for developers

For developers building agent stacks, the math is straightforward. If your workload runs thousands of small, repeated calls - coding agents, document parsing, automation pipelines - the per-token cost dominates far more than any single benchmark point. Gemini 3.7 Flash's $0.75/$3.75 pricing, paired with a genuine lead on enterprise-automation and coding benchmarks, makes it the obvious default to test first. Claude Sonnet 5 still makes sense if your workload leans on desktop or OS-level agent tasks, where it keeps its edge. But reliability has to factor in too. Right now Google's model isn't the one logging outages every few days.

None of this is permanent. Prices reset in January. Anthropic will keep patching its infrastructure, and Gemini 3.5 Pro is still coming. But for the next four months, the cheapest fast option on the market also happens to be one of the best performing ones. That combination doesn't come around often. Enterprises shopping for agent infrastructure right now know it.

**Also read:** [Taiwan Indicts Nine, Including an Nvidia Employee, Over AI Server Smuggling](https://startupfortune.com/taiwan-indicts-nine-including-an-nvidia-employee-over-ai-server-smuggling/) • [Nvidia Is in Talks to Back Perplexity AI at a $30 Billion Valuation](https://startupfortune.com/nvidia-is-in-talks-to-back-perplexity-ai-at-a-30-billion-valuation/) • [d-Matrix's Raptor Chip Aims to Break AI's Memory Wall at Hot Chips 2026](https://startupfortune.com/d-matrixs-raptor-chip-aims-to-break-ais-memory-wall-at-hot-chips-2026/)
