{"slug": "google-s-gemini-3-7-flash-beats-rivals-on-agent-benchmarks-at-half-the-price", "title": "Google's Gemini 3.7 Flash Beats Rivals on Agent Benchmarks at Half the Price", "summary": "Google released Gemini 3.7 Flash on August 13, 2026, topping key agent benchmarks at half the price of its predecessor, with input tokens at $0.75 per million and output at $3.75 per million through the end of 2026. The model scores 43.6% on FrontierCode 1.1, edging out Claude Sonnet 5's 42.7% and GPT-5.6 Terra's 41.3%, and 30.4% on AutomationBench, nearly triple Claude Sonnet 5's 10.7%. The launch comes as Anthropic, which reversed a planned September 1 price hike on August 11, faced ten outage incidents between August 12 and 19, part of 164 documented incidents since January.", "body_md": "*Google's new Gemini 3.7 Flash tops key agent benchmarks and costs half what its predecessor charged, landing right as Anthropic fights outages and reverses a planned price hike.*\n\nGoogle shipped Gemini 3.7 Flash on August 13, 2026, just three weeks after Gemini 3.6 Flash. The pitch: \"our most intelligent workhorse model yet for coding and agents.\" Bold words. This time, though, they're backed up. On FrontierCode 1.1 the model hits 43.6%, edging out Claude Sonnet 5 at 42.7% and GPT-5.6 Terra at 41.3%. On AutomationBench, which measures how well a model handles real enterprise workflow automation, it scores 30.4% - nearly triple Claude Sonnet 5's 10.7%.\n\nPrice is the real story here. Google is charging $0.75 per million input tokens and $3.75 per million output tokens through the end of 2026: half of what 3.6 Flash cost at launch. That intro rate doubles to $1.50 and $7.50 on January 1, 2027. But for now, if you're running an agent that fires off thousands of tool calls a day, that gap adds up fast.\n\n## Anthropic is on the back foot\n\nYou don't have to squint to see why the timing matters. Anthropic made its $2 input, $10 output pricing for Sonnet 5 permanent on August 11, according to reporting tracked by explainx.ai, scrapping a planned September 1 hike to $3 and $15. That's not a company raising prices from strength. It's a company backing off one, days before a competitor undercut it anyway.\n\nAnthropic has bigger problems than pricing, too. Claude logged ten separate outage incidents between August 12 and 19 alone, according to Tech Times, part of a run of 164 documented incidents since January. One evening outage in early August knocked out OAuth authentication along with model access. Paying customers couldn't even log in. Most enterprise SaaS contracts demand 99.9% uptime. Claude's real-world numbers work out to something closer to 14 to 23 hours of downtime per service each quarter.\n\n[How One Judge's Split Ruling on Anthropic Became AI's Copyright Rulebook](https://startupfortune.com/how-one-judges-split-ruling-on-anthropic-became-ais-copyright-rulebook/)\n\nJudge Alsup ruled that training AI on books is fair use, but that Anthropic's piracy of those books from shadow libraries was not, a split that cost the company $1.5 billion. Publishers are now using the same legal seam to sue Google over how Gemini was trained. - [how judges ruled on AI training copyrighted books](https://startupfortune.com/how-one-judges-split-ruling-on-anthropic-became-ais-copyright-rulebook/) - [why anthropic paid 1.5 billion for copyright infringement](https://startupfortune.com/how-one-judges-split-ruling-on-anthropic-became-ais-copyright-rulebook/)\n\n## Not a clean sweep\n\nFrankly, the picture isn't a clean sweep for Google. On DeepSWE v1.1, a benchmark for autonomous software engineering tasks, Gemini 3.7 Flash scores 65.3%, up sharply from 3.6 Flash's 49.0%. GPT-5.6 Terra still leads there at 69.6%. On Agent's Last Exam, which tests multimodal desktop and operating-system tasks, Claude Sonnet 5 comes out on top at 33.3% against Gemini 3.7 Flash's 26.3%. So this isn't a model that beats every rival on every measure. It's a model that's good enough on most of them, and cheap enough that the gap stops mattering for a lot of buyers.\n\nThat's the calculation enterprises are actually making right now. Nobody's asking which model wins gold across the board. They're asking which model gets the agent pipeline built without blowing the inference budget or going down mid-shift. A 30.4% AutomationBench score paired with sub-dollar input pricing solves a real procurement problem, even if Claude still wins the occasional headline benchmark.\n\nGoogle isn't just cutting prices to look generous, either. Gemini 3.6 Flash arrived only three weeks before 3.7 Flash. That tells you Google is iterating fast enough to keep resetting the price-performance bar before competitors catch up. Gemini 3.5 Pro, meanwhile, is still stuck in partner testing with no release date. Google is clearly leaning on Flash as its workhorse tier while the bigger model stays in the oven.\n\n## What this means for developers\n\nFor developers building agent stacks, the math is straightforward. If your workload runs thousands of small, repeated calls - coding agents, document parsing, automation pipelines - the per-token cost dominates far more than any single benchmark point. Gemini 3.7 Flash's $0.75/$3.75 pricing, paired with a genuine lead on enterprise-automation and coding benchmarks, makes it the obvious default to test first. Claude Sonnet 5 still makes sense if your workload leans on desktop or OS-level agent tasks, where it keeps its edge. But reliability has to factor in too. Right now Google's model isn't the one logging outages every few days.\n\nNone of this is permanent. Prices reset in January. Anthropic will keep patching its infrastructure, and Gemini 3.5 Pro is still coming. But for the next four months, the cheapest fast option on the market also happens to be one of the best performing ones. That combination doesn't come around often. Enterprises shopping for agent infrastructure right now know it.\n\n**Also read:** [Taiwan Indicts Nine, Including an Nvidia Employee, Over AI Server Smuggling](https://startupfortune.com/taiwan-indicts-nine-including-an-nvidia-employee-over-ai-server-smuggling/) • [Nvidia Is in Talks to Back Perplexity AI at a $30 Billion Valuation](https://startupfortune.com/nvidia-is-in-talks-to-back-perplexity-ai-at-a-30-billion-valuation/) • [d-Matrix's Raptor Chip Aims to Break AI's Memory Wall at Hot Chips 2026](https://startupfortune.com/d-matrixs-raptor-chip-aims-to-break-ais-memory-wall-at-hot-chips-2026/)", "url": "https://wpnews.pro/news/google-s-gemini-3-7-flash-beats-rivals-on-agent-benchmarks-at-half-the-price", "canonical_source": "https://startupfortune.com/googles-gemini-37-flash-beats-rivals-on-agent-benchmarks-at-half-the-price/", "published_at": "2026-08-24 09:30:20+00:00", "updated_at": "2026-08-24 09:42:59.597085+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-products", "ai-policy"], "entities": ["Google", "Gemini 3.7 Flash", "Anthropic", "Claude Sonnet 5", "GPT-5.6 Terra", "FrontierCode 1.1", "AutomationBench", "DeepSWE v1.1"], "alternates": {"html": "https://wpnews.pro/news/google-s-gemini-3-7-flash-beats-rivals-on-agent-benchmarks-at-half-the-price", "markdown": "https://wpnews.pro/news/google-s-gemini-3-7-flash-beats-rivals-on-agent-benchmarks-at-half-the-price.md", "text": "https://wpnews.pro/news/google-s-gemini-3-7-flash-beats-rivals-on-agent-benchmarks-at-half-the-price.txt", "jsonld": "https://wpnews.pro/news/google-s-gemini-3-7-flash-beats-rivals-on-agent-benchmarks-at-half-the-price.jsonld"}}