{"slug": "gemini-3-7-flash-vs-sonnet-5-vs-gpt-5-6-terra-real-wins", "title": "Gemini 3.7 Flash vs Sonnet 5 vs GPT-5.6 Terra: Real Wins", "summary": "Google's Gemini 3.7 Flash, launched August 13, 2026 at $0.75 per million input tokens and $3.75 per million output, wins 9 of 19 benchmark rows in its own model card, while GPT-5.6 Terra wins 6 and Claude Sonnet 5 wins 2 outright. Independent Artificial Analysis scores 3.7 Flash at 56 versus 52 for 3.6 Flash, a +4 gain. The card includes a blank Sonnet 5 cell for OSWorld-2.0, which Google says is a version mismatch with Anthropic's OSWorld-Verified, where Sonnet 5 scores 81.2%.", "body_md": "Gemini 3.7 Flash benchmarks look very different depending on which table you read. Google’s launch post led with five rows; the full model card publishes nineteen, compared across five models — Gemini 3.7 Flash, Gemini 3.6 Flash, Claude Sonnet 5, GPT-5.6 Terra, and Muse Spark 1.2. This post reads all nineteen.\n\nThe stakes are practical. Gemini 3.7 Flash shipped on August 13, 2026 at an introductory $0.75 per million input tokens and $3.75 per million output on Google’s API — a rate that undercuts every rival in its own comparison table, matched only by 3.6 Flash on the same introductory window. If the benchmark story holds, that price rewrites routing decisions. If it only holds on the five rows the vendor highlighted, it doesn’t. We covered the launch itself — pricing mechanics, positioning, and the asterisk on “half price” — in [the full launch rundown](/blog/gemini-3-7-flash-launch-half-price-workhorse-2026); this post is the benchmark deep-dive.\n\nWhat follows: the full 19-row table transcribed from the model card, a recomputed win tally, the rows where Sonnet 5 and even 3.6 Flash quietly beat the new model, the blank cell that most coverage repeated without investigating, the independent Artificial Analysis numbers that actually moved this time, and every price labelled by the surface it comes from. Every benchmark figure below is vendor-stated from Google’s card unless noted otherwise.\n\n- 01The real table is 19 rows and five models.Most launch coverage repeated the same five benchmarks Google led with. The model card publishes 19 rows — including rows Google loses, a fifth model column (Muse Spark 1.2), one blank cell, and one internal inconsistency.\n- 02It’s a split decision, not a sweep.By our tally of Google’s own table: 3.7 Flash takes 9 rows, GPT-5.6 Terra 6, Sonnet 5 2 outright (3 against the headline trio), 3.6 Flash 1, and Muse Spark 1.2 tops GDPVal-AA outright.\n- 03The independent signal finally moved.Artificial Analysis scores 3.7 Flash at 56 versus a recomputed 52 for 3.6 Flash at the same index vintage — a genuine +4 that breaks the Flash line’s “cheaper, not smarter” pattern from July.\n- 04The blank Sonnet 5 cell is a version mismatch, not a dodge.Google tested OSWorld-2.0; Anthropic publishes a different variant, OSWorld-Verified, where Sonnet 5 reportedly scores 81.2%. The two numbers must never be compared as if they were the same benchmark.\n- 05“$0.75 per million” is one of at least three live rates.Google’s intro price runs through December 31, 2026 and applies to 3.6 Flash too; the standard rate rises to $1.50/$7.50 from January 1, 2027 — the same rate 3.6 Flash’s standard pricing already carries. OpenRouter stacks its own time-limited 50%-off promo on top.\n\n## 01 — The SetupFive rows in the press cut, *nineteen* in the card.\n\nGoogle introduced Gemini 3.7 Flash as “our most intelligent workhorse model yet for coding and agents,” three weeks after [3.6 Flash’s launch](/blog/gemini-3-6-flash-launch-analysis-google-workhorse-2026) carried nearly identical positioning. The announcement post and almost every piece of launch coverage we located anchor on the same five benchmarks: FrontierCode, Terminal-bench 2.1, GDM-MRCR, OSWorld-2.0, and Harvey LAB-AA.\n\nThe model card is a different document. It compares five models across nineteen rows — and the two extra dimensions matter. The fifth column, Muse Spark 1.2, wins one row outright and is silently dropped by most coverage. And eight of the ten rows 3.7 Flash doesn’t win sit in the fourteen the press cut left out.\n\n*5* rows\n\nThe five benchmarks Google led with in its announcement — and the same five that launch coverage repeated. Two of the five are rows Google loses to GPT-5.6 Terra; one is a Google-authored benchmark.\n\n*19* rows × 5 models\n\nThe full published table: nineteen benchmark rows, a fifth model column, one blank cell on Sonnet 5’s computer-use row, and a 0.4pp inconsistency between the card and Google’s own announcement prose.\n\nNone of this means the table is dishonest — publishing rows you lose is more disclosure than most vendors manage. It means the table rewards close reading. The rest of this post is that reading.\n\n## 02 — The Full TableAll 19 rows, *tallied* honestly.\n\nThe table below transcribes every benchmark row from the Gemini 3.7 Flash model card, all five model columns included. Bold marks the best score in each row. Recomputing the tally across all nineteen rows: **3.7 Flash wins 9**, **GPT-5.6 Terra wins 6**, **Sonnet 5 wins 2** outright, **3.6 Flash wins 1** (against its own successor), and **Muse Spark 1.2 wins 1**. Against the two rivals Google names in its framing, Sonnet 5 takes three rows — GDPVal-AA, Agent’s Last Exam, and BioMysteryBench’s human-solvable subset — though on GDPVal-AA the whole table’s top score belongs to Muse Spark 1.2.\n\n| Benchmark | 3.7 Flash | 3.6 Flash | Sonnet 5 | GPT-5.6 Terra | Muse Spark 1.2 |\n|---|---|---|---|---|---|\n| Coding & software agents | |||||\n| FrontierCode 1.1Production code quality | 43.6% | 34.4% | 42.7% | 41.3% | — |\n| DeepSWE v1.1Long-horizon software engineering | 65.3% | 48.6% † | 53.8% | 69.6% | 54.9% |\n| Terminal-bench 2.1Agentic terminal coding | 85.8% | 78.0% | 80.4% | 87.4% | 82.9% |\n| Terminal-bench 3.0General agent capabilities | 14.9% | 5.4% | 14.6% | 20.8% | — |\n| WebDev Arena (Elo)Arena.ai human preference, vendor-reported | 1588 | 1538 | 1541 | 1523 | 1535 |\n| Enterprise & knowledge work | |||||\n| AutomationBenchWorkflow automation — Google’s private task set | 30.4% | 17.0% | 10.7% | 23.6% | — |\n| GDPVal-AA v2 (Elo)Knowledge work | 1525 | 1422 | 1598 | 1578 | 1628 |\n| Harvey LAB-AAComplex legal workflows | 90.7% | 85.1% | 90.1% | 85.2% | — |\n| GDP.pdfExpert PDF document comprehension | 34.0% | 22.0% | 28.0% | 24.7% | 16.0% |\n| Multimodal & long context | |||||\n| CharXiv Reasoning, no toolsChart reasoning | 84.5% | 85.2% | 77.0% | 85.9% | — |\n| CharXiv Reasoning, with toolsChart reasoning | 88.7% | 89.4% | 88.3% | — | — |\n| LVBenchLong video understanding | 85.4% | 84.2% | 68.5% | 78.9% | — |\n| GDM-MRCR v2, 128k *Long-context retrieval — Google-authored | 97.0% | 91.8% | 81.5% | 93.5% | — |\n| Computer use | |||||\n| OSWorld-2.0Agentic computer use | 47.9% | 33.8% | Not published ‡ | 50.2% | — |\n| Agent’s Last ExamMultimodal desktop / OS tasks, pass rate | 26.3% | 24.2% | 33.3% | 28.0% | — |\n| Science & expert reasoning | |||||\n| HLE-VerifiedMultidisciplinary expert reasoning | 53.6% | 51.2% | 31.0% | 51.1% | — |\n| BioMysteryBench, human-solvableBiology mysteries — solvable subset | 87.1% | 80.6% | 87.5% | 83.8% | — |\n| BioMysteryBench, human-difficultBiology mysteries — difficult subset | 43.5% | 41.2% | 34.1% | 49.4% | — |\n| LABBench2Real-world biology research tasks | 82.1% | 76.1% | 80.1% | 81.2% | — |\n\n* GDM-MRCR is a Google DeepMind-authored benchmark, self-reported on every Gemini release — read it as a vendor-designed eval, not an independent standard. † Google’s model card lists 48.6% for 3.6 Flash on DeepSWE v1.1; Google’s own announcement prose, published the same day, states 49.0% for the identical metric. We use the card as the table of record and flag the 0.4pp inconsistency rather than silently reconciling it. ‡ Not published in Google’s table. Anthropic reports Sonnet 5’s computer use on a different benchmark variant, OSWorld-Verified — see Section 05. “—” means the model card lists no score for that cell.\n\n*not equivalent*: 3.7 Flash’s ceiling is “high” on a three-rung ladder, while GPT-5.6 Terra’s published ceiling is “max” and Muse Spark 1.2’s is “xhigh” — deeper ladders whose labels don’t map one-to-one onto Google’s. Treat margins of two points or less as ties, and treat every row as a claim to verify on your own workload, not a verdict.\n\n## 03 — The WinsWhere 3.7 Flash *genuinely* wins.\n\nNine wins out of nineteen is a real result, and some of the wins are meaningful. The headline coding claim holds — narrowly: 43.6% on FrontierCode 1.1 against Sonnet 5’s 42.7%, a 0.9-point edge on production code quality. The WebDev Arena Elo of 1588 leads the field, and it is Arena.ai’s human-preference data rather than a private Google metric — though it reaches the card vendor-reported. The clearest daylight shows up on document and multimodal work: 34.0% on GDP.pdf against Sonnet 5’s 28.0%, 85.4% on LVBench long-video understanding against Sonnet 5’s 68.5% — video remains a Gemini strength that the text-first rivals don’t contest — and 53.6% on HLE-Verified, where Sonnet 5 posts a notably weak 31.0%, 22.6 points behind the leader.\n\n##### Harvey LAB-AA\n\n90.7% vs Sonnet 5’s 90.1% on complex legal workflows — close enough to call a statistical tie, though Google bolds it as a win. Terra sits at 85.2%, 3.6 Flash at 85.1%.\n\n##### AutomationBench\n\n30.4% vs Terra’s 23.6% and Sonnet 5’s 10.7% on enterprise workflow automation. The caveat: this is Google’s private, non-public task set — the card’s own notes say so.\n\n##### GDM-MRCR v2 · 128k\n\n97.0% vs Terra’s 93.5% and Sonnet 5’s 81.5%. Decisive — but GDM-MRCR is Google DeepMind’s own benchmark, self-reported on every Gemini release. A vendor-authored eval, however transparently named.\n\nNotice the pattern in those three caps: the narrowest win is on an independent benchmark, and the two most decisive wins are on a private task set and a Google-authored eval. That doesn’t make the numbers false — it makes them exactly the kind of numbers you weight down when comparing across vendors, the same discount you’d apply to any lab grading its own homework.\n\n## 04 — The LossesThe rows Google *didn’t* highlight.\n\nGPT-5.6 Terra is the strongest counterweight in Google’s own data: six rows, concentrated exactly where Google’s framing claims 3.7 Flash leads — agentic coding and computer use. Terra takes Terminal-bench 2.1 (87.4% vs 85.8%), Terminal-bench 3.0 (20.8% vs 14.9%, a hard new benchmark where the whole field scores low), DeepSWE v1.1 (69.6% vs 65.3%), and OSWorld-2.0 (50.2% vs 47.9%). Sonnet 5 takes Agent’s Last Exam outright at 33.3% — seven points clear of 3.7 Flash on multimodal desktop tasks — and edges BioMysteryBench’s human-solvable subset. On GDPVal-AA knowledge work, Sonnet 5’s 1598 Elo beats 3.7 Flash’s 1525 by 73 points, and Muse Spark 1.2 tops the whole row at 1628.\n\n#### Margins over 3.7 Flash on the rows it loses\n\nSource: Gemini 3.7 Flash model card, winning margins recomputed from Google’s own figures (GDPVal-AA’s Elo-scale losses excluded from the pp chart)The oddest entries are the two CharXiv rows, where **Gemini 3.6 Flash beats its own successor** — 85.2% vs 84.5% without tools and 89.4% vs 88.7% with tools. That is genuinely what Google’s table says, and it deserves credit for printing a same-vendor regression rather than trimming the rows. It is also a useful calibration: when a three-week successor loses to its predecessor on chart reasoning, the version bump bought capability in some places by spending it in others — normal for fast-cycle releases, invisible in a five-row press cut.\n\n## 05 — Table LiteracyThe blank cell and the *name collision*.\n\nTwo traps in this table will produce confidently wrong takes, and most launch coverage walked past both. The first is the blank Sonnet 5 cell on OSWorld-2.0. It is not a zero, and it is not evidence that Anthropic declined to disclose computer-use performance. Anthropic publishes Sonnet 5’s computer-use score on a *different* benchmark variant — OSWorld-Verified, where it reportedly scores 81.2% — and disclosed that it changed its OSWorld-Verified methodology between Sonnet 4.6 and Sonnet 5, restating Sonnet 4.6 to 78.5% in the process. Google tested OSWorld-2.0. Different variant, different methodology, different scale: the 81.2% and the 47.9%–50.2% range in Google’s table must never be cross-compared as the same number.\n\n*AutomationBench*in Google’s table (3.7 Flash: 30.4%) is Google’s private enterprise task set.\n\n*AutomationBench-AA*(3.7 Flash: 62.7%) is Artificial Analysis’s own, separate benchmark. Same-sounding name, different tasks, different scoring — merging them, or quoting the 62.7% as an improvement over the 30.4%, is a fabricated comparison.\n\nBoth traps are instances of the general failure mode we catalogued in [our guide to reading vendor benchmark tables](/blog/vendor-benchmark-tables-reading-disclosed-losses-2026): version mismatches and missing cells get read as verdicts when they are artifacts of who tested what, under which harness. Add this card’s own contribution to the genre — the DeepSWE figure for 3.6 Flash that reads 48.6% in the table and 49.0% in the announcement prose published the same day — and the lesson generalizes: even first-party numbers disagree with themselves at the margins, which is precisely why margins under a point should never drive a routing decision.\n\n## 06 — Independent SignalThe independent read: *52 → 56*, same vintage.\n\nThis is the part of the story that vendor tables can’t settle, and it is where 3.7 Flash earns its most defensible claim. Artificial Analysis scores Gemini 3.7 Flash (high) at **56 on its Intelligence Index** — a 4-point improvement over Gemini 3.6 Flash’s 52, with both scores computed at the same index vintage (v4.1.1, a nine-evaluation composite). That matters because of what came before: [our July read of 3.6 Flash](/blog/gemini-3-6-flash-benchmarks-vs-gpt-5-6-sonnet-5-kimi-k3) found a flat Intelligence Index — “cheaper, not smarter.” That pattern breaks here. AA also puts 3.7 Flash on its Intelligence-vs-Time Pareto frontier at 1.7 minutes per task, which it describes as “40% faster than GPT-5.6 Terra (max),” with output throughput around 340 tokens per second.\n\n#### AA Intelligence Index · same-vintage scores by effort configuration\n\nSource: Artificial Analysis Intelligence Index v4.1.1*50*— an earlier AA index vintage. AA has since updated its methodology to v4.1.1 and recomputed 3.6 Flash’s own score to 52. Comparing July’s 50 to today’s 56 conflates two vintages; the valid same-vintage comparison is AA’s own 52 → 56. The +4 is real either way — but only one arithmetic is honest.\n\nTwo qualifiers keep the independent read honest. First, the effort-ladder asymmetry from Section 02 applies here too: 3.7 Flash’s 56 is the top of a three-rung ladder (51/53/56), while Terra’s 57 comes at “max” and Muse Spark’s 57 at “xhigh” — each model’s respective ceiling, on ladders of different depth. Second, the ceiling is close but not reached: 56 is one point behind both, and VentureBeat notes Artificial Analysis places Claude Opus 5 at 63 on the same index — the frontier tier remains a different conversation. On AA’s sub-indices, as surfaced through OpenRouter’s AA-sourced panel, 3.7 Flash posts a Coding Index of 76.1 and an Agentic Index of 45.1; Arena.ai’s early human-preference results provisionally rank it ninth overall and eighth for web development. Independent evals exist this time — plural, and broadly consistent with each other.\n\n\"That is the metric enterprise teams will ultimately need to test: not price per million tokens in isolation, but cost per successfully completed task.\"— VentureBeat, Gemini 3.7 Flash launch coverage, August 13, 2026\n\n## 07 — PricingOne model, *three* live rates — label every price by surface.\n\n“Gemini 3.7 Flash costs $0.75 per million tokens” is true on exactly one surface. Google’s official API lists the introductory rate through December 31, 2026, and posts a standard rate of $1.50/$7.50 from January 1, 2027 — the same rate 3.6 Flash’s standard pricing already carries, and 3.6 Flash shares the intro window too, so the cliff as posted today reads Flash-line-wide rather than specific to the new model. The widely repeated “half price” framing is only true against the Flash line’s pre-August-13 workhorse rate. Meanwhile OpenRouter layers its own time-limited 50%-off promotion on top of Google’s already-discounted intro price. Each row below names its surface.\n\n| Pricing surface | Input $/1M | Output $/1M | Conditions |\n|---|---|---|---|\n| Gemini 3.7 Flash — three live rates for one model | |||\n| Google API, introductory | $0.75 | $3.75 | Through December 31, 2026. The same intro rate applies to 3.6 Flash. Context caching $0.075/1M in the same window. |\n| Google API, standard | $1.50 | $7.50 | The rate Google’s pricing page posts for the end of the intro window on January 1, 2027 — double the intro price, and the same $1.50/$7.50 that 3.6 Flash’s standard pricing already carries. |\n| OpenRouter promo | $0.375 | $1.875 | OpenRouter’s own “50% off for a limited time” banner, stacked on Google’s intro rate and expiring on OpenRouter’s schedule, not Google’s. |\n| The rivals in Google’s table — official list rates | |||\n| Claude Sonnet 5 | $2.00 | $10.00 | Now the permanent standard price — the increase to $3/$15 scheduled for September 1 was cancelled by Anthropic. Batch: $1.00/$5.00. |\n| GPT-5.6 Terra, short context | $2.00 | $12.00 | Standard list. Batch and Flex: $1.00/$6.00; Fast mode: $4.00/$24.00. |\n| GPT-5.6 Terra, long context | $4.00 | $18.00 | Per OpenAI’s own pricing docs, whole requests above a 272K-token threshold reprice — 2× on input, 1.5× on output — a structure Google’s and Anthropic’s flat tables don’t have. |\n| Muse Spark 1.2 | $1.25 | $4.25 | As listed in Google’s own model-card pricing row — cheaper than Sonnet 5 and Terra while tying Terra on the AA Intelligence Index. |\n\nOne footnote deserves its own hedge: the model card marks the Flash prices with an asterisk whose footnote text we could not retrieve from the page itself. The introductory-window reading is consistent with Google’s announcement and its official pricing page, so we treat it as the intro-price marker — but it is a presumption, not a verbatim confirmation. And on whether the January 1 doubling actually lands, there is a fresh precedent pointing the other way: Sonnet 5’s own $2/$10 rate was announced at launch as introductory pricing with a scheduled step-up to $3/$15 — and Anthropic cancelled that increase, making the intro rate permanent. Vendors have now demonstrated that a scheduled price cliff is a plan, not a promise. Budget for $1.50/$7.50 from January; don’t be surprised if the cliff moves.\n\n## 08 — Decision MatrixRouting the workloads this table *actually* settles.\n\nRead as a whole, the nineteen rows don’t crown a single model — they partition the workload space. Here is how we’d translate the full table, the independent AA read, and the surface-labelled prices into routing defaults, treating every margin of two points or less as a tie to be broken by price.\n\n##### High-volume code generation & web dev\n\n3.7 Flash wins FrontierCode narrowly and WebDev Arena’s Elo outright while costing a fraction of either rival on the intro rate. At $0.75/$3.75 through December, the price-per-win math is hard to argue with for volume work.\n\n##### Terminal agents & *computer use*\n\nGoogle’s own table gives GPT-5.6 Terra Terminal-bench 2.1 and 3.0, DeepSWE, and OSWorld-2.0 — the exact categories in 3.7 Flash’s positioning. Mind OpenAI’s published 272K long-context repricing on big agent transcripts.\n\n##### Judgment-heavy documents & desktop tasks\n\nSonnet 5 beats both named rivals on GDPVal-AA and wins Agent’s Last Exam outright, at a $2/$10 rate that is now permanent after the cancelled September increase. Its weak HLE-Verified row cuts the other way — test your own mix.\n\n##### Anything mission-critical\n\nEight of nineteen rows are decided by two points or less, effort configs aren’t comparable across vendors, and one cell is a version mismatch. Vendor tables shortlist candidates; your own task-level evals pick the winner.\n\nThe projection worth making: three-week release cycles mean this table is a snapshot, not a standings board. The durable takeaways are structural — Google is now competitive enough on coding benchmarks to force per-task price comparisons, OpenAI holds the hardest agentic rows, Anthropic holds judgment-heavy knowledge work, and every vendor’s pricing now carries dated conditions that move independently of capability. If your stack still routes every workload to one default model, that is the actual finding of this table — and it’s the kind of routing decision [our AI transformation engagements](/services/ai-transformation) start with, benchmarked on your workloads rather than the vendor’s. For the wider market context beyond these three vendors, see [our broader model comparison](/blog/chatgpt-vs-claude-vs-gemini-vs-grok-ai-comparison).\n\n## 09 — ConclusionA split decision — and an index that finally *moved*.\n\n### Nine wins out of nineteen is a strong workhorse, not a coronation.\n\nRead in full, Google’s own table says something more interesting than the launch framing: Gemini 3.7 Flash wins nine of nineteen rows, loses the hardest agentic benchmarks to GPT-5.6 Terra, cedes knowledge-work rows to Claude Sonnet 5 and Muse Spark 1.2, and even loses two chart-reasoning rows to *its own predecessor*. VentureBeat’s independent read matches ours: “In other words, Google’s own results do not show 3.7 Flash universally displacing higher-priced competitors. They instead suggest a model that has become substantially more competitive in coding and agent workloads while occupying a lower price tier.”\n\nThe strongest claim in the release isn’t in Google’s table at all — it’s Artificial Analysis moving its Intelligence Index from 52 to 56 at matched vintage, breaking the “cheaper, not smarter” pattern we documented in July. A real capability gain, delivered three weeks after the last one, at the lowest intro price in its comparison set, matched only by 3.6 Flash: that combination is the story, and it survives every caveat this post has raised.\n\nThe caveats still matter. Prices carry surfaces and expiry dates; benchmarks carry versions, authors, and effort configs; blank cells carry explanations. The teams that win with these releases aren’t the ones that pick whichever model’s vendor published the most flattering table this week — they’re the ones with the eval harness and the routing layer to re-test the frontier every few weeks and move traffic on their own numbers.", "url": "https://wpnews.pro/news/gemini-3-7-flash-vs-sonnet-5-vs-gpt-5-6-terra-real-wins", "canonical_source": "https://www.digitalapplied.com/blog/gemini-3-7-flash-vs-sonnet-5-gpt-5-6-terra-benchmarks", "published_at": "2026-08-14 00:00:00+00:00", "updated_at": "2026-08-14 15:07:03.203052+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-products", "ai-research"], "entities": ["Google", "Gemini 3.7 Flash", "GPT-5.6 Terra", "Claude Sonnet 5", "Muse Spark 1.2", "Artificial Analysis", "Anthropic", "OpenRouter"], "alternates": {"html": "https://wpnews.pro/news/gemini-3-7-flash-vs-sonnet-5-vs-gpt-5-6-terra-real-wins", "markdown": "https://wpnews.pro/news/gemini-3-7-flash-vs-sonnet-5-vs-gpt-5-6-terra-real-wins.md", "text": "https://wpnews.pro/news/gemini-3-7-flash-vs-sonnet-5-vs-gpt-5-6-terra-real-wins.txt", "jsonld": "https://wpnews.pro/news/gemini-3-7-flash-vs-sonnet-5-vs-gpt-5-6-terra-real-wins.jsonld"}}