cd /news/artificial-intelligence/google-releases-gemini-3-8-flash-bea… · home topics artificial-intelligence article
[ARTICLE · art-119096] src=officechai.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Google Releases Gemini 3.8 Flash, Beats Opus 5, GPT 5.6 Sol On Some Benchmarks At A Fraction Of The Price

Google released Gemini 3.8 Flash, priced at $0.75 per million input tokens and $3.75 per million output tokens through 2026, which outperforms Anthropic's Claude Opus 5 and OpenAI's GPT-5.6 Sol on several benchmarks including Vals Finance Agent v2 (61.4% vs 58.6% and 53.8%) and Harvey's Legal Agent Benchmark (10.0% vs 6.7% and 2.5%), while costing 6-7x less. However, Opus 5 still leads on heavier benchmarks like Terminal-bench 4.0 (51.8% vs 19.1%) and GDPVal-AA v2 (1824 vs 1545).

read4 min views1 publishedSep 2, 2026
Google Releases Gemini 3.8 Flash, Beats Opus 5, GPT 5.6 Sol On Some Benchmarks At A Fraction Of The Price
Image: Officechai (auto-discovered)

Google might be back in the AI game.

Google has released Gemini 3.8 Flash, the third update to its Flash model line in just five weeks, and the numbers it’s putting up are hard to ignore given the price tag attached to them.

The model is priced at $0.75 per million input tokens and $3.75 per million output tokens — a discounted introductory rate that Google says will hold through the end of 2026, with the regular pricing set at $1.50 and $7.50 respectively. That makes Gemini 3.8 Flash roughly 6-7x cheaper on input and output than Claude Opus 5, which is priced at $5 input / $25 output per million tokens, and meaningfully cheaper than GPT-5.6 Sol as well, which comes in at $4 input / $20 output.

And yet, on several benchmarks, the cheap little Flash model is beating both of them.

Gemini 3.8 Flash Benchmarks #

Gemini 3.8 Flash tops the table on a number of practical, agent-heavy benchmarks. On Vals Finance Agent v2, which tests financial analyst tasks, it scores 61.4%, ahead of Opus 5’s 58.6% and GPT-5.6 Sol’s 53.8%. On Harvey’s Legal Agent Benchmark, which covers complex legal workflows, it posts a 10.0% pass rate — more than double GPT-5.6 Sol’s 2.5% and comfortably ahead of Opus 5’s 6.7%. It also edges out Opus 5 on Terminal-bench 2.1, an agentic terminal coding benchmark, with 89.4% against Opus 5’s 89.1%.

The pattern continues elsewhere. On CharXiv Reasoning, which measures information synthesis from complex charts, Gemini 3.8 Flash hits 86.2%, ahead of Opus 5’s 83.7%. On LVBench, a long video understanding benchmark, it reaches 87.8% in agentic mode, well clear of Opus 5’s 75.4%. It narrowly leads on HLE-Verified, a multidisciplinary expert reasoning test, at 54.9% versus Opus 5’s 54.4%, and it comes out ahead on LABBench2, which evaluates real-world biology research tasks, at 86.2% against Opus 5’s 84.2%. On the harder tier of BioMysteryBench, labeled “Human Difficult,” Gemini 3.8 Flash scores 56.5%, ahead of both Opus 5’s 49.4% and GPT-5.6 Sol’s 44.7%.

That’s a lot of green across finance, legal, coding, video, and scientific research workflows — categories that matter a great deal to the enterprise customers Google is chasing with its Flash line.

To be clear, Gemini 3.8 Flash is not simply better than Opus 5 across the board — and Google isn’t claiming that either. Opus 5, which sits at the top of Anthropic’s lineup and costs considerably more to run, still leads on several of the heavier benchmarks. On DeepSWE v1.1, a long-horizon software engineering benchmark, Opus 5 scores 74.0% against Gemini 3.8 Flash’s 71.0%. On GDPVal-AA v2, an Elo-scored knowledge work benchmark, Opus 5 leads by a wide margin at 1824 versus 1545.

The biggest gap shows up on Terminal-bench 4.0, a general agent capabilities benchmark, where Opus 5 scores 51.8% against Gemini 3.8 Flash’s 19.1%. Opus 5 also leads on OSWorld-2.0, which tests agentic computer use, at 75.4% versus 59.0%, and on the easier “Human Solvable” tier of BioMysteryBench, at 90.1% versus 88.8%.

The gap on Terminal-bench 4.0 in particular stands out — this is one of the harder, more general agentic benchmarks in the set, and it’s where the difference between a flagship reasoning model and a fast, cheap Flash model shows up most clearly.

Compared to its predecessor Gemini 3.7 Flash, though, 3.8 Flash is an improvement almost everywhere on the sheet, from Terminal-bench 4.0 (19.1% vs 11.2%) to BioMysteryBench Human Difficult (56.5% vs 43.5%) to OSWorld-2.0 (59.0% vs 50.6%).

The Bigger Picture #

Google has been on an unusually aggressive release cadence with its Flash line this year — Gemini 3.6 Flash landed in July, 3.7 Flash followed in mid-August, and now 3.8 Flash has arrived just weeks later. The pattern suggests Google is treating Flash less like a traditional annual or bi-annual model release and more like a continuously iterated product, shipping incremental gains every few weeks rather than waiting for a single big leap.

The strategy also seems to be a deliberate pricing play. Rather than trying to beat Opus 5 or GPT-5.6 Sol outright across every benchmark, Google appears to be targeting the specific workflows — coding agents, financial analysis, legal review, long-video and document understanding — where a fast, cheap model can get close enough to frontier performance that the price difference becomes the deciding factor for businesses running these tasks at scale.

For companies running high-volume agentic workloads, a model that costs a fraction of Opus 5 while beating or matching it on more than half the benchmarks tested is likely to be an appealing proposition, even if it isn’t the strongest model on the market in absolute terms.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @google 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/google-releases-gemi…] indexed:0 read:4min 2026-09-02 ·