Gemini 3.8 Flash challenges Claude Opus 5 on key benchmarks at a fraction of the price Google DeepMind released Gemini 3.8 Flash on September 2, scoring 54.9% on HLE-Verified reasoning tests and outperforming Anthropic's Claude Opus 5 on engineering benchmarks at roughly one-seventh the token cost. Priced at $0.75 per million input tokens and $3.75 per million output tokens, Gemini 3.8 Flash undercuts Claude Opus 5's $5 and $25 rates, offering a compelling value proposition for businesses, though Claude Opus 5 retains an edge on agentic tasks. Photo: Steve A Johnson / Pexels Gemini 3.8 Flash challenges Claude Opus 5 on key benchmarks at a fraction of the price Google DeepMind's latest Flash model scores 54.9% on HLE-Verified reasoning tests and beats larger frontier models on engineering tasks, all at roughly one-seventh the token cost of Anthropic's flagship. Google DeepMind dropped Gemini 3.8 Flash on September 2, and it’s already picking fights above its weight class. The model outperforms Anthropic’s Claude Opus 5, a model that costs roughly seven times more per output token, on several critical benchmarks spanning coding, multi-step reasoning, and legal task performance. The kicker: Gemini 3.8 Flash is priced at $0.75 per million input tokens and $3.75 per million output tokens. Claude Opus 5 runs $5 per million input tokens and $25 per million output tokens. The benchmark breakdown Gemini 3.8 Flash posted a 54.9% score on HLE-Verified, a demanding multi-step reasoning evaluation that spans STEM, humanities, and other disciplines. The model also outperformed larger frontier models on the DeepSWE v1.1 benchmark, which measures performance on complex engineering problem-solving tasks. Claude Opus 5, which Anthropic launched on July 24, still leads on several benchmarks in agentic knowledge work and coding tasks. Independent evaluations as of early September show the two models trading leads depending on the category being tested. The relevant benchmarks for comparison include SWE-bench and Terminal-Bench, among others. On some of these evaluations, Claude Opus 5 maintains an edge, particularly in tasks requiring sustained autonomous reasoning over longer workflows. Three models in six weeks Gemini 3.8 Flash is the third Flash-series model Google DeepMind has released in just six weeks. The gap between Gemini 3.7 Flash and 3.8 Flash was approximately three weeks. Anthropic has positioned Claude Opus 5 as a premium product, pricing it at $25 per million output tokens compared to Flash’s $3.75. What this means for the AI market For businesses evaluating which models to integrate into their workflows, Gemini 3.8 Flash’s benchmark performance at its price point creates a compelling value proposition. A company running millions of API calls per day could see substantial cost reductions by switching from a premium model to Flash without sacrificing meaningful capability on most tasks. The math becomes especially attractive for software engineering workflows, where the DeepSWE results suggest Flash can handle complex coding challenges that would previously have required a more expensive model. Claude Opus 5 still commands respect on agentic benchmarks, and enterprises with demanding use cases in autonomous research, legal analysis, and multi-step planning may still find the premium worthwhile. Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy https://cryptobriefing.com/editorial-policy/ .