Google isn’t quite at AI frontier any more, but Gemini 3.7 Flash is proving to be surprisingly capable in terms of performance and cost.
The ARC Prize Foundation has published verified ARC-AGI results for Gemini 3.7 Flash, and the numbers are hard to ignore even for a lab that’s spent most of this year playing catch-up to Anthropic and OpenAI. Google’s model scores 84.6% on ARC-AGI-2 at $0.25 per task, and 95.5% on ARC-AGI-1 at just $0.12 per task. On the leaderboard’s cost-versus-score plot, that places Gemini 3.7 Flash almost alone in the lower-left region where high scores and low cost are both true at the same time, well ahead of everything else that costs under a dollar a task.
For context on what 84.6% actually means: it puts Gemini 3.7 Flash within a handful of points of Claude Fable 5 (Max), which sits close to the top of the ARC-AGI-2 board around the low-to-mid 90s, and in the same neighborhood as GPT-5.6 Sol (Max), currently the highest scorer on the chart. The gap in score between Gemini 3.7 Flash and these frontier reasoning models is real but small. The gap in price is not small at all. Fable 5 and Sol are both running at costs several multiples higher per task, deep in the territory where a single ARC-AGI-2 run costs dollars rather than cents. Gemini 3.7 Flash is doing most of the job those models do, for a fraction of what they charge to do it. ARC-AGI-2 exists specifically to punish models that lean on memorized patterns instead of actual reasoning. Each task hands a model a small number of visual examples and asks it to infer the underlying rule, then apply that rule to a new grid it hasn’t seen. It’s a benchmark built by François Chollet to resist the usual ways labs inflate scores, and it’s proven stubborn: progress on it has historically come in small increments, with even flagship reasoning models struggling to clear the 30-40% range until recently. A Flash-tier model, the kind of product line Google has traditionally positioned as its budget option, landing at 84.6% is not a small increment.
That’s also what makes this result different from the usual “cheaper model matches expensive one” story that’s become common this year. Gemini 3.7 Flash isn’t winning by being fast at simple tasks or by doing well on benchmarks that reward volume and speed rather than depth. ARC-AGI-2 is specifically the kind of test where cheap, high-throughput models tend to fall apart, because pattern-matching genuinely does not help. Google’s own framing of the release leaned into this, describing Gemini 3.7 Flash as standing out for its combination of low cost and high scores relative to other frontier models on both ARC-AGI-1 and ARC-AGI-2, and the leaderboard backs that framing up.
It also fits a pattern that’s shown up elsewhere this month. Gemini 3.7 Flash recently topped Artificial Analysis’s AA-AnalystAgent benchmark, a test built around the kind of spreadsheet and data-analysis work that shows up in actual business workflows, beating Claude Opus 5, GPT-5.5, and Fable 5 itself along the way. That result and this one point in the same direction: Google’s Flash line has quietly become a serious reasoning product, not just a speed play. When Google released Gemini 3.7 Flash a few weeks ago, the pitch was about price cuts and incremental gains over Gemini 3.6 Flash. The ARC-AGI numbers suggest the incremental framing undersold what shipped.
None of this means Gemini 3.7 Flash has closed the gap with the actual frontier. Fable 5 and Sol still score higher on ARC-AGI-2, and on harder, more novel benchmarks like ARC-AGI-3, where Anthropic’s Opus 5 tripled the previous best score, the newest reasoning-heavy flagships remain in a different class altogether. Google also hasn’t shipped a true Pro or Ultra successor since February, leaving Flash to carry the company’s frontier narrative largely on its own for months now. What Gemini 3.7 Flash does show is that the price of getting close to frontier-level abstract reasoning has dropped considerably faster than the price of the frontier itself, and for the vast majority of teams building products on top of these models, that curve matters more than who sits at the very top of the leaderboard.
For businesses running ARC-AGI-style reasoning tasks at any real volume, that difference between $0.25 and several dollars a task isn’t minor — it’s the difference between a workload that’s economical to run constantly and one that gets reserved for the cases that genuinely need it.