cd /news/artificial-intelligence/grok-4-7-takes-the-top-three-spots-o… · home › topics › artificial-intelligence › article
[ARTICLE · art-146142] src=cryptobriefing.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Grok 4.7 takes the top three spots on VulcanBench Frontier v4

XAI's Grok 4.7, launched September 21, 2026, took first, second and third place on the VulcanBench Frontier v4 leaderboard, scoring 93.15 at extra-high reasoning effort, 92.71 at high effort and 92.30 at medium effort, all above rival Claude Fable 5.1's 91.84. At extra-high effort the model passed all 23 behavioral-reconstruction tasks on the benchmark, which uses deterministic hidden tests to weigh functional correctness, code quality and complexity. Grok 4.7 keeps Grok 4.6's pricing of $2 per million input tokens and $6 per million output tokens, supports text and image inputs, and offers a 500,000-token context window.

by read3 min views4 publishedOct 6, 2026
Grok 4.7 takes the top three spots on VulcanBench Frontier v4
Image: Cryptobriefing (auto-discovered)

Photo: Tima Miroshnichenko / Pexels

xAI's newest coding model claimed first, second and third place on the leaderboard while keeping its price unchanged

xAI has a new coding model, and it arrived with a scoreboard. Grok 4.7, launched on September 21, 2026, now holds first, second and third place on the VulcanBench Frontier v4 leaderboard.

Three effort levels, three podium spots #

Grok 4.7 lets users pick how much reasoning effort the model spends on a task. Each setting was scored separately on VulcanBench Frontier v4, and each landed near the top.

At the extra-high effort level, Grok 4.7 scored 93.15, good for first place. The high effort setting followed at 92.71 in second, and the medium setting took third with 92.30.

The closest rival named in the results was Claude Fable 5.1, which scored 91.84. That puts even Grok 4.7’s medium setting ahead of it on this particular test.

The standout result came at extra-high effort. There, Grok 4.7 passed all 23 behavioral-reconstruction tasks on the benchmark, with strong code quality scores under the latest testing protocol.

A behavioral-reconstruction task asks a model to rebuild software so that it behaves exactly like an existing reference. Getting the output roughly right does not count, because the code has to match what the original actually does.

AI, tech, and the markets they move—in one daily briefing.

Daily. Free. Join 34,000+ readers across crypto, finance, and policy.

What VulcanBench actually measures #

VulcanBench Frontier v4 is designed to evaluate AI models on real-world software engineering problems. It weighs functional correctness, code quality and complexity rather than rewarding a model for producing code that merely looks plausible.

The platform leans on deterministic hidden tests and emphasizes transparent evaluation. Put simply, the model cannot see the answer key, and the same code always gets the same grade.

Under the hood and on the price tag #

Grok 4.7 is built on an extended base model compared with its predecessor, Grok 4.6. It supports text and image inputs, works with additional tools, and offers a context window of 500,000 tokens.

Pricing is unchanged from Grok 4.6. Grok 4.7 costs $2 per million input tokens and $6 per million output tokens.

xAI also reported that Grok 4.7 improved on Grok 4.6 across multiple other benchmarks, including CursorBench and Terminal-Bench. Both are aimed at coding work, with Terminal-Bench focused on tasks carried out in a command-line environment.

What this means for developers and the AI market #

For developers, the practical takeaway is choice. With three effort levels all scoring above 92, teams may be able to run Grok 4.7 at a lower setting for routine work and save the extra-high mode for harder problems. There are caveats worth keeping in mind. A single leaderboard captures one slice of performance, and the gap between first place and Claude Fable 5.1 is 1.31 points at the extra-high setting.

Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our

Editorial Policy.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @xai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/grok-4-7-takes-the-t…] indexed:0 read:3min 2026-10-06 · —