{"slug": "grok-4-7-takes-the-top-three-spots-on-vulcanbench-frontier-v4", "title": "Grok 4.7 takes the top three spots on VulcanBench Frontier v4", "summary": "XAI's Grok 4.7, launched September 21, 2026, took first, second and third place on the VulcanBench Frontier v4 leaderboard, scoring 93.15 at extra-high reasoning effort, 92.71 at high effort and 92.30 at medium effort, all above rival Claude Fable 5.1's 91.84. At extra-high effort the model passed all 23 behavioral-reconstruction tasks on the benchmark, which uses deterministic hidden tests to weigh functional correctness, code quality and complexity. Grok 4.7 keeps Grok 4.6's pricing of $2 per million input tokens and $6 per million output tokens, supports text and image inputs, and offers a 500,000-token context window.", "body_md": "Photo: Tima Miroshnichenko / Pexels\n\n# Grok 4.7 takes the top three spots on VulcanBench Frontier v4\n\nxAI's newest coding model claimed first, second and third place on the leaderboard while keeping its price unchanged\n\nxAI has a new coding model, and it arrived with a scoreboard. Grok 4.7, launched on September 21, 2026, now holds first, second and third place on the VulcanBench Frontier v4 leaderboard.\n\n## Three effort levels, three podium spots\n\nGrok 4.7 lets users pick how much reasoning effort the model spends on a task. Each setting was scored separately on VulcanBench Frontier v4, and each landed near the top.\n\nAt the extra-high effort level, Grok 4.7 scored 93.15, good for first place. The high effort setting followed at 92.71 in second, and the medium setting took third with 92.30.\n\nThe closest rival named in the results was Claude Fable 5.1, which scored 91.84. That puts even Grok 4.7’s medium setting ahead of it on this particular test.\n\nThe standout result came at extra-high effort. There, Grok 4.7 passed all 23 behavioral-reconstruction tasks on the benchmark, with strong code quality scores under the latest testing protocol.\n\nA behavioral-reconstruction task asks a model to rebuild software so that it behaves exactly like an existing reference. Getting the output roughly right does not count, because the code has to match what the original actually does.\n\n### AI, tech, and the markets they move—in one daily briefing.\n\nDaily. Free. Join 34,000+ readers across crypto, finance, and policy.\n\n## What VulcanBench actually measures\n\nVulcanBench Frontier v4 is designed to evaluate AI models on real-world software engineering problems. It weighs functional correctness, code quality and complexity rather than rewarding a model for producing code that merely looks plausible.\n\nThe platform leans on deterministic hidden tests and emphasizes transparent evaluation. Put simply, the model cannot see the answer key, and the same code always gets the same grade.\n\n## Under the hood and on the price tag\n\nGrok 4.7 is built on an extended base model compared with its predecessor, Grok 4.6. It supports text and image inputs, works with additional tools, and offers a context window of 500,000 tokens.\n\nPricing is unchanged from Grok 4.6. Grok 4.7 costs $2 per million input tokens and $6 per million output tokens.\n\nxAI also reported that Grok 4.7 improved on Grok 4.6 across multiple other benchmarks, including CursorBench and Terminal-Bench. Both are aimed at coding work, with Terminal-Bench focused on tasks carried out in a command-line environment.\n\n## What this means for developers and the AI market\n\nFor developers, the practical takeaway is choice. With three effort levels all scoring above 92, teams may be able to run Grok 4.7 at a lower setting for routine work and save the extra-high mode for harder problems.\n\nThere are caveats worth keeping in mind. A single leaderboard captures one slice of performance, and the gap between first place and Claude Fable 5.1 is 1.31 points at the extra-high setting.\n\n**Disclosure:** This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our\n\n[Editorial Policy](https://cryptobriefing.com/editorial-policy/).", "url": "https://wpnews.pro/news/grok-4-7-takes-the-top-three-spots-on-vulcanbench-frontier-v4", "canonical_source": "https://cryptobriefing.com/grok-4-7-tops-vulcanbench-frontier-v4/", "published_at": "2026-10-06 15:39:32+00:00", "updated_at": "2026-10-06 15:46:48.335438+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-products", "developer-tools"], "entities": ["xAI", "Grok 4.7", "VulcanBench Frontier v4", "Claude Fable 5.1", "Grok 4.6", "CursorBench", "Terminal-Bench", "Diego Almada Lopez"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/grok-4-7-takes-the-top-three-spots-on-vulcanbench-frontier-v4", "markdown": "https://wpnews.pro/news/grok-4-7-takes-the-top-three-spots-on-vulcanbench-frontier-v4.md", "text": "https://wpnews.pro/news/grok-4-7-takes-the-top-three-spots-on-vulcanbench-frontier-v4.txt", "jsonld": "https://wpnews.pro/news/grok-4-7-takes-the-top-three-spots-on-vulcanbench-frontier-v4.jsonld"}}