cd /news/artificial-intelligence/grok-4-6-hits-1753-elo-on-gdpval-aa-… · home topics artificial-intelligence article
[ARTICLE · art-93887] src=cryptobriefing.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Grok 4.6 hits 1,753 Elo on GDPVal-AA, claiming top non-Anthropic model spot

XAI's Grok 4.6, launched in early August 2026, scored 1,753 Elo on the GDPval-AA v2 leaderboard, placing third overall and becoming the highest-ranked model from any company other than Anthropic. The score, which carries a margin of ±21 Elo, marks a roughly 200-point jump from Grok 4.5's mid-1500s rating, achieved through refined supervised fine-tuning and updated reinforcement learning on the same 1.5 trillion-parameter V9 foundation. Grok 4.6 trails only Anthropic's Claude Opus 5 variants, which scored 1,849 and 1,817 Elo.

read2 min views4 publishedAug 12, 2026
Grok 4.6 hits 1,753 Elo on GDPVal-AA, claiming top non-Anthropic model spot
Image: Cryptobriefing (auto-discovered)

xAI's latest model jumps from the mid-1500s to third place overall on the real-world knowledge work leaderboard, trailing only Anthropic's Claude Opus 5 variants

xAI’s Grok 4.6 just handed the AI benchmark charts a reason to pay attention. Launched in early August 2026, the model scored 1,753 Elo on the GDPval-AA v2 leaderboard, placing third overall and cementing itself as the highest-ranked model from any company not named Anthropic.

That gap from its predecessor is not trivial. Grok 4.5 was sitting in the mid-1500 Elo range before this update, meaning Grok 4.6 added roughly 200 Elo points in a single generational jump.

What GDPval actually measures #

OpenAI introduced the GDPval benchmark in September 2025 with a specific goal: stop measuring AI against academic puzzles and start measuring it against the kind of work that actually generates economic value.

The benchmark evaluates models on tasks developed by professionals with an average of 14 years of experience across various industries. The Elo system used here works the same way it does in chess: models compete head-to-head on identical tasks, and ratings shift based on wins and losses.

Artificial Analysis runs the evaluations and maintains the leaderboard, applying the Elo methodology to keep comparisons consistent across model versions and providers. Grok 4.6’s score of 1,753 carries a margin of plus or minus 21 Elo points.

How Grok 4.6 got here #

Grok 4.6 is built on the same 1.5 trillion-parameter V9 foundation as Grok 4.5, so the score improvement did not come from simply throwing more compute at the problem. xAI attributes the gains to refined supervised fine-tuning and updated reinforcement learning approaches applied during post-training.

Grok 4.6 sits behind two Claude Opus 5 variants on the leaderboard, which scored 1,849 and 1,817 Elo respectively. The gap between Grok 4.6 and the second Claude Opus 5 variant is about 64 Elo points.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our

Editorial Policy.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @xai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/grok-4-6-hits-1753-e…] indexed:0 read:2min 2026-08-12 ·