Grok 4.7 Is Out — Benchmarks Show It’s Mid-Pack vs. Claude and GPT-6 XAI released Grok 4.7 on September 21 and pushed it to GitHub Copilot across every subscription plan the same day, with the 2.1 trillion parameter model featuring a 500K context window, improved self-verification, and extended reinforcement learning for multi-hour agentic tasks. On the Artificial Analysis Intelligence Index (v4.3.2), which aggregates ten benchmarks, Grok 4.7 scores 46, while Claude Fable 5.1 and GPT-6 Astra both score 53 — a 13% gap that narrows for specific workloads. xAI shipped Grok 4.7 on September 21 and pushed it to GitHub Copilot across every subscription plan the same day. The 2.1 trillion parameter model comes with a 500K context window, improved self-verification, and extended reinforcement learning aimed at multi-hour agentic tasks. The marketing frames it as xAI’s “most capable coding model yet.” Independent benchmarks tell a more complicated story. Where Grok 4.7 Actually Stands On the Artificial Analysis Intelligence Index v4.3.2 , which aggregates ten benchmarks, Grok 4.7 scores 46. Claude Fable 5.1 and GPT-6 Astra both score 53 — a 13% gap. That gap narrows for specific workloads, but … The post