xAI shipped Grok 4.7 on September 21 and pushed it to GitHub Copilot across every subscription plan the same day. The 2.1 trillion parameter model comes with a 500K context window, improved self-verification, and extended reinforcement learning aimed at multi-hour agentic tasks. The marketing frames it as xAI’s “most capable coding model yet.” Independent benchmarks tell a more complicated story. Where Grok 4.7 Actually Stands On the Artificial Analysis Intelligence Index (v4.3.2), which aggregates ten benchmarks, Grok 4.7 scores 46. Claude Fable 5.1 and GPT-6 Astra both score 53 — a 13% gap. That gap narrows for specific workloads, but […]
The post