SpaceXAI’s newest model, Grok 4.7, has posted a modest but meaningful gain on the Artificial Analysis Intelligence Index, scoring 46 — up 2 points from Grok 4.6’s 44 — while making a far bigger leap in agentic knowledge work and coding. The release lifts SpaceXAI into the top four AI labs on the benchmark, even as the model still trails several rivals on the headline score, including GPT-5.6 Sol, Meta’s Muse Spark 1.3, and Anthropic’s Claude Fable 5.
The Intelligence Index, rebalanced most recently to version 4.3, incorporates 10 evaluations spanning knowledge work, agentic tasks, coding, and scientific reasoning. Grok 4.7 was evaluated at xhigh reasoning effort.
Where The Score Sits #
At 46, Grok 4.7 sits one point behind GPT-5.6 Sol (47), two behind Muse Spark 1.3 (48), and four behind Claude Fable 5 (50). The top of the table remains crowded: Claude Fable 5.1 and GPT-6 Astra lead at 53 apiece, followed by Claude Opus 5 at 51. For SpaceXAI, though, the 2-point jump over Grok 4.6 is enough to move the lab into the index’s top four — a symbolic milestone even if the raw number keeps it in the middle of the frontier pack.
The Real Story: Agentic Knowledge Work #
Where Grok 4.7 genuinely joins the frontier is on AA-Briefcase, Artificial Analysis’s private benchmark for long-horizon agentic knowledge work. The model scores 1,657 Elo — a gain of 111 over Grok 4.6 (high) — placing it alongside Claude Opus 5 and Claude Fable 5.1 at the very top of the leaderboard.
The improvement is led by analytical quality. Grok 4.7 scores 1,994 Elo on that sub-metric versus 1,690 for its predecessor, while presentation quality sits at 1,499, roughly flat with Grok 4.6’s 1,519. AA-Briefcase evaluates whether models can complete realistic professional work end-to-end — from performing the analysis to delivering the requested work product.
The same pattern shows up on GDPval-AA, which measures performance on real-world work tasks like documents, spreadsheets, and slides. Grok 4.7 scores 1,695 Elo there, up 90 points from Grok 4.6 (high)’s 1,605.
A Leap In Coding #
On the Artificial Analysis Coding Agent Index, Grok 4.7 paired with the Grok Build harness scores 56 — up 9 points from Grok 4.6 (xhigh). Among models running in their native harnesses, that ranks fourth, behind only Claude Fable 5.1, GPT-6 Astra, and Claude Opus 5 — and ahead of GPT-5.6 Sol. For a company that has increasingly pitched Grok as a coding model, this is the more strategically important number.
Token Hunger And Tradeoffs #
The gains don’t come free. Grok 4.7 (xhigh) burns through approximately 81,000 output tokens per Intelligence Index task — more than double the 36,000 used by Grok 4.6 (high), and roughly triple the 27,000 used by GPT-6 Astra (max). That is 125% more tokens than its predecessor and 196% more than OpenAI’s flagship. Outside of agentic knowledge work, the model broadly matches Grok 4.6 on the index’s other tasks: it improves on Terminal-Bench 4.0 (+4.5 percentage points) and GDP.pdf (+3.0 p.p.), while regressing on AA-LCR (-3.7 p.p.) and AutomationBench-AA (-1.1 p.p.).
On speed, Artificial Analysis measured answer output at approximately 188 tokens per second for long prompts, with the model averaging about 7.1 minutes per Intelligence Index task.
Elsewhere, the specs are unchanged from Grok 4.6: a 500,000-token context window, and pricing of $2 per million input tokens and $6 per million output tokens, with cache hits discounted to $0.50 per million. Reasoning effort is configurable from low to xhigh, with the evaluation run at the maximum setting.
The Bottom Line #
Grok 4.7 is not a headline-grabbing jump on the composite intelligence score — it still trails three named rivals on that metric — but SpaceXAI will care more about where the model is climbing fastest. Being at the frontier on agentic knowledge work and fourth among native harnesses on coding, at a price well below Anthropic’s flagships, keeps Grok firmly in the conversation for enterprises evaluating models on actual work output rather than benchmark totals. Whether the token appetite at xhigh is a fair price for that capability is a question the efficiency-focused end of the market will be asking.