In brief
- xAI released Grok 4.7 on Monday, after at least five delays since late July.
- The model scored second to Claude Fable 5.1 on GDPval and AA-Briefcase, and second to GPT-6 Astra on EEBench, an electrical-engineering benchmark.
- Grok 4.7 is live immediately in the Grok app, Cursor, Grok Build, and the xAI API, running on 2.1 trillion parameters with supplemental training on SpaceX engineering data.
Elon Musk’s xAI released Grok 4.7 on Monday afternoon, its best model to date, calling it "a notable improvement over Grok 4.6 at the same price and speed."
The release comes after several apparent delays. Musk had walked the timeline back at least five times since late July: "four weeks out," then "a few weeks," then "3 to 4 weeks," then "10 days" on September 1, then "needs a few more days to cook" on September 11.
xAI says the model spends longer working through hard problems and double-checks its own answers more often than Grok 4.6 did, alongside what the company calls its strongest safety guardrails yet.
Musk followed up on X, calling Grok 4.7 "a strong combination of intelligence, speed & low cost." There's no waitlist this time—it's live now in the Grok app, Cursor, Grok Build, and the xAI API.
Grok 4.7 is here.
It's a notable improvement over Grok 4.6 at the same price and speed. pic.twitter.com/H3OTBbXyvO
— SpaceXAI (@SpaceXAI) September 21, 2026 Grok 4.7 packs 2.1 trillion parameters, up 40% from the 1.5 trillion in Grok 4.6, itself a refinement of Grok 4.5. The costs are $2 per million input tokens and $6 per million output tokens. Parameters are the internal knobs a model tunes during training, and more of them generally means more capacity to learn patterns, while tokens are the basic amount of information an AI model can either register or generate.
xAI also folded in supplemental training data pulled from SpaceX, Musk's rocket company: Starlink satellite telemetry, manufacturing records, and engineering failure logs. The pitch is a model that reasons better about hardware and physical systems than anything trained purely on internet text.
That said, benchmark scores tell a familiar story. GDPval measures how a model performs on real, economically valuable knowledge work—legal memos, spreadsheets, slide decks—using tasks vetted by working professionals in each field, and scores it as an Elo rating, the same head-to-head ranking system chess uses.
Grok 4.7 hit 1695 on GDPval. Claude Fable 5.1 topped the chart at 1735.
AA-Briefcase, built by Artificial Analysis, tests multi-hour office work that strings research, analysis, and document production into one long task, also scored on the Elo scale. Grok 4.7 posted 1657 against Fable 5.1's 1678. Same result, different test.
CursorBench 4.0, Cursor's benchmark for real coding tasks inside its editor, plots accuracy against the cost and token count each task burns through. Grok 4.7 lands in the middle: pricier per task than GPT-5.6 Sol's successor, GPT-6 Astra, and Claude Sonnet 5, but still short of Fable 5.1, which wins at every price point on the chart.
This isn't a new pattern for xAI. Grok 4.5 launched in July with the biggest training cluster in the industry and third-place scores behind Claude and OpenAI's models. Before it, Grok 4.20 traded reliability for speed and personality. Grok 4.6 also trailed the frontier pack on coding autonomy.
None of this makes Grok 4.7 a bad product for the millions of people who talk to it through X, the standalone app, or their Tesla's dashboard. It means the model most likely to answer your questions, or power your car's voice assistant, is running on a system that its own maker's benchmarks place a rung below the top of the ladder.
That gap is why the price tag matters more than the leaderboard position for most people. xAI has consistently undercut Anthropic and OpenAI on cost per token even as it trails them on raw capability, betting that "good enough, cheap, and everywhere" beats "best, but pricier" for the bulk of everyday use.
Musk had already set expectations lower days before launch, writing that Grok 4.7 should land "roughly on par with" Anthropic's Claude Opus 5.0, not the newer Opus 5.1, with multimodal performance still needing work.
In that same post he sketched the next three models: Grok 4.8 as a meaningful step up, Grok 4.9 in "Astra/Fable class," and Grok 5 as a possible frontier leader. Those upgrades don’t have a release date yet. “We shall see,” he wrote.