Even as Google has seemingly dropped out of the frontier AI race, Anthropic and OpenAI seem to have some fresh competition.
SpaceXAI has rolled out Grok 4.6, the latest update to its flagship model line, and the company is positioning it as a direct rival to the current top-tier models from Anthropic and OpenAI. Unlike the jump from Grok 4 to Grok 4.5, this release is being pitched less as a raw intelligence upgrade and more as a model built for sticking with long, multi-step jobs — the kind of extended coding sessions, research tasks, or app-building workflows that have become the real battleground for frontier labs this year.
According to SpaceXAI, Grok 4.6 is tuned to stay on task across long agentic runs, whether that means digging through a codebase, researching an unfamiliar topic over many steps, or taking a rough product idea all the way to a working, polished piece of software. The company says the model has also started showing more self-checking behaviour on longer tasks, verifying its own outputs before moving forward rather than just producing a single pass and stopping.
What’s new under the hood #
xAI says Grok 4.6 went through a longer secondary training phase than its predecessor, pulling in curated, model-generated data aimed at reasoning and technical depth, along with engineering-heavy datasets and tweaks to the optimizer and training pipeline. That gave the fine-tuning and reinforcement learning stages that followed a stronger base to build on.
For the supervised fine-tuning stage, xAI used Grok 4.5 to regenerate training trajectories across different reasoning efforts, agent setups, and domains including STEM, software engineering, and general knowledge work, filtering out low-quality examples using automated checks. On the reinforcement learning side, the model was trained across a broad spread of agentic tasks — general coding, knowledge work, and more specialised environments like kernel optimisation, web development, and CAD work. xAI also claims the model produces noticeably stronger first attempts at visual and interactive projects compared to Grok 4.5, often nailing the overall structure and design language of an app in a single pass rather than needing several rounds of back-and-forth.
Grok 4.6 Benchmarks #
xAI put Grok 4.6 up against OpenAI’s GPT-5.6 Sol and Anthropic’s Fable 5 across a set of coding and agentic benchmarks, and the numbers land Grok 4.6 in the same tier as both rather than clearly ahead or behind.
On the Artificial Analysis Intelligence Index — a composite score built from nine separate benchmarks — Grok 4.6 posted a 61, tying GPT-5.6 Sol Max exactly and landing just one point behind Fable 5 Max’s 62. Grok 4.5 High, by comparison, scored 56 on the same index.
The pattern holds across most of the other evals xAI published:
GDPVal-AA v2: Grok 4.6 led the pack here with 1753, ahead of Fable 5 Max (1741) and GPT-5.6 Sol Max (1728).CursorBench v3.2: Grok 4.6 scored 69.9%, just behind Fable 5 Max’s 70.5% and ahead of GPT-5.6 Sol Max’s 67.2%.DeepSWE v1.1: This was Grok 4.6’s weaker showing relatively speaking, at 65.9%, behind both GPT-5.6 Sol Max (73%) and Fable 5 Max (70%).FrontierCode v1.1 (Extended): Grok 4.6 hit 61.3%, trailing Fable 5 Max (64.9%) but ahead of GPT-5.6 Sol Max (60.6%).APEX-Agents: Grok 4.6 scored 57.5%, sitting between GPT-5.6 Sol Max (56.7%) and Fable 5 Max (59.2%).Terminal-Bench v3.0: Grok 4.6 clearly lagged here at 26%, well behind GPT-5.6 Sol Max (34.6%) and Fable 5 Max (34.1%).APEX-SWE: Grok 4.6 scored 56.4%, behind Fable 5 Max’s 58.8%; xAI did not list a comparable GPT-5.6 Sol Max figure.AA-Briefcase: Grok 4.6 topped this eval too, at 1577, narrowly ahead of Fable 5 Max (1574) and GPT-5.6 Sol Max (1502).Harvey LAB (Vals): Grok 4.6 led comfortably with 15.8%, well ahead of GPT-5.6 Sol Max’s 2.5% and Fable 5 Max’s 11.3%.
Across the full set, Grok 4.6 beats Grok 4.5 High on every single benchmark by a wide margin, and it’s genuinely competitive with GPT-5.6 Sol Max and Fable 5 Max — leading on a few evals, trailing on others, but never falling far behind. It’s worth noting xAI is self-reporting all of this, and it says competitor numbers are pulled from published system cards or public leaderboards rather than xAI running the rival models itself.
Grok 4.6 Pricing: Undercutting the competition #
Where Grok 4.6 makes its clearest pitch is on cost. xAI has priced the model at $2 per million input tokens and $6 per million output tokens — a price point the company is framing as roughly half of what comparable frontier models charge. There’s also a faster variant available at double that price for users who want lower latency.
Availability #
Grok 4.6 is live starting today inside Cursor and xAI’s own Grok Build, as well as through the API and third-party platforms including OpenRouter, Vercel, and Cloudflare. To get people trying the new model, xAI is offering 2x the usual included usage inside both Grok Build and Cursor for the first week after launch.
xAI says the model’s safety systems have also been recalibrated to match its expanded capabilities, with what the company describes as its most extensive round of pre-deployment safety testing to date, alongside ongoing post-deployment and third-party evaluation.