SpaceXAI Launches Grok 4.7, Beats GPT-5.6 Sol And Fable 5.1 On Some Benchmarks SpaceXAI launched Grok 4.7, priced at $2 per million input tokens and $6 per million output tokens, claiming it beats OpenAI's GPT-5.6 Sol on five of seven benchmarks and Anthropic's Fable 5.1 on three. SpaceXAI says Grok 4.7 is "twice as fast, at half the price of comparable models," with input tokens costing half of GPT-5.6 Sol's $4 and a fifth of Fable 5.1's $10, while its $6 output price compares with $20 for Sol and $50 for Fable 5.1. The model scores 1,695 Elo on GDPval, up from Grok 4.6's 1,605 but behind Fable 5.1's 1,735 and ahead of GPT-6 Astra's 1,542. SpaceXAI has launched Grok 4.7, the latest version of its flagship model, which the company is pitching as its most capable yet for coding and knowledge work. Priced at $2 per million input tokens and $6 per million output tokens, the same as Grok 4.6 https://officechai.com/ai/grok-4-6-benchmarks/ , the model beats OpenAI’s GPT-5.6 Sol https://officechai.com/ai/openai-launches-gpt-5-6-sol-beats-mythos-on-terminalbench/ and Anthropic’s Fable 5.1 https://officechai.com/ai/fable-5-1-benchmarks/ on a few benchmarks, while costing a fraction of what either does. SpaceXAI is calling it “twice as fast, at half the price of comparable models.” The release comes just over a month after Grok 4.6, and roughly two months after Grok 4.5 https://officechai.com/ai/spacexai-and-cursor-release-grok-4-5-beats-opus-4-8-and-gpt-5-5-on-some-benchmarks/ . Per SpaceXAI’s announcement, Grok 4.7 is built on a new, larger base model and was trained with a longer reinforcement learning run on a harder mix of tasks, weighted towards problems that take many hours to complete. The company says the model is better at verifying its own work and managing long contexts, and that it was also trained to natively understand the harness used by Grok Bot, SpaceXAI’s agent product, which should make it better at conversational tasks and general knowledge work. Grok 4.7 Benchmarks SpaceXAI compared Grok 4.7 at xHigh effort with Grok 4.6 High , GPT-5.6 Sol Max and Fable 5.1 Max . In the list below, the scores for Grok 4.7 are followed by those for Grok 4.6, GPT-5.6 Sol and Fable 5.1, in that order: - CursorBench 4.0 software engineering : 46.3% vs 40.4%, 41.7% and 51.8% - DeepSWE v1.1 software engineering : 71.0% vs 65.2%, 72.7% and 70.0% Grok 4.7’s score is at high effort - EEBench electrical engineering : 64.0% vs 53.0%, 39.4% and 56.4% - AA Briefcase v1.1 multi-hour office work : 1,657 vs 1,546, 1,487 and 1,678 - Terminal-Bench 4.0 multi-hour terminal work : 38.0% vs 20.3%, 37.3% and 57.9% - Harvey Legal Agent Benchmark legal work : 19.6% vs 15.8%, 2.5% and 6.7% - HealthBench Professional clinical reasoning : 56.7% vs 48.5%, 60.5% and 62.1% Grok 4.7 comes out ahead of GPT-5.6 Sol on five of the seven benchmarks, falling behind only on DeepSWE and HealthBench Professional. Against Fable 5.1 the picture is more mixed: Grok 4.7 wins on DeepSWE, EEBench and the Harvey legal benchmark, but Fable 5.1 still leads on CursorBench, AA Briefcase, Terminal-Bench and HealthBench Professional. The Terminal-Bench gap is the largest, with Fable 5.1 scoring almost 20 points higher. Even so, Grok 4.7’s Terminal-Bench score is nearly double Grok 4.6’s, and the model improves on its predecessor on every benchmark in the table. On GDPval, which tests AI on tasks done by professionals such as lawyers, nurses and financial analysts, Grok 4.7 scores 1,695 Elo, up from Grok 4.6’s 1,605. That’s behind Fable 5.1’s 1,735, but ahead of the 1,542 posted by OpenAI’s newer GPT-6 Astra https://officechai.com/ai/gpt-6-astra-benchmarks/ . Where it makes the most sense: price Grok 4.7’s pricing is where the gap with rivals is starkest. Its input tokens cost half of GPT-5.6 Sol’s $4 and a fifth of Fable 5.1’s $10, while its $6 output price compares with $20 for Sol and $50 for Fable 5.1 — less than a third and about an eighth respectively. SpaceXAI also plotted CursorBench 4.0 scores against the average cost of completing a task, and says Grok 4.7 is at the frontier of price-performance. Going by the chart, Grok 4.7’s top score of around 46% comes at roughly $6 per task, and its curve sits at or above the other models’ at lower spend levels. That includes Claude Opus 5 https://officechai.com/ai/claude-opus-5-benchmarks/ , which appears to need roughly twice as much per task to reach a similar score, as well as GPT-5.6 Sol and Claude Sonnet 5. Fable 5.1 does pull ahead at higher budgets, hitting 51.8% at a cost of around $17 per task, nearly three times as much. Safety and cybersecurity SpaceXAI says Grok 4.7 was built with an entirely new safeguard stack, and calls it the strongest model it has tested on refusals and jailbreak resistance. In dual-use areas like cybersecurity and biology, the company says the model leads on both usefulness for benign tasks and safe refusals for dangerous ones. It tops LatchBio’s biosafety benchmark at 62.4%, and on HackerBench v0.3, SpaceXAI’s own benchmark for risky and malicious cyber tasks, it lets through only 3.3% of risky dual-use prompts while rarely blocking legitimate security work. The company has also begun giving select cybersecurity partners invite-only access to Grok 4.7’s red-team capabilities for defence research. Availability Grok 4.7 is available now in Cursor and Grok Build, and through the Grok API, third-party coding harnesses, and model routers and cloud platforms. A fast variant with twice the output speed is available at twice the price. A few caveats As with most model launches, the numbers come from the company itself. CursorBench, the benchmark SpaceXAI leads with, is built by Cursor, which SpaceX finished acquiring last month. And while Grok 4.7 is compared with GPT-5.6 Sol, OpenAI’s newer GPT-6 Astra, which launched earlier this month and is priced well above Sol, appears only in the GDPval chart. Independent evaluations of the model will show how these results hold up outside SpaceXAI’s own testing.