The larger model targets longer coding and knowledge-work agents, with API rates of $2 input and $6 output per million tokens.
By [Ryan Merket](https://runtimewire.com/author/ryan-merket)
· Published
Primary source: [SpaceXAI on X](https://x.com/spacexai/status/2102069815225586149?s=46)
Why it matters #
SpaceXAI is using Grok 4.7 to compete on completed-task economics: a larger model, unchanged token rates and immediate distribution through Cursor, Grok Build and the API.
SpaceXAI (@SpaceXAI), the AI operation inside Elon Musk (@elonmusk)'s SpaceX, launched Grok 4.7 on September 21st, pitching a larger base model that can stay on difficult coding and knowledge-work assignments longer without raising its standard token prices.
The release arrived 40 days after Grok 4.6, another quick turn in Musk's push to make model development part of SpaceX's broader infrastructure operation. SpaceX acquired xAI on February 2nd, and its June prospectus described the acquisition as the foundation of a newly formed AI segment. That filing said SpaceX planned to allocate substantial capital to compute infrastructure while integrating xAI's personnel, commercial products and data center operations.
Grok 4.7 is the clearest product result of that strategy so far: spend heavily on training, ship models directly into widely used coding surfaces and compete on the cost of completing long jobs rather than on chatbot subscriptions alone.
A larger model for longer jobs
In its Grok 4.7 launch post, SpaceXAI said the model uses a larger base than Grok 4.6 and received a longer reinforcement-learning run weighted toward problems that take hours to finish. SpaceXAI also trained Grok 4.7 to understand the Grok Bot agent harness natively, an effort aimed at improving conversations, research, document production and other work outside software repositories.
SpaceXAI's developer documentation lists a 500,000-token context window and a May 2026 knowledge cutoff. Grok 4.7 accepts text and images, produces text and supports function calling, web search, X search and code execution. Developers can select low, medium, high or xhigh reasoning effort.
Standard API pricing remains $2 per million input tokens and $6 per million output tokens, matching Grok 4.6. SpaceXAI also offers a fast version at twice those rates and says it provides twice the output speed. The fast version is limited to Cursor and Grok Build rather than the public API.
Grok 4.7 is available through the SpaceXAI API, Cursor, Grok Build, OpenRouter, Vercel and Cloudflare. It becomes the default model in Grok Build, while Cursor is making it available across its plans. SpaceXAI also offers a US regional API endpoint carrying a 10% token-price premium.
The benchmark case rests on price
SpaceXAI's strongest argument is cost-adjusted performance. Its launch table gives Grok 4.7 an xhigh score of 46.3% on CursorBench 4.0, up from Grok 4.6's 40.4% at high effort. Those settings differ, limiting the usefulness of the direct comparison.
Cursor's own September results provide a cleaner high-effort comparison. Cursor scored Grok 4.7 at 43.9%, against 40.4% for Grok 4.6, and calculated an average task cost of $4.69 for Grok 4.7 versus $5.20 for Grok 4.6. Fable 5.1 and Opus 5 scored higher, at 49.2% and 44.7%, but cost about $9 per task in Cursor's test. Cursor cautions that small score differences may fall within normal evaluation variance.
The pattern carries into SpaceXAI's other reported tests. Grok 4.7 scored 71% on DeepSWE v1.1 at high effort, close to GPT-5.6 Sol's 72.7% and above the 70% result listed for Fable 5.1. SpaceXAI reported larger gains over Grok 4.6 on Terminal-Bench 4.0, electrical engineering tasks and the Harvey legal-agent benchmark. The results come from different harnesses and reasoning settings, so they establish a competitive range rather than a single model ranking.
SpaceXAI puts safeguards into the release pitch
SpaceXAI says Grok 4.7 uses an entirely new safeguard stack and is its strongest model for jailbreak resistance and calibrated refusals. According to SpaceXAI's evaluation, Grok 4.7 scored 62.4% on LatchBio's biosafety benchmark. SpaceXAI also reported that the model allowed 3.3% of risky dual-use prompts through on HackerBench v0.3, SpaceXAI's own cyber-safety test.
Those figures remain SpaceXAI-reported results. They extend the safety work around Grok 4.6, which LatchBio separately found performed strongly at distinguishing disguised hazardous biology requests from legitimate research tasks.
Grok 4.7 enters a market where model providers increasingly ship inside coding tools on release day and charge according to the tokens consumed by autonomous work. SpaceXAI has matched that distribution from the start. The larger bet is that better self-checking and longer task persistence will cut retries enough to matter more than the unchanged token price. Cursor's early cost-per-task result supports that case, although production workloads will determine whether the advantage survives outside controlled evaluations.