cd /news/large-language-models/spacexai-releases-grok-4-7-for-codin… · home topics large-language-models article
[ARTICLE · art-136439] src=testingcatalog.com ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

SpaceXAI releases Grok 4.7 for coding and knowledge work

SpaceXAI released Grok 4.7, its most capable model for coding and knowledge work, now available in Cursor, Grok Build, the Grok API, third-party coding harnesses, model routers, and cloud platforms at $2 per million input tokens and $6 per million output tokens, matching Grok 4.6 pricing. Grok 4.7 scored 46.3% on CursorBench 4.0, up from Grok 4.6's 40.4%, and 71.0% on DeepSWE v1.1 at high effort, while its 56.7% HealthBench Professional score trailed GPT-5.6 Sol Max and Fable 5.1 Max. SpaceXAI said the model uses an entirely new safeguard stack, scoring 62.4% on LatchBio's biosafety benchmark and allowing only 3.3% of risky dual-use prompts through on HackerBench v0.3.

by read2 min views1 publishedSep 21, 2026
SpaceXAI releases Grok 4.7 for coding and knowledge work
Image: Testingcatalog (auto-discovered)

SpaceXAI has released Grok 4.7, its most capable model for coding and knowledge work, with access now open in Cursor, Grok Build, the Grok API, third-party coding harnesses, model routers, and cloud platforms. The company positions it as a faster, lower-cost rival to other frontier systems. Pricing starts at $2 per million input tokens and $6 per million output tokens, matching Grok 4.6, while a fast variant offers twice the output speed at twice the price. Grok Build also offers free access to try the model.

Grok 4.7 runs on a new, larger base model trained through a longer reinforcement learning cycle and a harder task mix weighted toward work that can take many hours. It is designed to stay on difficult assignments for longer, check its output more carefully, and manage extended context. Native training on the Grok Bot harness also targets conversational tasks and general knowledge work.

The benchmark results show clear gains over Grok 4.6. Grok 4.7 scored 46.3% on CursorBench 4.0, up from 40.4%, and reached 71.0% on DeepSWE v1.1 at high effort. It also posted 64.0% on EEBench, 1,657 on AA Briefcase v1.1, 38.0% on Terminal-Bench 4.0, and 19.6% on the Harvey Legal Agent Benchmark. Its 56.7% HealthBench Professional score trailed GPT-5.6 Sol Max and Fable 5.1 Max, showing that its lead is not universal. Results from GDPval and AA Briefcase point to stronger document and presentation work across professional tasks for lawyers, nurses, and financial analysts.

Safety is another major part of the release. SpaceXAI says Grok 4.7 uses an entirely new safeguard stack and is its strongest model yet for refusals and jailbreak resistance. It scored 62.4% on LatchBio’s biosafety benchmark and allowed only 3.3% of risky dual-use prompts through on HackerBench v0.3, while rarely rejecting legitimate security work. Select cybersecurity partners are also receiving invite-only access to its red-team capabilities for defense research.

For SpaceXAI, Grok 4.7 ties the Grok model line more closely to its own coding products while extending access across external developer tools and infrastructure. The combination of longer task execution, self-checking, competitive token pricing, and tighter safeguards makes the release a direct push for coding teams and knowledge workers choosing among frontier models.

── more in #large-language-models 4 stories · sorted by recency
x.ai · · #large-language-models
Grok 4.7
── more on @spacexai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/spacexai-releases-gr…] indexed:0 read:2min 2026-09-21 ·