SpaceXAI's latest model scored 56 on a new benchmark for AI cyber defense, matching MiMo-V2.6-Pro and edging past GPT-6 Luna
Grok 4.7 now sits at the top of the Artificial Analysis Cyber Index, a new leaderboard built to measure how well AI agents defend software. It posted a composite score of 56. It shares the top spot with MiMo-V2.6-Pro, so this is less a coronation and more a co-headlining tour.
The timing matters. The index launched on or around September 28, 2026, which makes Grok 4.7 one of the first models to claim a top position on a benchmark aimed squarely at enterprise security work.
How the scoreboard shakes out #
The ranking is tight at the top. Grok 4.7 and MiMo-V2.6-Pro each scored 56 to tie for first place.
GPT-6 Luna landed in third with a score of 53.
The Cyber Index tests AI agents on a set of enterprise cyber defense tasks. Models must find vulnerabilities, reproduce them, and then ship working patches inside real codebases.
Grok 4.7’s composite score rests on strong results in specific sub-benchmarks. On CWE-Bench-AA, it recorded a 68% pass@1 rate, which tied for the lead on that test.
Pass@1 measures whether a model gets the task right on its first attempt. No retries, no do-overs, which is roughly how a security team would want an automated tool to behave.
AI, tech, and the markets they move—in one daily briefing.
Daily. Free. Join 34,000+ readers across crypto, finance, and policy.
On the CyberGym-E2E-AA patching task, Grok 4.7 hit a 74% success rate. That sub-test focuses on the end-to-end job of fixing flaws, so a strong showing there speaks directly to the defensive use case the index was designed around.
The price of being at the top #
Performance came with a hefty bill. Grok 4.7 costs $11.67 per task on the index, a figure considered high next to cheaper competing models.
Elon Musk highlighted the ranking on social media after the results came out.
Background: a new benchmark and a new parent company #
Grok 4.7 comes from SpaceXAI, the entity formed after the merger with xAI. The release follows its predecessor, Grok 4.6, which was benchmarked earlier in September 2026.
The index itself did not launch alone. Artificial Analysis paired it with the Artificial Analysis Cyber Index Alliance, a group that includes IBM and NVIDIA.
What this means #
For enterprises shopping for AI security tools, the results offer a useful but incomplete picture. A tie at 56 tells buyers that Grok 4.7 and MiMo-V2.6-Pro are performing at a similar level on this test, which means cost, integration, and reliability will likely decide which one gets deployed. The $11.67 per-task figure is the number to watch. If organizations are willing to pay a premium for top-tier results, Grok 4.7’s ranking could translate into demand from security teams that prioritize accuracy over budget.
Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our