Grok 4.6 tops Newsroom Reliability v0.2 benchmark at 0.78 SpaceXAI's Grok 4.6 topped the Newsroom Reliability v0.2 benchmark with a mean score of 0.78, outperforming OpenAI's GPT-5.6 Luna Pro (0.77) and GPT-5.6 Sol Pro (0.76), Meta's Muse Spark 1.2 (0.73), Anthropic's Claude Opus 4.8 (0.71), and Google's Gemini 3.7 Flash (0.67). The benchmark evaluated 50 open-ended tasks graded by gpt-5.4, with costs per task ranging from $0.0005 for GPT-5.6 Luna Pro to $0.0215 for GPT-5.6 Sol Pro. Leaderboard | Rank | Model | Mean score | Est. cost/task | | 1 | SpaceXAI: Grok 4.6 | 0.78 | $0.0056 | | 2 | OpenAI: GPT-5.6 Luna Pro | 0.77 | $0.0005 | | 3 | OpenAI: GPT-5.6 Sol Pro | 0.76 | $0.0215 | | 4 | Meta: Muse Spark 1.2 | 0.73 | $0.0044 | | 5 | Anthropic: Claude Opus 4.8 | 0.71 | $0.0211 | | 6 | Google: Gemini 3.7 Flash | 0.67 | $0.0014 | How we scored it Every model answered the same 50-task battery from Newsroom Reliability v0.2 , one task at a time, with no tools and no retries on content. 50 open-ended tasks were graded 0–10 by gpt-5.4 against a fixed rubric — blind to which model wrote the answer — and normalized to the same 0–1 scale. Reasoning effort was pinned to max for every model whose lane exposes a control OpenAI-style reasoning effort , Anthropic extended thinking, Gemini thinking config . Models marked vendor default on the interactive board expose no such control and ran as shipped. A model's mean score averages its graded tasks; a generation failure counts as 0. Cost per task is estimated from each model's published per-token pricing "—" where pricing isn't public , so treat it as directional, not billing-exact. Explore every prompt, answer, and per-task grade in the interactive leaderboard /benchmarks/newsroom-reliability-v0-2-leaderboard-2 .