Grok 4.6 tops BBEH Mini benchmark, scoring 0.67 on 460 tasks SpaceXAI's Grok 4.6 topped the BBEH Mini benchmark with a mean score of 0.67 across 460 tasks at an estimated cost of $0.0034 per task, according to a RuntimeWire evaluation. Anthropic's Claude Opus 4.8 followed at 0.62 ($0.0485 per task), and OpenAI's GPT-5.6 Sol Pro scored 0.59 ($0.0089 per task). The benchmark, sourced from Google DeepMind's BBEH (Apache-2.0), scored each model deterministically against ground truth answers. Grok 4.6 tops BBEH Mini benchmark, scoring 0.67 on 460 tasks In a 460-task BBEH Mini evaluation, SpaceXAI’s Grok 4.6 led the field with a 0.67 score at $0.0034 per task. Anthropic’s Claude Opus 4.8 followed at 0.62, while OpenAI’s GPT-5.6 Sol Pro posted 0.59. By RuntimeWire Staff /author/runtimewire-staff · Published Leaderboard | Rank | Model | Mean score | Est. cost/task | | 1 | SpaceXAI: Grok 4.6 | 0.67 | $0.0034 | | 2 | Anthropic: Claude Opus 4.8 | 0.62 | $0.0485 | | 3 | OpenAI: GPT-5.6 Sol Pro | 0.59 | $0.0089 | How we scored it Every model answered the same 460-task battery from BBEH Mini BBEH — Google DeepMind Apache-2.0 , one task at a time, with no tools and no retries on content. 460 tasks carry ground truth exact, numeric, multiple-choice, or the benchmark's official matcher and were scored deterministically against the reference answer — correct is 1, incorrect is 0. A model's mean score averages its graded tasks; a generation failure counts as 0. Cost per task is estimated from each model's published per-token pricing "—" where pricing isn't public , so treat it as directional, not billing-exact. Explore every prompt, answer, and per-task grade in the interactive leaderboard /benchmarks/bbeh-mini-leaderboard . Reader comments Conversation for this story loads after sign-in.