{"slug": "grok-4-6-tops-bbeh-mini-benchmark-scoring-0-67-on-460-tasks", "title": "Grok 4.6 tops BBEH Mini benchmark, scoring 0.67 on 460 tasks", "summary": "SpaceXAI's Grok 4.6 topped the BBEH Mini benchmark with a mean score of 0.67 across 460 tasks at an estimated cost of $0.0034 per task, according to a RuntimeWire evaluation. Anthropic's Claude Opus 4.8 followed at 0.62 ($0.0485 per task), and OpenAI's GPT-5.6 Sol Pro scored 0.59 ($0.0089 per task). The benchmark, sourced from Google DeepMind's BBEH (Apache-2.0), scored each model deterministically against ground truth answers.", "body_md": "# Grok 4.6 tops BBEH Mini benchmark, scoring 0.67 on 460 tasks\n\n**In a 460-task BBEH Mini evaluation, SpaceXAI’s Grok 4.6 led the field with a 0.67 score at $0.0034 per task. Anthropic’s Claude Opus 4.8 followed at 0.62, while OpenAI’s GPT-5.6 Sol Pro posted 0.59.**\n\nBy [RuntimeWire Staff](/author/runtimewire-staff)\n· Published\n\n### Leaderboard\n\n| Rank |\nModel |\nMean score |\nEst. cost/task |\n| 1 |\nSpaceXAI: Grok 4.6 |\n0.67 |\n$0.0034 |\n| 2 |\nAnthropic: Claude Opus 4.8 |\n0.62 |\n$0.0485 |\n| 3 |\nOpenAI: GPT-5.6 Sol Pro |\n0.59 |\n$0.0089 |\n\n### How we scored it\n\nEvery model answered the same 460-task battery from **BBEH Mini** (BBEH — Google DeepMind (Apache-2.0)), one task at a time, with no tools and no retries on content.\n\n460 tasks carry ground truth (exact, numeric, multiple-choice, or the benchmark's official matcher) and were scored deterministically against the reference answer — correct is 1, incorrect is 0.\n\nA model's **mean score** averages its graded tasks; a generation failure counts as 0. **Cost per task** is estimated from each model's published per-token pricing (\"—\" where pricing isn't public), so treat it as directional, not billing-exact.\n\nExplore every prompt, answer, and per-task grade in the [interactive leaderboard](/benchmarks/bbeh-mini-leaderboard).\n\n## Reader comments\n\nConversation for this story loads after sign-in.", "url": "https://wpnews.pro/news/grok-4-6-tops-bbeh-mini-benchmark-scoring-0-67-on-460-tasks", "canonical_source": "https://runtimewire.com/article/grok-4-6-tops-bbeh-mini-benchmark-scoring-0-67-on-460-tasks", "published_at": "2026-08-14 17:04:14+00:00", "updated_at": "2026-08-14 17:07:44.734747+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-products"], "entities": ["SpaceXAI", "Grok 4.6", "Anthropic", "Claude Opus 4.8", "OpenAI", "GPT-5.6 Sol Pro", "Google DeepMind", "RuntimeWire"], "alternates": {"html": "https://wpnews.pro/news/grok-4-6-tops-bbeh-mini-benchmark-scoring-0-67-on-460-tasks", "markdown": "https://wpnews.pro/news/grok-4-6-tops-bbeh-mini-benchmark-scoring-0-67-on-460-tasks.md", "text": "https://wpnews.pro/news/grok-4-6-tops-bbeh-mini-benchmark-scoring-0-67-on-460-tasks.txt", "jsonld": "https://wpnews.pro/news/grok-4-6-tops-bbeh-mini-benchmark-scoring-0-67-on-460-tasks.jsonld"}}