cd /news/artificial-intelligence/grok-4-6-tops-bbeh-mini-benchmark-sc… · home topics artificial-intelligence article
[ARTICLE · art-97095] src=runtimewire.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Grok 4.6 tops BBEH Mini benchmark, scoring 0.67 on 460 tasks

SpaceXAI's Grok 4.6 topped the BBEH Mini benchmark with a mean score of 0.67 across 460 tasks at an estimated cost of $0.0034 per task, according to a RuntimeWire evaluation. Anthropic's Claude Opus 4.8 followed at 0.62 ($0.0485 per task), and OpenAI's GPT-5.6 Sol Pro scored 0.59 ($0.0089 per task). The benchmark, sourced from Google DeepMind's BBEH (Apache-2.0), scored each model deterministically against ground truth answers.

read1 min views2 publishedAug 14, 2026
Grok 4.6 tops BBEH Mini benchmark, scoring 0.67 on 460 tasks
Image: Runtimewire (auto-discovered)

In a 460-task BBEH Mini evaluation, SpaceXAI’s Grok 4.6 led the field with a 0.67 score at $0.0034 per task. Anthropic’s Claude Opus 4.8 followed at 0.62, while OpenAI’s GPT-5.6 Sol Pro posted 0.59.

By RuntimeWire Staff · Published

Leaderboard

| Rank | Model | Mean score | Est. cost/task | | 1 | SpaceXAI: Grok 4.6 | 0.67 | $0.0034 | | 2 | Anthropic: Claude Opus 4.8 | 0.62 | $0.0485 | | 3 | OpenAI: GPT-5.6 Sol Pro | 0.59 | $0.0089 |

How we scored it

Every model answered the same 460-task battery from BBEH Mini (BBEH — Google DeepMind (Apache-2.0)), one task at a time, with no tools and no retries on content.

460 tasks carry ground truth (exact, numeric, multiple-choice, or the benchmark's official matcher) and were scored deterministically against the reference answer — correct is 1, incorrect is 0.

A model's mean score averages its graded tasks; a generation failure counts as 0. Cost per task is estimated from each model's published per-token pricing ("—" where pricing isn't public), so treat it as directional, not billing-exact.

Explore every prompt, answer, and per-task grade in the interactive leaderboard.

Reader comments #

Conversation for this story loads after sign-in.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @spacexai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/grok-4-6-tops-bbeh-m…] indexed:0 read:1min 2026-08-14 ·