# Grok 4.6 tops BBEH Mini benchmark, scoring 0.67 on 460 tasks

> Source: <https://runtimewire.com/article/grok-4-6-tops-bbeh-mini-benchmark-scoring-0-67-on-460-tasks>
> Published: 2026-08-14 17:04:14+00:00

# Grok 4.6 tops BBEH Mini benchmark, scoring 0.67 on 460 tasks

**In a 460-task BBEH Mini evaluation, SpaceXAI’s Grok 4.6 led the field with a 0.67 score at $0.0034 per task. Anthropic’s Claude Opus 4.8 followed at 0.62, while OpenAI’s GPT-5.6 Sol Pro posted 0.59.**

By [RuntimeWire Staff](/author/runtimewire-staff)
· Published

### Leaderboard

| Rank |
Model |
Mean score |
Est. cost/task |
| 1 |
SpaceXAI: Grok 4.6 |
0.67 |
$0.0034 |
| 2 |
Anthropic: Claude Opus 4.8 |
0.62 |
$0.0485 |
| 3 |
OpenAI: GPT-5.6 Sol Pro |
0.59 |
$0.0089 |

### How we scored it

Every model answered the same 460-task battery from **BBEH Mini** (BBEH — Google DeepMind (Apache-2.0)), one task at a time, with no tools and no retries on content.

460 tasks carry ground truth (exact, numeric, multiple-choice, or the benchmark's official matcher) and were scored deterministically against the reference answer — correct is 1, incorrect is 0.

A model's **mean score** averages its graded tasks; a generation failure counts as 0. **Cost per task** is estimated from each model's published per-token pricing ("—" where pricing isn't public), so treat it as directional, not billing-exact.

Explore every prompt, answer, and per-task grade in the [interactive leaderboard](/benchmarks/bbeh-mini-leaderboard).

## Reader comments

Conversation for this story loads after sign-in.
