cd /news/artificial-intelligence/cafebench-sol-6-1-is-great-value-but… · home › topics › artificial-intelligence › article
[ARTICLE · art-142797] src=getdot.ai ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

CafeBench: Sol 6.1 is great value, but not quite Opus level

OpenAI's GPT-6.1 Sol at medium reasoning effort placed second on Dot's Cafe Bench with an average profit of +$186,449 at a cost of $13.09, behind Anthropic's Claude Opus 5.5 (medium) at +$226,546, while Anthropic's Claude Sonnet 5.5 lost money at low and medium effort and cost more than Opus 5.5 at high effort (+$83,929, $15.65). Dot reported that higher effort did not increase GPT-6.1 Sol's profit, with the high-effort run averaging +$174,098 at $27.20 and taking 214 minutes, the slowest of any model tested. The Pareto frontier now runs from GPT-6 Luna to GPT-6.1 Sol (low) to Opus 5.5, based on two runs per setting in the same simulation as the prior week.

by read3 min views2 publishedSep 30, 2026
CafeBench: Sol 6.1 is great value, but not quite Opus level
Image: source

All Posts Last week we published Cafe Bench: Dot gets the books of a four-café chain and runs it for a year, and the score is how much more profit it makes. This week OpenAI released GPT-6.1 Sol and Anthropic released Claude Sonnet 5.5, so we ran both at low, medium and high reasoning effort.

GPT-6.1 Sol at medium effort is now second on the board, behind Opus 5.5. Sonnet 5.5, on the other hand, is just bad: at low and medium effort, GPT-6 Luna beats it on both profit and cost. The Pareto frontier now runs from GPT-6 Luna to GPT-6.1 Sol (low) to Opus 5.5.

More effort did not buy more profit from GPT-6.1 Sol. This could be a fluke, but various public benchmarks show the same behavior, which is very interesting. A high-effort run of GPT-6.1 Sol also took about three and a half hours, the slowest of any model we have tested.

Sonnet 5.5, on the other hand, gets a huge boost from medium to high effort, the only setting where it made money. However, it still falls off the Pareto frontier, as it costs more than Opus 5.5 while earning less.

The full board, with this week's models in bold:

Model Year A Year B Average Cost Time
Claude Opus 5.5 (medium) +$236,638 +$216,454 +$226,546 $8.57 22 min
GPT-6.1 Sol (medium) +$184,813 +$188,084 +$186,449 $13.09 98 min
GPT-6.1 Sol (high) +$200,084 +$148,113 +$174,098 $27.20 214 min
GPT-6 Astra (medium) +$129,942 +$184,064 +$157,003 $56.34 132 min
Claude Fable 5.1 (medium) +$75,072 +$204,709 +$139,890 $12.81 27 min
GPT-6.1 Sol (low) +$138,793 +$73,347 +$106,070 $3.97 38 min
GPT-6 Sol (medium) +$106,889 +$90,213 +$98,551 $7.41 28 min
Claude Sonnet 5.5 (high) −$17,574 +$185,433 +$83,929 $15.65 40 min
GPT-6 Astra (low) +$87,107 +$63,128 +$75,117 $22.68 51 min
GPT-5.6 Sol (medium) −$19,374 −$73,116 −$46,245 $9.40 28 min
GPT-6 Luna (medium) −$22,696 −$98,103 −$60,400 $1.86 40 min
Claude Sonnet 5.5 (medium) −$271,768 +$32,737 −$119,516 $3.38 12 min
Claude Sonnet 5.5 (low) +$4,934 −$341,267 −$168,167 $2.17 9 min
GPT-5.6 Luna (medium) −$226,377 −$339,861 −$283,119 $1.11 10 min

Things to keep in mind #

  • Two runs per setting. With Sonnet 5.5's years this far apart, its ranking could move a lot with more runs.
  • Same world as last week. Nothing about the simulation changed, so every row here is directly comparable with thefirst post .

Anand Ani

Anand is a founding AI engineer at Dot. Builder at heart, part-time nerd, and a chess player on the side.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/cafebench-sol-6-1-is…] indexed:0 read:3min 2026-09-30 · —