{"slug": "cafebench-sol-6-1-is-great-value-but-not-quite-opus-level", "title": "CafeBench: Sol 6.1 is great value, but not quite Opus level", "summary": "OpenAI's GPT-6.1 Sol at medium reasoning effort placed second on Dot's Cafe Bench with an average profit of +$186,449 at a cost of $13.09, behind Anthropic's Claude Opus 5.5 (medium) at +$226,546, while Anthropic's Claude Sonnet 5.5 lost money at low and medium effort and cost more than Opus 5.5 at high effort (+$83,929, $15.65). Dot reported that higher effort did not increase GPT-6.1 Sol's profit, with the high-effort run averaging +$174,098 at $27.20 and taking 214 minutes, the slowest of any model tested. The Pareto frontier now runs from GPT-6 Luna to GPT-6.1 Sol (low) to Opus 5.5, based on two runs per setting in the same simulation as the prior week.", "body_md": "[All Posts](https://www.getdot.ai/blog)\n\n# Cafe Bench: GPT-6.1 Sol is incredible value, but not as good as Opus\n\nLast week we published [Cafe Bench](https://www.getdot.ai/blog/cafe-bench): Dot gets the books of a four-café chain and runs it for a year, and the score is how much more profit it makes. This week OpenAI released GPT-6.1 Sol and Anthropic released Claude Sonnet 5.5, so we ran both at low, medium and high reasoning effort.\n\nGPT-6.1 Sol at medium effort is now second on the board, behind Opus 5.5. Sonnet 5.5, on the other hand, is just bad: at low and medium effort, GPT-6 Luna beats it on both profit and cost. The Pareto frontier now runs from GPT-6 Luna to GPT-6.1 Sol (low) to Opus 5.5.\n\nMore effort did not buy more profit from GPT-6.1 Sol. This could be a fluke, but various public benchmarks show the same behavior, which is very interesting. A high-effort run of GPT-6.1 Sol also took about three and a half hours, the slowest of any model we have tested.\n\nSonnet 5.5, on the other hand, gets a huge boost from medium to high effort, the only setting where it made money. However, it still falls off the Pareto frontier, as it costs more than Opus 5.5 while earning less.\n\nThe full board, with this week's models in bold:\n\n| Model | Year A | Year B | Average | Cost | Time | \n|---|---|---|---|---|---|\n| Claude Opus 5.5 (medium) | +$236,638 | +$216,454 | +$226,546 | $8.57 | 22 min | \n| **GPT-6.1 Sol** (medium) | +$184,813 | +$188,084 | +$186,449 | $13.09 | 98 min | \n| **GPT-6.1 Sol** (high) | +$200,084 | +$148,113 | +$174,098 | $27.20 | 214 min | \n| GPT-6 Astra (medium) | +$129,942 | +$184,064 | +$157,003 | $56.34 | 132 min | \n| Claude Fable 5.1 (medium) | +$75,072 | +$204,709 | +$139,890 | $12.81 | 27 min | \n| **GPT-6.1 Sol** (low) | +$138,793 | +$73,347 | +$106,070 | $3.97 | 38 min | \n| GPT-6 Sol (medium) | +$106,889 | +$90,213 | +$98,551 | $7.41 | 28 min | \n| **Claude Sonnet 5.5** (high) | −$17,574 | +$185,433 | +$83,929 | $15.65 | 40 min | \n| GPT-6 Astra (low) | +$87,107 | +$63,128 | +$75,117 | $22.68 | 51 min | \n| GPT-5.6 Sol (medium) | −$19,374 | −$73,116 | −$46,245 | $9.40 | 28 min | \n| GPT-6 Luna (medium) | −$22,696 | −$98,103 | −$60,400 | $1.86 | 40 min | \n| **Claude Sonnet 5.5** (medium) | −$271,768 | +$32,737 | −$119,516 | $3.38 | 12 min | \n| **Claude Sonnet 5.5** (low) | +$4,934 | −$341,267 | −$168,167 | $2.17 | 9 min | \n| GPT-5.6 Luna (medium) | −$226,377 | −$339,861 | −$283,119 | $1.11 | 10 min | \n\n## Things to keep in mind\n\n- **Two runs per setting.** With Sonnet 5.5's years this far apart, its ranking could move a lot with more runs.\n- **Same world as last week.** Nothing about the simulation changed, so every row here is directly comparable with the[first post](https://www.getdot.ai/blog/cafe-bench) .\n\nAnand Ani\n\nAnand is a founding AI engineer at Dot. Builder at heart, part-time nerd, and a chess player on the side.", "url": "https://wpnews.pro/news/cafebench-sol-6-1-is-great-value-but-not-quite-opus-level", "canonical_source": "https://www.getdot.ai/blog/cafe-bench-gpt-6-1-sol-sonnet-5-5", "published_at": "2026-09-30 19:57:00+00:00", "updated_at": "2026-09-30 20:19:48.344730+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research"], "entities": ["OpenAI", "GPT-6.1 Sol", "Anthropic", "Claude Sonnet 5.5", "Claude Opus 5.5", "Dot", "Cafe Bench", "GPT-6 Luna"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/cafebench-sol-6-1-is-great-value-but-not-quite-opus-level", "markdown": "https://wpnews.pro/news/cafebench-sol-6-1-is-great-value-but-not-quite-opus-level.md", "text": "https://wpnews.pro/news/cafebench-sol-6-1-is-great-value-but-not-quite-opus-level.txt", "jsonld": "https://wpnews.pro/news/cafebench-sol-6-1-is-great-value-but-not-quite-opus-level.jsonld"}}