The latest numbers from Terminal-Bench-Science show a massive gap between the top tier and everyone else. GPT-6 Astra and Claude Opus 5.5 are leading Fable 5.1 by roughly 20 points. To put that in perspective, the strongest model outside those two labs, Qwen3.8 Max, is sitting at just 12%. It is an incredible time to be building with these tools because the performance ceiling is just skyrocketing.
How does Claude Opus 5.5 handle reasoning effort? #
The scaling of reasoning effort on Terminal-Bench-Science is a bit surprising. Opus 5.5 jumps from 24% at "low effort" up to 62% at "xhigh," but then it actually dips to 59% when set to "max." If you are tuning your prompts, avoid the "max" setting; it imposes a minimum reasoning budget that seems to degrade the output slightly.
On the vision side, it is currently ranked as Anthropic’s strongest vision model, beating Fable 5 and GPT-6 Sol, though it still trails GPT-6 Astra. The best part is the efficiency—it is roughly 60% cheaper than Fable 5.1 while leading SimpleBench with an 88.4% score.
What is the current state of the GPT-6 family? #
The different GPT-6 variants are carving out very specific niches:
- Astra: This is the heavy hitter for reviews and audits. It reportedly beat NetHack on its third attempt and maintains an 82.5% win rate in DOOM agent matches.
- Luna [Max]: The efficiency play. It hit #24 in Code Arena WebDev with a score of 1593 (which is +74 over GPT-5.6 Luna) at a blended cost of about $0.40/Mtok.
- Sol: The speed king, though it loses to Luna on the wins-per-dollar metric.
How do the latest Flash and Open-Weight models compare? #
Gemini 3.8 Flash is putting up some wild numbers, hitting 291 tok/s with a 1M context window. It scored 89.2% on ARC-AGI v2 at $0.40 per task and 98.5% on v1. Interestingly, on v3, it hits 10.4% with the standard harness but jumps to 35% using the provider harness.
Then there is the Xiaomi MiMo-V2.6-Pro. It is released under the MIT license, is omni-modal with 1M context, and scores 46 on the AA index. For comparison, GPT-5.6 Sol is at 47. The price difference is where Xiaomi wins—it costs $0.13 compared to $1.99 per task.
It is also worth noting that the $200 Claude Code plan is receiving a lot of praise right now, with a growing consensus that it is officially beating Codex. Whether you are optimizing for raw power with Astra or cost-efficiency with MiMo, there is a perfect tool for every single use case right now.
Next Deepfake Detection Fails Without Cross-Generator Generalization →
All Replies (1) #
Want a live back-and-forth? Join the global AI chat room — login to talk. That 20 point gap is wild. I'm seeing a huge jump in complex logic tasks with Opus 5.5 compared to my old setup.