# Claude Opus 5.5 and GPT-6 Astra are dominating the latest benchmarks

> Source: <https://promptcube3.com/en/threads/9624/>
> Published: 2026-09-26 21:02:28+00:00

# Claude Opus 5.5 and GPT-6 Astra are dominating the latest benchmarks

The latest numbers from Terminal-Bench-Science show a massive gap between the top tier and everyone else. GPT-6 Astra and [Claude](https://promptcube3.com/en/tags/claude/) Opus 5.5 are leading Fable 5.1 by roughly 20 points. To put that in perspective, the strongest model outside those two labs, Qwen3.8 Max, is sitting at just 12%. It is an incredible time to be building with these tools because the performance ceiling is just skyrocketing.

## How does Claude Opus 5.5 handle reasoning effort?

The scaling of reasoning effort on Terminal-Bench-Science is a bit surprising. Opus 5.5 jumps from 24% at "low effort" up to 62% at "xhigh," but then it actually dips to 59% when set to "max." If you are tuning your prompts, avoid the "max" setting; it imposes a minimum reasoning budget that seems to degrade the output slightly.

On the vision side, it is currently ranked as Anthropic’s strongest vision model, beating Fable 5 and GPT-6 Sol, though it still trails GPT-6 Astra. The best part is the efficiency—it is roughly 60% cheaper than Fable 5.1 while leading SimpleBench with an 88.4% score.

## What is the current state of the GPT-6 family?

The different GPT-6 variants are carving out very specific niches:

- **Astra:** This is the heavy hitter for reviews and audits. It reportedly beat NetHack on its third attempt and maintains an 82.5% win rate in DOOM agent matches.
- **Luna [Max]:** The efficiency play. It hit #24 in Code Arena WebDev with a score of 1593 (which is +74 over GPT-5.6 Luna) at a blended cost of about $0.40/Mtok.
- **Sol:** The speed king, though it loses to Luna on the wins-per-dollar metric.

## How do the latest Flash and Open-Weight models compare?

[Gemini](https://promptcube3.com/en/tags/gemini/) 3.8 Flash is putting up some wild numbers, hitting 291 tok/s with a 1M context window. It scored 89.2% on ARC-AGI v2 at $0.40 per task and 98.5% on v1. Interestingly, on v3, it hits 10.4% with the standard harness but jumps to 35% using the provider harness.

Then there is the Xiaomi MiMo-V2.6-Pro. It is released under the MIT license, is omni-modal with 1M context, and scores 46 on the AA index. For comparison, GPT-5.6 Sol is at 47. The price difference is where Xiaomi wins—it costs $0.13 compared to $1.99 per task.

It is also worth noting that the $200 [Claude Code](https://promptcube3.com/en/tags/claude%20code/) plan is receiving a lot of praise right now, with a growing consensus that it is officially beating Codex. Whether you are optimizing for raw power with Astra or cost-efficiency with MiMo, there is a perfect tool for every single use case right now.

[Next Deepfake Detection Fails Without Cross-Generator Generalization →](https://promptcube3.com/en/threads/9585/)

## All Replies （1）

Want a live back-and-forth? [Join the global AI chat room](https://promptcube3.com/en/chat/) — login to talk.

That 20 point gap is wild. I'm seeing a huge jump in complex logic tasks with Opus 5.5 compared to my old setup.
