The results are predictable: GPT-5.6 (or its equivalent iteration) takes the lead on GPQA, but Opus 5 dominates the Humanity Last metrics. This creates a fragmented landscape where "the best model" depends entirely on which specific test you value most.
This raises a real concern for anyone building an AI workflow. If a model is over-fitted to a benchmark, its real-world performance might not actually match those high scores. When we see these leaps in performance on paper, it's often just the result of targeted optimization rather than a fundamental breakthrough in reasoning.
For those of us doing actual prompt engineering, the takeaway is to trust your own internal evals over the marketing slides. A model that scores 90% on a specialized exam might still hallucinate on a basic deployment task in your specific codebase.
[Brolly: My minimalist weather workflow 2h ago](/en/news/3358/)
[Trump's Plane Switch: Security Implications 2h ago](/en/news/3350/)
Anthropic's recruitment strategy isn't enough to sway everyone 3h ago
[Apple's AI Strategy: Why Hardware Integration Wins 3h ago](/en/news/3316/)
[Stop Pretending to Be Human: System Prompt Guide 4h ago](/en/news/3306/)
[Philosophers vs. Anthropic: The AI Industry's Blind Spot 5h ago](/en/news/3282/)
[Next Brolly: My minimalist weather workflow →](/en/news/3358/)