Companies are optimizing models for specific benchmarks
OpenAI is optimizing its GPT-5.6 model for the GPQA Diamond benchmark, while Anthropic is optimizing Opus 5 for the Humanity's Last Exam benchmark, with GPT-5.6 winning on GPQA Diamond and Opus 5 winn…