Companies are optimizing models for specific benchmarks OpenAI is optimizing its GPT-5.6 model for the GPQA Diamond benchmark, while Anthropic is optimizing Opus 5 for the Humanity's Last Exam benchmark, with GPT-5.6 winning on GPQA Diamond and Opus 5 winning on Humanity's Last Exam. Openai is optimizing for gpqa diamond and anthropic is optimizing for humanity last exam. gpt 5.6 wins on gpqa and opus 5 wins on humanity last exam Comments URL: https://news.ycombinator.com/item?id=49044813 Points: 1 Comments: 0