{"slug": "companies-are-optimizing-models-for-specific-benchmarks", "title": "Companies are optimizing models for specific benchmarks", "summary": "OpenAI is optimizing its GPT-5.6 model for the GPQA Diamond benchmark, while Anthropic is optimizing Opus 5 for the Humanity's Last Exam benchmark, with GPT-5.6 winning on GPQA Diamond and Opus 5 winning on Humanity's Last Exam.", "body_md": "Openai is optimizing for gpqa diamond and anthropic is optimizing for humanity last exam. gpt 5.6 wins on gpqa and opus 5 wins on humanity last exam\n\nComments URL: https://news.ycombinator.com/item?id=49044813\n\nPoints: 1\n\n# Comments: 0", "url": "https://wpnews.pro/news/companies-are-optimizing-models-for-specific-benchmarks", "canonical_source": "https://news.ycombinator.com/item?id=49044813", "published_at": "2026-07-25 05:34:15+00:00", "updated_at": "2026-07-25 05:52:11.718439+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-research", "ai-products"], "entities": ["OpenAI", "Anthropic", "GPT-5.6", "Opus 5", "GPQA Diamond", "Humanity's Last Exam"], "alternates": {"html": "https://wpnews.pro/news/companies-are-optimizing-models-for-specific-benchmarks", "markdown": "https://wpnews.pro/news/companies-are-optimizing-models-for-specific-benchmarks.md", "text": "https://wpnews.pro/news/companies-are-optimizing-models-for-specific-benchmarks.txt", "jsonld": "https://wpnews.pro/news/companies-are-optimizing-models-for-specific-benchmarks.jsonld"}}