04:00
2026-09-01
arxiv.org
artificial-intelligence
A collective capability boundary in frontier large language models on guideline-conformant and case-specific oncology decision-making
A new benchmark, the Oncology Decision Boundary Benchmark (ODBB), found that 42.1% of 2,005 oncology decision points were answered incorrectly by all nine frontier large language models evaluated, incβ¦