18:32
2026-07-31
arxiv.org
artificial-intelligence
Orca-Bench: How Ready Are Language Model Agents for Oncall?
A new benchmark, ORCA-bench, shows that frontier language model agents achieve only 25.3% root cause analysis accuracy on Medium-difficulty oncall tasks and 10.0% on Hard tasks, with the best performaβ¦