{"slug": "orca-bench-how-ready-are-language-model-agents-for-oncall", "title": "Orca-Bench: How Ready Are Language Model Agents for Oncall?", "summary": "A new benchmark, ORCA-bench, shows that frontier language model agents achieve only 25.3% root cause analysis accuracy on Medium-difficulty oncall tasks and 10.0% on Hard tasks, with the best performance coming from Claude Fable 5, according to a paper submitted to arXiv on July 30, 2026. The benchmark, which pairs a live OpenTelemetry-instrumented microservice system with 1,079 RCA tasks, reveals that the weakest model hallucinates an implausible root cause in 40% of incident reports, and removing source-code access degrades every metric, indicating a significant gap before agents can be trusted with production reliability.", "body_md": "# Computer Science > Computation and Language\n\n[Submitted on 30 Jul 2026]\n\n# Title:ORCA-bench: How Ready Are Language Model Agents for Oncall?\n\n[View PDF](/pdf/2607.28545)\n\n[HTML (experimental)](https://arxiv.org/html/2607.28545v1)\n\nAbstract:Large language models can write, patch, and search code, but oncall root cause analysis (RCA) demands something different: reasoning over noisy metrics, logs, traces, and source code, starting from ambiguous user-facing reports, often hours after the incident began. We introduce ORCA-bench, a benchmark that puts general-purpose coding agents in a production-fidelity oncall setting. ORCA-bench pairs a live OpenTelemetry-instrumented microservice system--exposing six days of metrics, logs, and traces through real telemetry interfaces (Prometheus, Jaeger, and OpenSearch via Grafana) and full source-code access--with 1,079 RCA tasks that systematically vary report specificity, time-to-detection, and co-occurring fault scenarios. Ground-truth symptoms are curated and signed off by expert SREs, and our LLM-as-judge is independently re-scored by humans (Cohen's $\\kappa_w=0.90$). Across five frontier agents, the best RCA Accuracy is 25.3% on Medium-difficulty tasks (the realistic-input setting) and 10.0% on Hard--a gap that remains even with Claude Fable 5. The weakest model hallucinates an implausible root cause in 40% of incident reports, and removing source-code access degrades every metric. Crucially, these are performances on a curated 50 GB / six-day testbed with tasks investigated in isolation on a system whose code and instrumentation are public. Since real production systems are order of magnitudes larger, more dynamic, and more idiosyncratic, the gap we report is a lower bound on the engineering investment required before frontier coding agents can be safely entrusted with production reliability. We release the public set at[this https URL].\n\n### Current browse context:\n\ncs.CL\n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer\n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers\n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps\n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations\n\n*(*[What are Smart Citations?](https://www.scite.ai/))# Code, Data and Media Associated with this Article\n\nalphaXiv\n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers\n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub\n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub\n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face\n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast\n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower\n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender\n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [ Learn more about arXivLabs](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/orca-bench-how-ready-are-language-model-agents-for-oncall", "canonical_source": "https://arxiv.org/abs/2607.28545", "published_at": "2026-07-31 18:32:43+00:00", "updated_at": "2026-07-31 18:52:56.326824+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-agents", "ai-research"], "entities": ["ORCA-bench", "OpenTelemetry", "Prometheus", "Jaeger", "OpenSearch", "Grafana", "Claude Fable 5", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/orca-bench-how-ready-are-language-model-agents-for-oncall", "markdown": "https://wpnews.pro/news/orca-bench-how-ready-are-language-model-agents-for-oncall.md", "text": "https://wpnews.pro/news/orca-bench-how-ready-are-language-model-agents-for-oncall.txt", "jsonld": "https://wpnews.pro/news/orca-bench-how-ready-are-language-model-agents-for-oncall.jsonld"}}