{"slug": "incident-arena-getting-agents-to-the-last-nine-of-reliability", "title": "Incident-Arena: Getting agents to the last nine of reliability", "summary": "Researchers introduced Incident-Arena, a human-built benchmark of 20 tasks that deploy real-world open source production applications to ephemeral Kubernetes clusters, inject faults from the config layer through underlying images, and apply sustained load profiles, per the arXiv paper 2610.00648v1. Frontier models scored below 64.3% across the 20 tasks and 3 application substrates, with agent trials averaging 2.81M tokens and 41 turns, and failures spanning diagnosis and localization errors, incomplete repairs, and unsafe regressions. The benchmark targets agentic site-reliability-engineering (SRE), a field the authors say existing benchmarks limit through toy repositories, non-standard framework implementations, and simple static verifiers.", "body_md": "arXiv:2610.00648v1 Announce Type: new \nAbstract: AI coding agents are ubiquitous in engineering workflows amongst industry and academia. Yet, despite their use in app coding, relatively less attention has been paid to their ability to execute on production incident response. This emerging field, termed agentic site-reliability-engineering (SRE) contains benchmarks limited by (1) unrealistic environments, typically toy repositories (2) non-standard framework implementations and (3) simple static verifiers. We introduce Incident-Arena, a human-built benchmark of 20 carefully selected tasks grounded in real-world deployed open source software. Each task deploys a production application to an ephemeral Kubernetes cluster, injecting a fault from the config layer through underlying images, and a sustained load profile given the task requirements. We also present a novel verification method, going beyond static checks to functional verifiers, holding systems level metrics stable, while ensuring repairs are done safely. Agent trials run an average of 2.81M tokens and 41 turns, going beyond existing benchmarks, demonstrating agentic long horizon reasoning. Across 20 tasks and 3 application substrates, frontier models score below 64.3%, with failures extending from diagnosis/localization errors, through incomplete repairs and unsafe regressions.", "url": "https://wpnews.pro/news/incident-arena-getting-agents-to-the-last-nine-of-reliability", "canonical_source": "https://www.machinebrief.com/news/incident-arena-getting-agents-to-the-last-nine-of-reliabilit-r74t", "published_at": "2026-10-02 04:00:00+00:00", "updated_at": "2026-10-02 04:46:03.438397+00:00", "lang": "en", "topics": ["ai-agents", "ai-safety", "mlops", "ai-research", "artificial-intelligence"], "entities": ["Incident-Arena", "arXiv", "Kubernetes"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/incident-arena-getting-agents-to-the-last-nine-of-reliability", "markdown": "https://wpnews.pro/news/incident-arena-getting-agents-to-the-last-nine-of-reliability.md", "text": "https://wpnews.pro/news/incident-arena-getting-agents-to-the-last-nine-of-reliability.txt", "jsonld": "https://wpnews.pro/news/incident-arena-getting-agents-to-the-last-nine-of-reliability.jsonld"}}