ITBench-AA: Frontier Models Score Below 50% on the First Benchmark for Agentic Enterprise IT Tasks — by Artificial Analysis and IBM

wpnews.pro

cd /news/ai-agents/itbench-aa-frontier-models-score-bel… · home › topics › ai-agents › article

[ARTICLE · art-15562] src=huggingface.co ↗ pub=2026-05-27T17:20Z topic=ai-agents verified=true sentiment=· neutral

ITBench-AA: Frontier Models Score Below 50% on the First Benchmark for Agentic Enterprise IT Tasks — by Artificial Analysis and IBM

Artificial Analysis and IBM Research launched ITBench-AA, the first benchmark for agentic enterprise IT tasks, revealing that all frontier AI models scored below 50% on Site Reliability Engineering challenges involving Kubernetes incident response. Claude Opus 4.7 led at 47%, followed by GPT-5.5 at 46% and Qwen3.7 Max at 42%, with the benchmark exposing that longer investigation trajectories did not yield higher accuracy. The new evaluation, based on IBM's ITBench dataset and implemented through an open-source harness, marks the least saturated agentic benchmark in Artificial Analysis's suite, with plans to expand into FinOps and CISO tasks.

read4 min views11 publishedMay 27, 2026

Enterprise ArticlePublished May 27, 2026 Artificial Analysis and IBM Research are launching ITBench-AA, the first in a new series of benchmarks evaluating models on agentic enterprise IT tasks, starting with Site Reliability Engineering tasks where frontier models score below 50% ITBench-AA’s SRE tasks benchmark model performance on Kubernetes incident response, where models and agents must diagnose live systems by reading logs, tracing dependencies, and identifying root-cause entities across complex infrastructure. The underlying ITBench dataset has been developed by IBM Research, leveraging IBM’s deep expertise in enterprise IT operations. Artificial Analysis has worked closely with IBM over the last 6 months to develop an implementation of the dataset for frontier AI evaluation, beginning with Site Reliability Engineering (SRE) and expanding to Financial Operations (FinOps) and Chief Information Security Officer (CISO) tasks over time.

Key findings: #

Claude Opus 4.7 (Adaptive Reasoning, Max Effort) leads at 47%, followed by GPT-5.5 (xhigh) at 46% and Qwen3.7 Max at 42%.
All frontier models score below 50%, making ITBench-AA SRE one of the least saturated agentic benchmarks in our suite. For context, frontier models score considerably higher on Terminal-Bench.
Turn counts vary nearly 3x and longer trajectories do not translate to higher accuracy. GPT-5.5 (xhigh) averages 31 turns per task at 46%, while Gemini 3.1 Pro Preview averages 83 turns at 30%. Models that over-investigate tend to surface upstream fault-injection mechanisms or co-occurring symptoms as false positives.
GLM-5.1 (Reasoning) leads open weights models at 40%, effectively tied with Gemini 3.5 Flash (high). DeepSeek V4 Pro (Reasoning, Max Effort) follows at 38%, with Gemma 4 31B (Reasoning) at 37%, ahead of Gemini 3.1 Pro Preview at 30%.

ITBench-AA SRE overview: #

59 SRE tasks in total: 40 public tasks and 19 brand new, held-out tasks
Each task provides a Kubernetes incident snapshot containing alerts, events, traces, metrics, logs, and application topology. The model must identify the minimal set of independent root-cause Kubernetes entities responsible for the incident.
Faults span typical SRE failure modes including infrastructure, service, application, and chaos-injected incidents, such as resource quota exhaustion, rollout failures, connection pool exhaustion, and network partitions. Methodology details:
Agentic harness: each task is solved by the model running in our open-source Stirrup reference harness, with shell access to a sandboxed file system containing the relevant logs and snapshots. 100-turn cap per task, 3 repeats per task.
Models and agents submit a list of root-cause entities (Kubernetes Deployments, Services, Pods, etc.) they believe caused the incident. Each submission is compared against a ground-truth set of root causes provided by IBM Research.
Scoring uses average precision at full recall: if a model misses any of the ground-truth root causes, it scores 0.0 for that repeat. If it identifies all of them, it is awarded a score equal to its precision - the share of its submitted entities that are actual root causes, i.e. true positives / (true positives + false positives). The headline score is the average across 59 tasks × 3 repeats.
The harness (Stirrup) is held constant across all evaluated models, allowing an apples-to-apples comparison between models.

Highlights #

Tasks require agents to investigate Kubernetes incident snapshots through shell commands and submit a structured JSON diagnosis identifying the responsible root-cause entities. In one public SRE task, the agent sees user-facing failures in the frontend path. It uses shell commands to inspect the offline snapshot: reviewing alerts shows the incident window, then traces/logs narrow the failure to frontend traffic. Topology pins down the affected services, and Kubernetes manifests reveal a network policy blocking the frontend. The successful diagnosis identifies the responsible root-cause entity: otel-demo/NetworkPolicy/frontend-block-all-ports.
More turns do not mean better answers. Models that submit additional contributing entities beyond the true root cause get penalized: identifying the correct root cause but adding upstream mechanisms (e.g., a chaos-mesh controller) or co-occurring symptoms counts as a false positive under recall-gated precision. This is why some models with long trajectories underperform terser ones: Gemini 3.1 Pro Preview averages 83 turns and scores 30%, while Gemma 4 31B (Reasoning) averages 58 turns and scores 37%.
Open weights models sit on the cost frontier of ITBench-AA SRE. Gemma 4 31B (Reasoning) scores 37% at $0.14 per task, outperforming Gemini 3.1 Pro Preview ($2.23 per task, 30%) on both score and cost. GLM-5.1 (Reasoning) scores 40% at $1.23 per task, matching Gemini 3.5 Flash (high) ($1.70) on score at lower cost. Claude Opus 4.7 (Adaptive Reasoning, Max Effort) leads the leaderboard at 47% but is the most expensive at $5.38 per task.

ITBench-AA is built in partnership with @IBMResearch based on their ITBench benchmark.

- For more information see: ITBench paper on arXiv:
[https://arxiv.org/abs/2502.05352](https://arxiv.org/abs/2502.05352) - GitHub:
[https://github.com/itbench-hub/ITBench](https://github.com/itbench-hub/ITBench) - ITBench-AA leaderboard:
[https://artificialanalysis.ai/evaluations/itbench-aa](https://artificialanalysis.ai/evaluations/itbench-aa) - ITBench-AA HuggingFace repo:
[https://huggingface.co/datasets/ArtificialAnalysis/ITBench-AA/tree/main/sre](https://huggingface.co/datasets/ArtificialAnalysis/ITBench-AA/tree/main/sre)

source & further reading

huggingface.co — original article Add thamilvendhan/signalbench to the Benchmark allow-list Need feedback on my research Attachements not available on https://agents-course-unit4-scoring.hf.space/docs#

~/api · this article 200

$curl api.wpnews.pro/v1/news/itbench-aa-frontier-mode…

Read original on huggingface.co → huggingface.co/blog/ibm-research/itbench-aa

mentioned entities

Artificial Analysis

IBM Research

ITBench-AA

Claude Opus 4.7

GPT-5.5

Qwen3.7 Max

IBM

Kubernetes

metadata

slugitbench-aa-frontier-models-score-below-50-on-the-first-benchmark-for-agentic-it

topic#ai-agents

secondary4 topics

sentimentneutral

canonicalhuggingface.co

navigation

← prevBosses blinded by confidence abo…

next →AlpineGate Moves AgentFactory to…

── more in #ai-agents 4 stories · sorted by recency

pub.towardsai.net · 11 Jul · #ai-agents

Grok 4.5 Is xAI's Coding Comeback. The Price Is the Shock.

dev.to · 11 Jul · #ai-agents

$60 Billion for a Dataset: Why Grok 4.5 Just Killed the "Clever Architecture" Myth

evolvinglab.ai · 12 Jul · #ai-agents

AI Should Build Its Own Research World Model

dev.to · 12 Jul · #ai-agents

How to Stop AI Agent Cost Blowups Before They Happen

── more on @artificial analysis 3 stories trending now

wpnews · 30 May · #ai-safety

Nightcord Security Analysis Report - Threat Investigation

wpnews · 8 Jul · #artificial-intelligence

SpaceXAI unveils Grok 4.5 AI model ahead of July 2026 public release

wpnews · 8 Jul · #artificial-intelligence

xAI Launches Grok 4.5 With Pricing Built to Undercut Anthropic's Opus 4.8

sponsored brought to you by zahid.host 4,200+ EU-deployed projects

reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main

→ Live at https://your-agent.zahid.host ✓

Get free account → Pricing

from €0/mo · no card required