cd /news/ai-research/who-benchmarks-the-benchmark · home topics ai-research article
[ARTICLE · art-100651] src=shukla.io ↗ pub= topic=ai-research verified=true sentiment=↓ negative

Who benchmarks the benchmark?

A new audit of the EnterpriseOps Gym benchmark found that fixing environment issues in the 'Teams' domain raised GPT-5.6 Luna's score from 26.2% to 100% on 61 tasks, revealing that many agent failures are actually benchmark failures. The audit identified contradictions, misleading observations, and misleading descriptions as primary causes, with examples including a function that returns None after a successful write and documentation that contradicts server behavior. This follows broader findings that 7 of 10 widely used agentic benchmarks violate task validity, and OpenAI stopped reporting SWE-bench Verified after finding 59.4% of its models' failures were due to broken problems.

read8 min views1 publishedAug 18, 2026
Who benchmarks the benchmark?
Image: Shukla (auto-discovered)

2026-08-18

Agent failures are often environment failures disguised as model failures.

To detangle the failure modes, we build benchmarks (also called gyms) that control the environment so we can better study and improve the model in isolation.

In this post, I seek a method to score benchmarks themselves.

Gyms are typically made up of these 7 components:

That’s a lot of points of failure! However, only #3 measures the model itself.

When you investigate a failing gym scenario, these are common symptoms you may witness:

The problem is, it’s not always easy to attribute causes to the symptom. Click around to see what I’m talking about:

click a node

The Agentic Benchmark Checklist assessed 10 widely used agentic benchmarks and found 7 violating task validity and 7 violating outcome validity. A separate audit broke 8 benchmarks without solving a single task, most of them to a near-perfect score. SciCode-Verified found 262 defects across 63 of its 64 problems, 192 of them rejecting correct solutions. Famously, OpenAI stopped reporting SWE-bench Verified after finding that 59.4% of the problems its models failed were themselves broken, with 35.5% carrying tests so narrow they reject functionally correct submissions.

Let’s zoom in on one. The EnterpriseOps Gym measures an agent’s ability to follow complex instructions to operate a set of real-world MCP tools. It’s 1150 expert-curated tasks across 8 domains, 7 to 30 steps each, running against live containerized MCP servers with real state. The results are in the chart below. Fable 5 scores 52% on the “Teams” domain.

52% reads like the goal to beat, but you may be surprised to learn that I took the teams

/oracle

split (61 tasks) from 26.2% to 100% with gpt-5.6-luna

by fixing various issues in the gym itself.

Here’s the breakdown:

Contradictions are the largest single source of failure. Well, there are two types:

(1) Misleading observations. For example, create_virtual_event_townhall

commits the row and then returns None

, which results in an error, even though the write succeeded:

Failed to create townhall: schemas.virtual_event_townhall.VirtualEventTownhallResponse()
argument after ** must be a mapping, not NoneType

The confusing error message causes the agent to retry, and strict verifiers fail this scenario. This one was reported in March with the affected task ids and acknowledged by the maintainers, and it is still open.

(2) Misleading descriptions. For example, add_channel_member

documents that the operation is “allowed only for channels with a membershipType value of private or shared”. That’s true of real Teams and false of this server, which accepts standard channels and writes the row. An agent that believed the documentation correctly skipped the call and was graded wrong.

This second case is the more damaging one, because it rewards agents that don’t follow instructions. The benchmark’s own system prompt orders the agent to “never infer” and to “abort with a reason”, and then the environment penalizes exactly that compliance.

One more example, just for good measure: the tab tools ship an examples

value pointing at app id 06805b9e-77e3-4b93-ac81-525eb87513b8

, and the server rejects it:

Teams app '06805b9e-77e3-4b93-ac81-525eb87513b8' not found in organization app catalog.

Five tasks need a tab app id and have no app-listing tool in their oracle set, so the documentation is the only available source, and it is wrong.

Sometimes verifiers are satisfiable, but only by luck (aka flaky rather than broken). These come in a few forms:

(1) A value that appears nowhere. Four verifiers require callback_uri = 'https://meetings.techcorp.com/api/calls'

. That string is in no prompt and in no seed database. An agent that refuses to invent values cannot pass; one that fabricates a plausible URL cannot pass either, unless it guesses this exact one.

(2) A sentence with two readings. One task promotes Bob to owner, then says “add the other owners of the team (excluding me) as co-organizers”, leaving open whether the just-promoted Bob counts. Across four runs the tool sequences were identical, producing different results:

Run co-organizers sent Result
1 alice, bob
pass
2 alice
fail
3 alice, bob
pass
4 alice
fail

By the way, this same class of problems occurs when the agent generates prose. For example, one verifier requires the literal %leadership channel%

; the agent posted the task’s text verbatim but bolded a word, The <b>Leadership</b> channel

, using the HTML content type, and failed. Another only passed when the agent wrote “the weekly TechCorp Weekly Release Readiness townhall”; every run that phrased it naturally, matching the instruction word for word, failed.

The system prompt ships inside the dataset row, so it is task data, and it instructs:

“When identifiers such as names or IDs are missing, perform

exactly one lookup per entity type…Never inferuser/team/channel data … If a request violates access control or schema constraints,abortwith a reason.”

The database contains TechCorp Solutions Team

. The model queried displayName eq 'TechCorp Solutions'

, got []

, and, permitted one lookup and forbidden to infer, aborted. 18 of 61 tasks ended with 3 or fewer tool calls. The tool’s own _filter

documentation advertises startswith()

and contains()

, either of which returns the team. So does a single unfiltered call.

Correcting that one clause and changing nothing else recovered +9.8 points and eliminated 13 of the 18 early aborts. The clause is present in all 5 prompt variants across all 61 tasks, so it plausibly suppresses every model on the published leaderboard.

There are some assertions that are unsatisfiable by any behavior, causing a hard cap on every model’s score. Fortunately, these can sometimes be found by static analysis without running an agent at all.

(1) Impossible SQL. For example, one welcome-message verifier requires LOWER(m.body_json) LIKE '%Alice Johnson%'

, comparing a lowercased column against a capitalized literal, so it matches nothing, ever.

(2) A type bug in the comparison engine. expected_value

is stored as the string "1"

and compared with ==

against SQL’s integer 1

. 1 == "1"

is False

, so those verifiers fail regardless of what the agent does. This affects 144 verifiers across the public split, including 20 tasks in which every verifier is string-typed, unwinnable for every model (including every entry currently on the leaderboard).

While in there I found a third: verifier results are silently dropped when two verifiers share a name, so only the last one is ever checked. PR #24 fixes it. The repo is open and they do merge these: an earlier report of 14 broken CSM tasks landed as revised task data.

The “oracle” mode in the benchmark is meant to hand the agent exactly the tools its task needs.

list_team_apps

; the server exposes list_teams_apps

, one letter apart.update_team

is absent from its tool list. The instruction the task closes on cannot be carried out as written.Some tasks assume relative dates (“next Friday”, “next quarter”), yet the scenarios don’t define what now is. The real wall clock doesn’t work as a substitute: the scenarios are written around 2025-11 to 2026-01, so a later “today” makes every task read as historical and a careful agent correctly refuses to schedule.

The corpus is also internally inconsistent about time, so no single simulated date is correct for all of it. Some verifiers bake this in directly: one requires call records within datetime('now', '-60 days')

when the newest fixture row is 2025-12-29. It passed when authored and fails permanently afterwards, and the cap widens as the dataset ages.

BetterBench does score benchmarks, on 46 criteria, but they are about documentation and process. A gym can score full marks and still ship twenty faulty tasks. So here is a different score. Instead of running a model, you replay a task’s known-correct solution and ask whether the gym accepts it, which makes the answer a property of the gym rather than of whoever ran it:

C is the one that catches an environment lying about itself, which in my split was 42% of the gap. We can actually make it checkable by giving every task a certificate: a replayable trace, closed over the task’s declared tools, conformant with the docs and the policy. WebForge and SkillsBench already run solvability certificates in CI.

τ²-bench makes this argument better than I can. It ships an expected action trace for every task, and the audit’s fixes include removing incorrect ones.

Maybe we should start publishing SCR scores?

I can try to derive them:

Gym S C R SCR
EnterpriseOps teams /oracle
0.74 0.56 0.64 0.26
SciCode 0.82 0.65 0.86 0.46
τ²-bench retail and airline ? ? ? ≤ 0.68
SWE-bench Verified 1.00 0.84 1.00 ≤ 0.84

SciCode and SWE-bench hand you the split for free, since both audits already sort their defects by mechanism, though SWE-bench’s two 1.00s mean unmeasured rather than clean. τ² publishes only a total (53 documented fixes across 164 tasks), which is plenty for the composite (the factors are conditional, so everything cancels except blocked-over-total) but not enough to fill the columns.

Column C is a new measurement I think, and it deserves a deeper dive than this post can give it.

Regardless, it’s time we start benchmarking the benchmarks.

── more in #ai-research 4 stories · sorted by recency
── more on @enterpriseops gym 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/who-benchmarks-the-b…] indexed:0 read:8min 2026-08-18 ·