WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents Researchers introduced WhatWorkedBench, a benchmark that measures "experimental understanding" in AI agents by testing the accuracy of their predictions about component changes after budgeted experimentation. Agents inspect code, select measurements, and submit a response as part of the evaluation. AI research agents need reliable knowledge of how their experiments change outcomes. We introduce WhatWorkedBench to measure experimental understanding, the accuracy of predictions about component changes after budgeted experimentation. Agents inspect code, select measurements, and submit a response