# WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents

> Source: <https://aiflash.com/news/125329/>
> Published: 2026-09-24 02:30:06+00:00

AI research agents need reliable knowledge of how their experiments change outcomes. We introduce WhatWorkedBench to measure experimental understanding, the accuracy of predictions about component changes after budgeted experimentation. Agents inspect code, select measurements, and submit a response
