AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents

arXiv · AI, language, vision and robotics · article · Sep 23, 2026 · UTC

AI research agents need reliable knowledge of how their experiments change outcomes. We introduce WhatWorkedBench to measure experimental understanding, the accuracy of predictions about component changes after budgeted experimentation. Agents inspect code, select measurements, and submit a response surface, a table predicting scores for every configuration of component settings. Exhaustive CPU execution supplies reference effects for changing each component while holding the others fixed. These effects capture combinations of changes across 36 tasks from 30 data sources and 8 workflow types,

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-24T01:22:21.678Z. This is not the publication date.