When is a benchmark conclusion identified across evaluator meanings? A researcher revising a paper on benchmark evaluation frameworks acknowledged that the current arXiv version lags behind the working manuscript, which now introduces a claim-relative inference layer separating evidence stopping, semantic multiplicity, and claim-resolution separation. The updated framework, tested on a pinned Inspect Evals commit with a 124-unit finite frame, treats missing historical data as typed evidence stops rather than benchmark defects, and the researcher confirmed that reported quantities can shift while winners remain identified, as illustrated by an AgentDojo support example. Thank you — this is an unusually helpful response, especially the independent AgentDojo/AutoML reconstructions and the way you separated provenance, defensibility, equivalence, and admission. We are iterating the paper and executable audit quite rapidly, and the current arXiv version is already substantially behind the working manuscript. A few things in your comment are strikingly aligned with changes that are already in the next revision. One small preview: the new version is organized around a claim-relative inference layer rather than a generic robustness label. It now separates evidence stopping — when the historical substrate needed to replay a claim is unavailable, semantic multiplicity — when more than one grounded evaluator meaning remains live, and claim-resolution separation — when exact values, winners, complete orders, and pairwise relations have different identified sets. We also now explicitly separate origin, admission, and execution : source exposure establishes provenance, but not automatically same-target admissibility. The executable record carries support/missingness decisions, endpoint type, origin/admission, exposure state, identified sets, minimum witnesses, and the stable pairwise backbone. The empirical side has expanded quite a bit too. At a pinned Inspect Evals commit we now have a mechanically defined 124-unit finite frame with a terminal disposition for every unit; most stop before claim-level replay because some historical observation, judge/state, support, or semantic-registry object is missing. We treat those as typed evidence stops rather than zero effects or benchmark defects. Your AgentDojo support example is particularly useful to us because it illustrates exactly the kind of result we want the framework to preserve: a reported quantity can move substantially while the winner remains identified. Two suggestions in your post that are not yet as explicit as they should be in our current working version are the evidence horizon for provenance judgments and a more framework-native family-freeze receipt. We are taking both seriously. So thank you — this is very close to the kind of external pressure test we hoped the public thread would produce. The next preprint revision should make the current direction much clearer.