Where Scientific Search Agents Fail: Decision-Checkpoint Auditing of Exposure and Inspection Attempts A new arXiv paper (2609.38670v1) introduces decision checkpoints that record observations and tool actions during inference to distinguish whether a scientific-search agent failed to encounter a target paper, attempted to inspect it, or returned an accepted answer after inspection. Across five conditions on 540 answerable AutoResearchBench Deep questions in a fixed, target-enriched environment, keyword search achieved 24.6% accuracy versus 17.8% for raw search, while read-first recorded 27.4% more evidence-search calls than keyword search. Target inspection attempts occurred on 199 questions under read-first and 191 under keyword search, with both conditions reaching 24.6% accuracy. arXiv:2609.38670v1 Announce Type: new Abstract: Final-answer accuracy does not reveal whether a scientific-search agent failed to encounter a target paper, attempt to inspect it, or return an accepted answer after inspection. We introduce decision checkpoints that record observations and tool actions without benchmark labels during inference, then join target identities and evaluator labels to assign outcome categories from recorded events. Across five conditions on 540 answerable AutoResearchBench Deep questions in a fixed, target-enriched environment, keyword search achieves 24.6\% accuracy, compared with 17.8\% for raw search. The keyword condition has fewer incorrect answers with neither target exposure nor inspection, but more incorrect answers after the target is exposed and left uninspected. Compared with keyword search, read-first has 27.4\% more recorded evidence-search calls. Target inspection attempts occur on 199 questions under read-first and 191 under keyword search; both conditions achieve 24.6\% accuracy. The checkpoint protocol makes these question-level differences explicit, distinguishing target exposure and inspection from aggregate accuracy and total tool use.