Beyond Final Decisions: A Process-Centric Benchmark for Transparent AI-Assisted Peer Review A new arXiv paper (2609.05947v1) introduces a process-centric diagnostic benchmark for AI-assisted peer review that evaluates whether model decisions are supported by sufficient and reliable review evidence, rather than only judging final decision accuracy. Using process-aligned data converted from PeerRead, NLPeer ARR-22, and OpenReview-ICLR, experiments across three datasets and six models found that Gold-process variables generally carry higher decision value and that the Gold–Predicted gap remains stable across datasets and random seeds for the main analysis model. The authors report that although model-generated intermediate review texts show relatively high local consistency across adjacent stages, final decisions are not consistently supported by the preceding review evidence, positioning the benchmark as a transparent, auditable diagnostic tool for systems designed to assist rather than replace human reviewers. arXiv:2609.05947v1 Announce Type: new Abstract: Peer review is central to quality control in science. However, existing evaluations of AI-assisted peer review mainly focus on the overall quality of generated reviews or the accuracy of final decisions. They therefore provide limited evidence about whether model decisions are supported by sufficient and reliable review evidence. We introduce a process-centric diagnostic benchmark for AI-assisted peer review. It uses x,$z s$,$z c$,$z r$,y to represent the paper content, summary, critique, suggestion, and decision. We convert heterogeneous review records from PeerRead, NLPeer ARR-22, and OpenReview-ICLR into process-aligned data. Our benchmark uses direct decision prediction from the paper content Direct as its baseline. It compares the decision value of Gold-process variables and Predicted-process variables, and conducts stage-level evaluation, chain-consistency evaluation, and interventional sensitivity analysis. Experiments across three datasets and six models show that Gold-process variables generally have higher decision value. For the main analysis model, the Gold--Predicted gap remains stable across datasets and random seeds. This gap is also reproduced in most model--dataset combinations. Although model-generated intermediate review texts show relatively high local consistency across adjacent stages, the final decisions are not consistently supported by the preceding review evidence. Our benchmark targets AI systems designed to assist rather than replace human reviewers. It provides a transparent and auditable diagnostic tool for evaluating the reliability of their review processes.