Evaluating the Impact of Reviewer Guideline Design on LLM-Based Automated Peer Review A new study on arXiv (2607.22553v1) finds that official conference guidelines produce LLM-based automated peer review results most consistent with human judgments, while reviewer-imitating guidelines generated from high-quality human reviews are less effective. The research also shows that enforcing strict rubric-style scoring consistently degrades performance, emphasizing the need for subjective and holistic scoring in automated review. arXiv:2607.22553v1 Announce Type: new Abstract: Peer review is an essential process in scientific research, yet the growing workload has made its automation increasingly necessary. In this study, we analyze how different types of reviewer guidelines, such as official conference guidelines and reviewer-imitating ones generated from high-quality human reviews using LLMs, affect automated peer review. Our experiments show that official conference guidelines produce review results most consistent with human judgments, suggesting that evaluation criteria refined through conference practice serve as effective guidance for automated reviewing as well. In contrast, reviewer-imitating guidelines were generally less effective than official conference guidelines. Furthermore, enforcing strict rubric-style scoring consistently degraded performance, highlighting the importance of allowing subjective and holistic scoring.