{"slug": "how-closely-do-llm-reviews-align-with-human-peer-review", "title": "How Closely Do LLM Reviews Align with Human Peer Review?", "summary": "A study comparing reviews from OpenAI GPT-5.4, Google Gemini 3.1 Pro Preview, and Anthropic Claude Opus 4.6 with human reviews for 300 ICLR 2026 submissions found that all three LLMs distinguished accepted from rejected papers but none reproduced the oral versus poster distinction. The models showed provider-specific scoring patterns, with Gemini assigning higher ratings and LLMs more frequently identifying missing baseline comparisons while humans raised computational-efficiency concerns.", "body_md": "arXiv:2608.03659v1 Announce Type: new\nAbstract: Large language models (LLMs) are increasingly used to generate scientific reviews, yet existing evaluations rarely examine whether different providers align with both conference decisions and human reviewing priorities within the same controlled setting. We compare reviews from OpenAI GPT-5.4, Google Gemini 3.1 Pro Preview, and Anthropic Claude Opus 4.6 with human reviews and final decisions for 300 topic-matched ICLR 2026 submissions, equally divided among oral, poster, and rejected papers. Each model reviewed every paper using identical instructions and rating scales after decision information was removed. Our study contributes a cross-provider analysis of three complementary dimensions: alignment with broad and fine-grained decision categories, differences in recommendation-scale usage, and thematic agreement in identified weaknesses. All three LLMs distinguished accepted from rejected papers, but none reproduced the oral versus poster distinction present in human ratings. Scoring patterns were provider-specific: Gemini assigned systematically higher ratings, while OpenAI and Claude were closer to humans for rejected and poster papers but more critical of oral papers. Human and LLM reviews also differed in emphasis, with LLMs more frequently identifying missing baseline comparisons and humans more often raising computational-efficiency concerns. These results show that broad decision alignment does not imply agreement with finer human judgments or reviewing priorities.", "url": "https://wpnews.pro/news/how-closely-do-llm-reviews-align-with-human-peer-review", "canonical_source": "https://www.machinebrief.com/news/how-closely-do-llm-reviews-align-with-human-peer-review-7hre", "published_at": "2026-08-05 04:00:00+00:00", "updated_at": "2026-08-05 06:36:11.357870+00:00", "lang": "en", "topics": ["large-language-models", "ai-research"], "entities": ["OpenAI", "Google", "Anthropic", "ICLR"], "alternates": {"html": "https://wpnews.pro/news/how-closely-do-llm-reviews-align-with-human-peer-review", "markdown": "https://wpnews.pro/news/how-closely-do-llm-reviews-align-with-human-peer-review.md", "text": "https://wpnews.pro/news/how-closely-do-llm-reviews-align-with-human-peer-review.txt", "jsonld": "https://wpnews.pro/news/how-closely-do-llm-reviews-align-with-human-peer-review.jsonld"}}