cd /news/large-language-models/more-than-mimicking-reviewers-evalua… · home topics large-language-models article
[ARTICLE · art-125510] src=machinebrief.com ↗ pub= topic=large-language-models verified=true sentiment=· neutral

More Than Mimicking Reviewers: Evaluating LLMs for Pre-Submission Peer Review

A study of an author-facing LLM system that generates pre-submission peer review feedback found that independent sampling covered 44.9% of historical issues on a ten-paper diagnostic, while deduplication and refill reached 78.7% strict and 84.9% seriousness-weighted coverage at 3.6x more requests and 5.2x more tokens. The research, posted as arXiv:2609.05788v1, drew on 3,398 manuscripts with accessible pre-review versions from 10,000 ICLR 2026 submissions. A hidden Top-32 Oracle preserved the full 79.3% weighted coverage of a 256-candidate pool, but paper-only selectors retained only 40-44%, with ablations identifying representative selection and matcher sensitivity as the main sources of the gap.

by read1 min views1 publishedSep 10, 2026

arXiv:2609.05788v1 Announce Type: new Abstract: Peer-review feedback often arrives too late for authors to make meaningful revisions. We study an author-facing LLM system that moves part of this stress test before submission: it generates a broad pool of atomic concerns and compresses them into a short report. We evaluate agreement with historical reviews and, separately, the possible validity of concerns they omit. From 10,000 ICLR 2026 submissions, we use 3,398 manuscripts with accessible versions that predate review. On a ten-paper diagnostic, independent sampling covers 44.9% of historical issues; deduplication and refill reaches 78.7% strict and 84.9% seriousness-weighted coverage, at 3.6$\times$ more requests and 5.2$\times$ more tokens. A hidden Top-32 Oracle preserves the full 79.3% weighted coverage of a 256-candidate pool, but paper-only selectors retain only 40--44%. LLM review therefore provides broad coverage with a large candidate pool but compresses poorly; ablations identify representative selection and matcher sensitivity as the main sources of this gap.

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/more-than-mimicking-…] indexed:0 read:1min 2026-09-10 ·