cd /news/artificial-intelligence/firstpass-a-multi-domain-multi-round… · home topics artificial-intelligence article
[ARTICLE · art-113794] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

FIRSTPASS: A Multi-Domain, Multi-Round Peer Review Dataset Grounded in Real Editorial Outcomes

Researchers released FIRSTPASS, the first large-scale peer review dataset built on complete multi-round editorial dialogues from Nature Communications' mandatory transparent peer review, comprising 3,668 records across biology, chemistry, neuroscience, physics, and earth science. Each record includes referee reports, author responses, and updated assessments, with outcome labels derived from editorial decisions, providing ground truth for training AI systems to critique scientific work beyond computer science. Expert reviews average 2,155 words, and all data and pipelines are released for reproducible benchmarking.

read1 min views1 publishedAug 28, 2026

arXiv:2608.26129v1 Announce Type: new Abstract: Scientific peer review datasets have trained AI systems exclusively on Computer Science and Machine Learning venues, producing models that critique ablation studies yet have never seen a biology reviewer demand contamination controls or a chemist question Nuclear Magnetic Resonance (NMR) spectral assignments. We introduce FIRSTPASS, the first large-scale peer review dataset built on complete multi-round editorial dialogues from a multidisciplinary high-impact journal. Curated from Nature Communications mandatory transparent peer review (instituted November 2022), FIRSTPASS comprises 3,668 records spanning five scientific domains (biology, chemistry, neuroscience, physics, and earth science), capturing the full iterative structure of scientific validation: initial referee reports, author point-by-point responses, and updated reviewer assessments. Each record carries an outcome label derived directly from editorial decisions (STANDARD for two-round review; EXTENDED for three or more rounds), providing ground truth absent in all prior corpora. An automated audit confirms 100% content integrity. Expert reviews average 2,155 words, substantially denser than conference venue reviews. All data, parsing pipelines, and evaluation scripts are released to enable reproducible benchmarking of AI scientific judgment across disciplines.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @firstpass 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/firstpass-a-multi-do…] indexed:0 read:1min 2026-08-28 ·