arXiv:2609.05505v1 Announce Type: new Abstract: Systematic reviews require sustained human judgment across thousands of records, yet existing evaluations of large language models (LLMs) typically examine review stages in isolation. We introduce SciLitBench, a multi-stage benchmark spanning title and abstract screening, full-text screening, and schema-guided data extraction, with 42,981 retrieved records, 1,012 full texts, and annotations for 888 included papers. Across 22 open-weight LLMs from six model families, explicit inclusion and exclusion criteria improve title and abstract screening $F_2$ by 28.8%, while researcher-authored rationales improve full-text screening by 15%. Data extraction reveals a different reliability regime: performance declines from 0.97 accuracy for publication year to 0.37 Jaccard overlap for computational approach, while the strongest models recover only 30% of annotated evaluation evidence and 25% of limitations. SciLitBench identifies a practical boundary between high-recall screening and evidence-complete extraction and provides a reproducible resource for evaluating LLM-assisted evidence synthesis.
SciLitBench: Benchmark and Design Principles for LLM-Powered Systematic Literature Reviews
A new multi-stage benchmark called SciLitBench, spanning 42,981 retrieved records, 1,012 full texts, and annotations for 888 included papers, finds that explicit inclusion and exclusion criteria improve title and abstract screening F2 by 28.8% across 22 open-weight LLMs from six model families, while researcher-authored rationales improve full-text screening by 15%. Data extraction proves less reliable: accuracy falls from 0.97 for publication year to 0.37 Jaccard overlap for computational approach, and the strongest models recover only 30% of annotated evaluation evidence and 25% of limitations. The benchmark's authors identify a practical boundary between high-recall screening and evidence-complete extraction in LLM-assisted systematic literature reviews.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.