cd /news/large-language-models/scilitbench-benchmark-and-design-pri… · home topics large-language-models article
[ARTICLE · art-125400] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

SciLitBench: Benchmark and Design Principles for LLM-Powered Systematic Literature Reviews

A new multi-stage benchmark called SciLitBench, spanning 42,981 retrieved records, 1,012 full texts, and annotations for 888 included papers, finds that explicit inclusion and exclusion criteria improve title and abstract screening F2 by 28.8% across 22 open-weight LLMs from six model families, while researcher-authored rationales improve full-text screening by 15%. Data extraction proves less reliable: accuracy falls from 0.97 for publication year to 0.37 Jaccard overlap for computational approach, and the strongest models recover only 30% of annotated evaluation evidence and 25% of limitations. The benchmark's authors identify a practical boundary between high-recall screening and evidence-complete extraction in LLM-assisted systematic literature reviews.

by read1 min views2 publishedSep 10, 2026

arXiv:2609.05505v1 Announce Type: new Abstract: Systematic reviews require sustained human judgment across thousands of records, yet existing evaluations of large language models (LLMs) typically examine review stages in isolation. We introduce SciLitBench, a multi-stage benchmark spanning title and abstract screening, full-text screening, and schema-guided data extraction, with 42,981 retrieved records, 1,012 full texts, and annotations for 888 included papers. Across 22 open-weight LLMs from six model families, explicit inclusion and exclusion criteria improve title and abstract screening $F_2$ by 28.8%, while researcher-authored rationales improve full-text screening by 15%. Data extraction reveals a different reliability regime: performance declines from 0.97 accuracy for publication year to 0.37 Jaccard overlap for computational approach, while the strongest models recover only 30% of annotated evaluation evidence and 25% of limitations. SciLitBench identifies a practical boundary between high-recall screening and evidence-complete extraction and provides a reproducible resource for evaluating LLM-assisted evidence synthesis.

── more in #large-language-models 4 stories · sorted by recency
── more on @scilitbench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/scilitbench-benchmar…] indexed:0 read:1min 2026-09-10 ·