04:00
2026-09-10
arxiv.org
large-language-models
SciLitBench: Benchmark and Design Principles for LLM-Powered Systematic Literature Reviews
A new multi-stage benchmark called SciLitBench, spanning 42,981 retrieved records, 1,012 full texts, and annotations for 888 included papers, finds that explicit inclusion and exclusion criteria improβ¦