cd /news/large-language-models/a-benchmark-framework-for-screening-… · home › topics › large-language-models › article
[ARTICLE · art-140739] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

A Benchmark Framework for Screening Automation in Systematic Reviews

A new arXiv paper (2609.30298v1) introduces SRBench, a benchmark dataset of 45,064 labeled entries drawn from 32 curated secondary studies, for evaluating large language model performance in systematic review screening. The paper also proposes an evaluation framework that accounts for class imbalance — the natural prevalence of excluded articles relative to included articles — and releases PromptSR, a tool for prompt experimentation, experiment management, and result analysis in LLM-based screening. The authors argue existing evaluation approaches rely on traditional metrics that can be misleading on highly imbalanced systematic review screening datasets.

by read1 min views1 publishedSep 28, 2026

arXiv:2609.30298v1 Announce Type: new Abstract: Systematic reviews (SR) are essential for evidence-based research, but their screening phase is highly time-consuming and labor-intensive. Large language models (LLMs) offer a promising opportunity to reduce this workload by assisting with article relevance classification. However, existing evaluation approaches often rely on traditional metrics that may be misleading for highly imbalanced SR screening datasets.This paper presents a benchmark dataset of $45,064$ labeled entries for evaluating LLM performance in SR screening across 32 curated secondary studies. It proposes an evaluation framework that accounts for class imbalance, i.e., the natural prevalence of excluded articles relative to included articles in SRs. It also introduces PromptSR, a tool designed to support prompt experimentation, experiment management, and result analysis for LLM-based screening. We also present a use case demonstrating the application of SRBench and PromptSR.

── more in #large-language-models 4 stories · sorted by recency
── more on @srbench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/a-benchmark-framewor…] indexed:0 read:1min 2026-09-28 · —