{"slug": "litreview-arena-evaluating-literature-review-agents-with-battle-style-peer", "title": "LitReview Arena: Evaluating Literature Review Agents with Battle-Style Peer Review Platform", "summary": "Researchers introduced LitReview Arena, a battle-style evaluation platform for literature review agents, and found that even the strongest current systems win only 23.0% of decisive matches against human drafts on overall utility, while agentic LLMs like Sonar Deep Research outperform base language models by over 60%. The platform collected approximately 3,000 expert judgments across five literature-review-specific criteria, and the authors released an expert-calibrated evaluator, LitJudge, which improves alignment with human experts to Spearman's rho=0.78, comparable to inter-expert consistency. Code and data are publicly available at https://github.com/VanellopeAsher/LitReview-Arena.", "body_md": "arXiv:2608.21374v1 Announce Type: new\nAbstract: Literature reviews are essential to scientific progress, but rigorously evaluating automatically generated reviews remains difficult because many aspects of research utility depend on expert judgment rather than reference-overlap metrics. We introduce LitReview Arena, a battle-style evaluation platform with a structured protocol tailored to literature review quality: domain experts with AI paper-writing experience compare anonymized drafts, are matched to topics within their expertise, and provide dimension-wise outcomes over five literature-review-specific criteria. From this protocol, we collect approximately 3k expert judgments, each containing five dimension-wise outcomes, and show that even the strongest current systems win only 23.0% of decisive matches against human drafts on overall utility, while agentic LLMs such as Sonar Deep Research substantially outperform base language models by over 60%. We further find that existing LLM-as-a-judge methods are substantially misaligned with human experts (Spearman's rho=0.467), especially on synthesis-heavy criteria such as paper structure and research suggestions. Using the collected preference data, we provide an expert-calibrated evaluator, LitJudge, which improves alignment to Spearman's rho=0.78, comparable to inter-expert consistency; code and data are publicly available at https://github.com/VanellopeAsher/LitReview-Arena.", "url": "https://wpnews.pro/news/litreview-arena-evaluating-literature-review-agents-with-battle-style-peer", "canonical_source": "https://arxiv.org/abs/2608.21374", "published_at": "2026-08-25 04:00:00+00:00", "updated_at": "2026-08-25 04:13:57.532261+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-tools"], "entities": ["LitReview Arena", "LitJudge", "Sonar Deep Research", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/litreview-arena-evaluating-literature-review-agents-with-battle-style-peer", "markdown": "https://wpnews.pro/news/litreview-arena-evaluating-literature-review-agents-with-battle-style-peer.md", "text": "https://wpnews.pro/news/litreview-arena-evaluating-literature-review-agents-with-battle-style-peer.txt", "jsonld": "https://wpnews.pro/news/litreview-arena-evaluating-literature-review-agents-with-battle-style-peer.jsonld"}}