cd /news/large-language-models/evaluating-whether-llms-can-reliably… · home › topics › large-language-models › article
[ARTICLE · art-142991] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Evaluating Whether LLMs Can Reliably Connect the DOTs?

A new arXiv paper (2609.38406v1) introduces a ~9.2K-instance multi-domain benchmark for narrative text infilling and reports that model scale does not reliably predict infilling quality, with Gemma-2-2B scoring highest qualitatively at 4.02/5 and outperforming DeepSeek-Qwen-32B (3.77/5, 6.6%) and LLaMA-3.3-70B (3.71/5, 8.3%). The study evaluated 20 instruction-tuned open-source LLMs from 1.5B to 70B parameters across four narrative types — encyclopedic text, commonsense stories, news articles, and visual narratives — and found chain-of-thought reasoning yielded only a marginal +0.6% improvement. The authors conclude that short narratives and domain characteristics are stronger predictors of task difficulty than infill position for current LLMs.

by read1 min views1 publishedOct 1, 2026

arXiv:2609.38406v1 Announce Type: new Abstract: Access to real-world information is often noisy and fragmented. Constructing a coherent narrative from such fragments requires models to reconstruct missing spans within a broader storyline, commonly referred to as text infilling, while preserving consistency with both the local context and the global storyline. Despite using text infilling as a pre-training objective in many Large Language Models (LLMs), their actual performance on real-world narrative infilling remains underexplored. In this paper, we address this gap by introducing a multi-domain benchmark of ~9.2K instances for narrative infilling, constructed by masking one to three sentences across four narrative types: encyclopedic text, commonsense stories, news articles, and visual narratives. Using this benchmark, we evaluate 20 instruction-tuned open-source LLMs ranging from 1.5B to 70B parameters across varying levels of instruction specificity and reasoning guidance. Outputs are assessed using standard automatic metrics and a qualitative framework covering five narrative dimensions. Results show that model scale does not reliably predict infilling quality: Gemma-2-2B achieves the highest qualitative score (4.02/5), outperforming models over ten times larger, including DeepSeek-Qwen-32B (3.77/5, 6.6%) and LLaMA-3.3-70B (3.71/5, 8.3%). We further find that explicit reasoning offers limited benefits as chain-of-thought reasoning yields only a marginal improvement (+0.6%). Additionally, short narratives and domain characteristics emerge as stronger predictors of task difficulty than infill position alone for narrative infilling in current LLMs.

── more in #large-language-models 4 stories · sorted by recency
── more on @gemma-2-2b 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/evaluating-whether-l…] indexed:0 read:1min 2026-10-01 · —