Evaluating Whether LLMs Can Reliably Connect the DOTs? A new arXiv paper (2609.38406v1) introduces a ~9.2K-instance multi-domain benchmark for narrative text infilling and reports that model scale does not reliably predict infilling quality, with Gemma-2-2B scoring highest qualitatively at 4.02/5 and outperforming DeepSeek-Qwen-32B (3.77/5, 6.6%) and LLaMA-3.3-70B (3.71/5, 8.3%). The study evaluated 20 instruction-tuned open-source LLMs from 1.5B to 70B parameters across four narrative types — encyclopedic text, commonsense stories, news articles, and visual narratives — and found chain-of-thought reasoning yielded only a marginal +0.6% improvement. The authors conclude that short narratives and domain characteristics are stronger predictors of task difficulty than infill position for current LLMs. arXiv:2609.38406v1 Announce Type: new Abstract: Access to real-world information is often noisy and fragmented. Constructing a coherent narrative from such fragments requires models to reconstruct missing spans within a broader storyline, commonly referred to as text infilling, while preserving consistency with both the local context and the global storyline. Despite using text infilling as a pre-training objective in many Large Language Models LLMs , their actual performance on real-world narrative infilling remains underexplored. In this paper, we address this gap by introducing a multi-domain benchmark of ~9.2K instances for narrative infilling, constructed by masking one to three sentences across four narrative types: encyclopedic text, commonsense stories, news articles, and visual narratives. Using this benchmark, we evaluate 20 instruction-tuned open-source LLMs ranging from 1.5B to 70B parameters across varying levels of instruction specificity and reasoning guidance. Outputs are assessed using standard automatic metrics and a qualitative framework covering five narrative dimensions. Results show that model scale does not reliably predict infilling quality: Gemma-2-2B achieves the highest qualitative score 4.02/5 , outperforming models over ten times larger, including DeepSeek-Qwen-32B 3.77/5, 6.6% and LLaMA-3.3-70B 3.71/5, 8.3% . We further find that explicit reasoning offers limited benefits as chain-of-thought reasoning yields only a marginal improvement +0.6% . Additionally, short narratives and domain characteristics emerge as stronger predictors of task difficulty than infill position alone for narrative infilling in current LLMs.