{"slug": "evaluating-whether-llms-can-reliably-connect-the-dots", "title": "Evaluating Whether LLMs Can Reliably Connect the DOTs?", "summary": "A new arXiv paper (2609.38406v1) introduces a ~9.2K-instance multi-domain benchmark for narrative text infilling and reports that model scale does not reliably predict infilling quality, with Gemma-2-2B scoring highest qualitatively at 4.02/5 and outperforming DeepSeek-Qwen-32B (3.77/5, 6.6%) and LLaMA-3.3-70B (3.71/5, 8.3%). The study evaluated 20 instruction-tuned open-source LLMs from 1.5B to 70B parameters across four narrative types — encyclopedic text, commonsense stories, news articles, and visual narratives — and found chain-of-thought reasoning yielded only a marginal +0.6% improvement. The authors conclude that short narratives and domain characteristics are stronger predictors of task difficulty than infill position for current LLMs.", "body_md": "arXiv:2609.38406v1 Announce Type: new \nAbstract: Access to real-world information is often noisy and fragmented. Constructing a coherent narrative from such fragments requires models to reconstruct missing spans within a broader storyline, commonly referred to as text infilling, while preserving consistency with both the local context and the global storyline. Despite using text infilling as a pre-training objective in many Large Language Models (LLMs), their actual performance on real-world narrative infilling remains underexplored. In this paper, we address this gap by introducing a multi-domain benchmark of ~9.2K instances for narrative infilling, constructed by masking one to three sentences across four narrative types: encyclopedic text, commonsense stories, news articles, and visual narratives. Using this benchmark, we evaluate 20 instruction-tuned open-source LLMs ranging from 1.5B to 70B parameters across varying levels of instruction specificity and reasoning guidance. Outputs are assessed using standard automatic metrics and a qualitative framework covering five narrative dimensions. Results show that model scale does not reliably predict infilling quality: Gemma-2-2B achieves the highest qualitative score (4.02/5), outperforming models over ten times larger, including DeepSeek-Qwen-32B (3.77/5, 6.6%) and LLaMA-3.3-70B (3.71/5, 8.3%). We further find that explicit reasoning offers limited benefits as chain-of-thought reasoning yields only a marginal improvement (+0.6%). Additionally, short narratives and domain characteristics emerge as stronger predictors of task difficulty than infill position alone for narrative infilling in current LLMs.", "url": "https://wpnews.pro/news/evaluating-whether-llms-can-reliably-connect-the-dots", "canonical_source": "https://arxiv.org/abs/2609.38406", "published_at": "2026-10-01 04:00:00+00:00", "updated_at": "2026-10-01 04:19:25.445946+00:00", "lang": "en", "topics": ["large-language-models", "natural-language-processing", "ai-research", "machine-learning"], "entities": ["Gemma-2-2B", "DeepSeek-Qwen-32B", "LLaMA-3.3-70B", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/evaluating-whether-llms-can-reliably-connect-the-dots", "markdown": "https://wpnews.pro/news/evaluating-whether-llms-can-reliably-connect-the-dots.md", "text": "https://wpnews.pro/news/evaluating-whether-llms-can-reliably-connect-the-dots.txt", "jsonld": "https://wpnews.pro/news/evaluating-whether-llms-can-reliably-connect-the-dots.jsonld"}}