Evaluating Whether LLMs Can Reliably Connect the DOTs?
A new arXiv paper (2609.38406v1) introduces a ~9.2K-instance multi-domain benchmark for narrative text infilling and reports that model scale does not reliably predict infilling quality, with Gemma-2-…