{"slug": "don-t-let-me-ask-for-it-llms-show-deficiencies-in-active-multi-turn-information", "title": "Don't Let Me Ask for It: LLMs Show Deficiencies in Active Multi-Turn Information Acquisition for Abductive Inference", "summary": "A new study from arXiv (2608.03388v1) introduces the Alien Abduction game, an interactive probe revealing that large language models (LLMs) struggle with active multi-turn information acquisition for abductive inference. The study found that providing evidence upfront yields higher success rates than distributing it across turns, and models perform better when examples are provided by an oracle than when they select their own queries, though their final hypotheses are more consistent with self-selected evidence. These findings suggest LLMs may form hypotheses that fit self-selected evidence without sufficiently distinguishing alternatives, and may struggle to validate, refine, or know when to stop.", "body_md": "arXiv:2608.03388v1 Announce Type: new\nAbstract: Abductive reasoning requires forming hypotheses that explain observed evidence and revising them as new evidence becomes available. While large language models (LLMs) are often evaluated on whether they solve abductive reasoning tasks correctly, less is known about how they acquire evidence, update their hypotheses, and decide when to stop. We introduce Alien Abduction game, an interactive probe for studying these behaviours under different interaction modes. The modes vary in whether evidence is provided upfront or across turns, and whether queries are selected by the model or examples are provided by the oracle. Across models, providing evidence upfront leads to higher success rates than distributing it across turns. In multi-turn settings, some models commit before using the available evidence, while others exhaust the turn budget without converging. Models also achieve higher success rates when examples are provided by the oracle than when they select their own queries, although their final hypotheses are more consistent with the evidence they selected. These findings suggest that models may form hypotheses that fit self-selected evidence without sufficiently distinguishing them from alternatives, and may struggle to validate and refine their hypotheses or determine when to stop.", "url": "https://wpnews.pro/news/don-t-let-me-ask-for-it-llms-show-deficiencies-in-active-multi-turn-information", "canonical_source": "https://www.machinebrief.com/news/dont-let-me-ask-for-it-llms-show-deficiencies-in-active-mult-7ema", "published_at": "2026-08-05 04:00:00+00:00", "updated_at": "2026-08-05 05:03:58.326892+00:00", "lang": "en", "topics": ["large-language-models", "artificial-intelligence", "ai-research"], "entities": ["arXiv"], "alternates": {"html": "https://wpnews.pro/news/don-t-let-me-ask-for-it-llms-show-deficiencies-in-active-multi-turn-information", "markdown": "https://wpnews.pro/news/don-t-let-me-ask-for-it-llms-show-deficiencies-in-active-multi-turn-information.md", "text": "https://wpnews.pro/news/don-t-let-me-ask-for-it-llms-show-deficiencies-in-active-multi-turn-information.txt", "jsonld": "https://wpnews.pro/news/don-t-let-me-ask-for-it-llms-show-deficiencies-in-active-multi-turn-information.jsonld"}}