cd /news/large-language-models/don-t-let-me-ask-for-it-llms-show-de… · home topics large-language-models article
[ARTICLE · art-87207] src=machinebrief.com ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Don't Let Me Ask for It: LLMs Show Deficiencies in Active Multi-Turn Information Acquisition for Abductive Inference

A new study from arXiv (2608.03388v1) introduces the Alien Abduction game, an interactive probe revealing that large language models (LLMs) struggle with active multi-turn information acquisition for abductive inference. The study found that providing evidence upfront yields higher success rates than distributing it across turns, and models perform better when examples are provided by an oracle than when they select their own queries, though their final hypotheses are more consistent with self-selected evidence. These findings suggest LLMs may form hypotheses that fit self-selected evidence without sufficiently distinguishing alternatives, and may struggle to validate, refine, or know when to stop.

read1 min views1 publishedAug 5, 2026

arXiv:2608.03388v1 Announce Type: new Abstract: Abductive reasoning requires forming hypotheses that explain observed evidence and revising them as new evidence becomes available. While large language models (LLMs) are often evaluated on whether they solve abductive reasoning tasks correctly, less is known about how they acquire evidence, update their hypotheses, and decide when to stop. We introduce Alien Abduction game, an interactive probe for studying these behaviours under different interaction modes. The modes vary in whether evidence is provided upfront or across turns, and whether queries are selected by the model or examples are provided by the oracle. Across models, providing evidence upfront leads to higher success rates than distributing it across turns. In multi-turn settings, some models commit before using the available evidence, while others exhaust the turn budget without converging. Models also achieve higher success rates when examples are provided by the oracle than when they select their own queries, although their final hypotheses are more consistent with the evidence they selected. These findings suggest that models may form hypotheses that fit self-selected evidence without sufficiently distinguishing them from alternatives, and may struggle to validate and refine their hypotheses or determine when to stop.

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/don-t-let-me-ask-for…] indexed:0 read:1min 2026-08-05 ·