cd /news/large-language-models/information-discernment-in-large-lan… · home topics large-language-models article
[ARTICLE · art-69587] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=↓ negative

Information Discernment in Large Language Models

A new study from arXiv introduces Learn2Discern (L2D), a benchmark revealing that large language models (LLMs) fail at source and truth discernment when integrating external knowledge, performing near chance across 13 models and nearly 670K trials. A pre-registered user study (n=299) confirms that users endorse three normative axioms and report that violations reduce trust and usage intent. The authors identify simple inference-time interventions that improve both forms of discernment.

read1 min views1 publishedJul 23, 2026

arXiv:2607.19355v1 Announce Type: new Abstract: LLMs are increasingly used with external knowledge sources like the internet. Do they weigh information appropriately -- updating more for reliable sources (source discernment) and more when claims bring priors closer to the truth (truth discernment)? We formalize this as information discernment and introduce Learn2Discern (L2D), an experimental framework and benchmark grounded in three normative axioms with interpretable metrics. To establish external validity, a pre-registered, quota-matched user study (n=299) confirms that real LLM users endorse all three axioms and report that violations reduce their trust and usage intent. Across 13 models and nearly 670K trials, we find consistent failures across both dimensions: models perform near chance on source and truth discernment, rely on source popularity twice as much as source reliability, and update roughly equally whether a claim improves or worsens their position relative to the ground truth. Models integrate external knowledge most effectively on datasets where their priors are already the most accurate. Newer and larger models improve truth discernment but not source discernment, a blind spot that model complexity does not address. We identify simple inference-time interventions that improve both forms of discernment. We release our dataset and survey as a testbed for a core alignment property that scales in importance as LLMs replace traditional search.

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/information-discernm…] indexed:0 read:1min 2026-07-23 ·