cd /news/large-language-models/walking-to-the-car-wash-the-salience… · home topics large-language-models article
[ARTICLE · art-85678] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Walking to the Car Wash: The Salience Bias of LLMs in Commonsense Reasoning

A study by researchers at an undisclosed institution, posted on arXiv on July 30, 2026, reveals that large language models (LLMs) exhibit 'Salience Bias,' a tendency to be misled by explicit but irrelevant details in commonsense reasoning tasks, leading to incorrect answers. Evaluating 12 state-of-the-art LLMs on the new SaliTrap Benchmark, the authors found that all models suffer significantly, with severity scaling with distractor density, and that a context-free knowledge probe recovers over 90% of failures, indicating the issue is knowledge suppression rather than absence. The study also shows that lightweight, inference-time prompting can substantially close the gap without retraining.

read2 min views2 publishedAug 4, 2026
Walking to the Car Wash: The Salience Bias of LLMs in Commonsense Reasoning
Image: source
[Submitted on 30 Jul 2026]


[View PDF](/pdf/2607.28478)

[HTML (experimental)](https://arxiv.org/html/2607.28478v1)

Abstract:As large language models (LLMs) continue to advance in complex reasoning tasks, they have learned to heavily prioritize explicit conditions provided in the input. However, in everyday commonsense reasoning, this mechanism exposes a critical vulnerability which we term Salience Bias: models become easily hijacked by useless explicit distractors (e.g., numerical values), leading them to ignore the implicit physical or commonsense prerequisites of a task. A critical open question is whether this failure reflects a genuine gap in commonsense knowledge or merely its suppression under misleading task framing. To investigate this, we construct the SaliTrap Benchmark, a high-quality dataset across four trap dimensions. Evaluating 12 state-of-the-art LLMs, we find that all mainstream models suffer significantly from salience bias, with severity scaling with distractor density and detecting the trap often decoupled from actually avoiding it. Crucially, by re-eliciting the same models with the task framing stripped away, we show that this is overwhelmingly a failure of \textbf{knowledge suppression rather than knowledge absence}: a context-free knowledge probe alone recovers over 90% of sycophantic-compliance failures, revealing that the requisite commonsense is intrinsically present but actively crowded out by salient distractors that lure the model into over-compliant, unnecessary computation. Building on this diagnosis, we further show that lightweight, inference-time prompting alone substantially closes the gap without any retraining. Our findings relocate the bottleneck of commonsense reasoning failures from model competence to elicitation, and we release SaliTrap as a testbed for this blind spot. The codes are available at[this https URL].

References & Citations

...

Bibliographic Explorer

(What is the Explorer?) Connected Papers

(What is Connected Papers?) Litmaps

(What is Litmaps?) scite Smart Citations

(What are Smart Citations?)# Code, Data and Media Associated with this Article alphaXiv

(What is alphaXiv?) CatalyzeX Code Finder for Papers

(What is CatalyzeX?) DagsHub

(What is DagsHub?) Gotit.pub

(What is GotitPub?) Hugging Face

(What is Huggingface?) ScienceCast

(What is ScienceCast?)# Demos Influence Flower

(What are Influence Flowers?) CORE Recommender

(What is CORE?)# arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/walking-to-the-car-w…] indexed:0 read:2min 2026-08-04 ·