cd /news/artificial-intelligence/do-multimodal-llms-see-before-they-r… · home topics artificial-intelligence article
[ARTICLE · art-118567] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Do Multimodal LLMs See Before They Read? Diagnosing Contextual Sycophancy

A new arXiv study (2609.00067v1) introduces a 998-case diagnostic showing that external text can override conflicting image evidence in multimodal large language models, a failure termed multimodal contextual sycophancy. On abnormal images paired with Gemini-generated false text, GPT-5.1 scores 7.9% under joint conditioning, 49.7% when a context-blind witness report is scored directly, 63.7% under a two-call witness-arbiter pipeline, and 84.2% under System-2 Visual Arbitration (S2VA), which withholds text from the witness. Across six models, S2VA improves over the direct witness report by 19.7 to 44.1 points, with all paired 95% confidence intervals excluding zero.

read1 min views1 publishedSep 2, 2026

arXiv:2609.00067v1 Announce Type: new Abstract: External text can override conflicting image evidence in multimodal large language models, a failure we call multimodal contextual sycophancy. We introduce a 998-case diagnostic that independently varies visual evidence, commonsense priors, and external text, and probe when this failure arises by moving the information boundary around a context-blind visual witness. On abnormal images paired with Gemini-generated false text, GPT-5.1 scores 7.9% under joint conditioning, 49.7% when the context-blind witness report is scored directly, 63.7% under a matched two-call witness-arbiter pipeline that exposes the witness to the text, and 84.2% under System-2 Visual Arbitration (S2VA), which withholds the text from the witness. Across six models, S2VA improves over the direct witness report by 19.7 to 44.1 points, with all paired 95% confidence intervals excluding zero. The best information boundary is not uniform: textual context scaffolds some models, and a GPT-4o-regenerated subset changes the relative ordering of joint conditioning, Witness-Only, and S2VA. Contextual sycophancy is therefore sensitive to when text is introduced, as well as to the model and context source.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/do-multimodal-llms-s…] indexed:0 read:1min 2026-09-02 ·