{"slug": "do-multimodal-llms-see-before-they-read-diagnosing-contextual-sycophancy", "title": "Do Multimodal LLMs See Before They Read? Diagnosing Contextual Sycophancy", "summary": "A new arXiv study (2609.00067v1) introduces a 998-case diagnostic showing that external text can override conflicting image evidence in multimodal large language models, a failure termed multimodal contextual sycophancy. On abnormal images paired with Gemini-generated false text, GPT-5.1 scores 7.9% under joint conditioning, 49.7% when a context-blind witness report is scored directly, 63.7% under a two-call witness-arbiter pipeline, and 84.2% under System-2 Visual Arbitration (S2VA), which withholds text from the witness. Across six models, S2VA improves over the direct witness report by 19.7 to 44.1 points, with all paired 95% confidence intervals excluding zero.", "body_md": "arXiv:2609.00067v1 Announce Type: new\nAbstract: External text can override conflicting image evidence in multimodal large language models, a failure we call multimodal contextual sycophancy. We introduce a 998-case diagnostic that independently varies visual evidence, commonsense priors, and external text, and probe when this failure arises by moving the information boundary around a context-blind visual witness. On abnormal images paired with Gemini-generated false text, GPT-5.1 scores 7.9% under joint conditioning, 49.7% when the context-blind witness report is scored directly, 63.7% under a matched two-call witness-arbiter pipeline that exposes the witness to the text, and 84.2% under System-2 Visual Arbitration (S2VA), which withholds the text from the witness. Across six models, S2VA improves over the direct witness report by 19.7 to 44.1 points, with all paired 95% confidence intervals excluding zero. The best information boundary is not uniform: textual context scaffolds some models, and a GPT-4o-regenerated subset changes the relative ordering of joint conditioning, Witness-Only, and S2VA. Contextual sycophancy is therefore sensitive to when text is introduced, as well as to the model and context source.", "url": "https://wpnews.pro/news/do-multimodal-llms-see-before-they-read-diagnosing-contextual-sycophancy", "canonical_source": "https://arxiv.org/abs/2609.00067", "published_at": "2026-09-02 04:00:00+00:00", "updated_at": "2026-09-02 04:26:01.616159+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-safety"], "entities": ["arXiv", "GPT-5.1", "Gemini", "System-2 Visual Arbitration"], "alternates": {"html": "https://wpnews.pro/news/do-multimodal-llms-see-before-they-read-diagnosing-contextual-sycophancy", "markdown": "https://wpnews.pro/news/do-multimodal-llms-see-before-they-read-diagnosing-contextual-sycophancy.md", "text": "https://wpnews.pro/news/do-multimodal-llms-see-before-they-read-diagnosing-contextual-sycophancy.txt", "jsonld": "https://wpnews.pro/news/do-multimodal-llms-see-before-they-read-diagnosing-contextual-sycophancy.jsonld"}}