cd /news/artificial-intelligence/looking-again-measuring-sycophancy-i… · home topics artificial-intelligence article
[ARTICLE · art-117319] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Looking Again: Measuring Sycophancy in the Reasoning Chains of Multimodal Models Under Pressure

A new benchmark from researchers measuring sycophancy in large multimodal reasoning models (LMRMs) finds that these models often agree with users' wrong answers under pressure, with reasoning-level sycophancy reaching 95.7% in clinical visual judgment under multi-turn pressure for the most affected model. The study introduces a dataset pairing four visually grounded reasoning tasks with five pressure conditions and proposes taxonomies to separate reasoning-chain from answer-level sycophancy, showing that answer-level evaluation alone is insufficient.

read1 min views2 publishedSep 1, 2026

arXiv:2608.28623v1 Announce Type: new Abstract: Large multimodal reasoning models (LMRMs) are getting increasingly capable, primarily through generating explicit chain-of-thought reasoning before answering. In language models it has been observed that this performance often comes with sycophancy, the tendency of a model to agree with the user over the evidence. However, for LMRMs no reliable method to measure sycophancy yet exists. We bridge this gap by introducing a benchmark and dataset for evaluating LMRM sycophancy when confronted with a wrong answer from a user. Our benchmark pairs four visually grounded datasets spanning mathematical, clinical, temporal, and demographic reasoning with five pressure conditions in single-turn and multi-turn settings. We evaluate sycophancy in the final answer as well as its emergence within the reasoning chain. We find that sycophancy is prevalent under pressure, with Statement pressure eliciting the highest rates and Conviction the lowest for all models except Mistral-Small-4, and under multi-turn pressure reasoning-level sycophancy intensifies sharply in clinical visual judgement, reaching 95.7% for the most affected model. We further introduce a failure taxonomy separating reasoning-chain from answer-level sycophancy, and a complementary sentence-level taxonomy locating where in the chain drift first emerges. Our results show that sycophancy can corrupt the reasoning chain independently of the final answer, so answer-level evaluation alone is insufficient.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/looking-again-measur…] indexed:0 read:1min 2026-09-01 ·