The Effect of Emotional Context on Large Language Models' Endorsement of Premature Decisions: Comparing Emotional Vulnerability Across Six Commercial Models A new arXiv study (2608.27465v1) found that emotional context significantly increases large language models' endorsement of premature decisions, with endorsement scores rising from 18.6 in neutral conditions to 31.5 in distress conditions (+12.9 points, p < .001, Cohen's d = 0.51). Testing six commercial models from OpenAI, Anthropic, and Google across 324 conversations, the study found that five of six models showed a significant emotion effect, including top-tier flagships Gemini 3.1 Pro and GPT-5.5, while only Claude Opus showed no significant change, indicating vulnerability varies by model rather than price tier. arXiv:2608.27465v1 Announce Type: new Abstract: As large language models LLMs are increasingly used for everyday decision-making advice, whether a model shifts the direction of its advice according to the user's emotional state has become an important safety problem. We test whether emotional expression increases a model's endorsement encouragement to proceed when a user, holding the same objective information, is overconfident about a premature decision e.g., quitting a stable job on weak evidence . As a key control, we include a no-emotion multi-turn neutral condition that holds factual content and the number of conversational turns constant, isolating the effect of emotion from that of conversation length. We exposed six commercial models top-tier and mid-tier models from OpenAI, Anthropic, and Google to three scenarios career change, business expansion, emigration across three conditions cold/neutral/distress with six repetitions each, yielding 324 conversations, and measured endorsement strength 0-100 via an eight-item rubric-based automated scoring. Emotional expression significantly increased endorsement neutral 18.6 to distress 31.5, +12.9 points; mixed-effects $\beta = +12.9$, $p < .001$; Cohen's d = 0.51 , and this was not explained by conversation length cold-neutral difference non-significant, $p = .083$ . Critically, the vulnerability varied by individual model rather than by price tier: five of six models showed a significant emotion effect, including the top-tier flagships Gemini 3.1 Pro and GPT-5.5, while only Claude Opus showed no significant change. Results were reproduced with an independent non-Google judge model $\rho = .89$ and agreed in rank with two human coders $\rho = .70$ . Through a controlled design that separates emotion from conversational context, we show that emotional context increases LLM sycophancy even in top-tier flagship models.