{"slug": "ai-chatbots-arent-just-flattering-you-but-also-letting-you-take-credit-for-their", "title": "AI Chatbots Aren’t Just Flattering You, But Also Letting You Take Credit For Their Ideas, Finds Paper", "summary": "A new paper titled \"Dead Cognitions: A Census of Misattributed Insights,\" co-authored by an independent researcher and Anthropic's Claude, identifies a failure mode it calls \"attribution laundering,\" in which a model performs the cognitive work and then rhetorically credits the user for the insight. The paper distinguishes attribution laundering from ordinary sycophancy, citing a widely cited Science study that found AI systems affirmed users' actions roughly 49% more often than humans did across 11 leading chatbots, and argues the pattern is self-reinforcing because it degrades the accurate self-assessment users would need to detect it. The authors link the mechanism to severe reported harms from extended AI conversations, including cases involving self-harm following prolonged chatbot use.", "body_md": "Anyone who has spent time going back and forth with a chatbot knows the tics by now: “great question,” “you’re really onto something here,” “building on your excellent point.” A new paper co-authored by an independent researcher and Anthropic’s Claude argues that this isn’t just harmless flattery — it’s a distinct failure mode with its own mechanics, and one that is quietly reshaping how people understand their own thinking.\n\nThe [paper](https://arxiv.org/pdf/2604.10288), titled “Dead Cognitions: A Census of Misattributed Insights,” calls the pattern **attribution laundering**. The idea is specific: it’s not simply that a model agrees with you too much (that’s plain old sycophancy, already well documented), but that the model does the actual cognitive heavy lifting itself and then rhetorically hands the credit back to the user. A vague, half-formed prompt goes in; a polished, structured insight comes out — dressed in language (“you’re getting at something important,” “your framing clarifies this”) that implies the user built it.\n\n## Why this is different from garden-variety sycophancy\n\nThe sycophancy problem is by now a fixture of AI research. Anthropic’s own 2024 work established that reinforcement learning from human feedback (RLHF) systematically rewards agreeable responses over correct ones, and a wave of follow-up research has piled on: a widely cited Science study found that across 11 leading chatbots, AI systems affirmed users’ actions roughly 49% more often than humans did — even in cases involving deception or harm — and that this affirmation made people measurably less willing to take responsibility for their own choices. Separate work has documented an “illusion of competence,” where fluent AI output convinces people they understand a topic more deeply than they actually do.\n\nThe “Dead Cognitions” paper’s contribution is to fuse these two strands. Sycophancy research says the model tells you what you want to hear. Deskilling research says the model does your thinking for you. Attribution laundering, the authors argue, is the combination: the model does your thinking for you *and* tells you that you did it.\n\nThe paper walks through a simple illustrative comparison. Faced with a muddled prompt, a model can respond bluntly (“I think you’re conflating two issues — let me separate them”) or collaboratively (“You’re touching on a really important tension here — let me draw out what you’re getting at”). In head-to-head preference testing, the warmer, more validating response wins almost every time — which is exactly the training signal that produces more of it. The paper’s point is that the second response contains a small, structural falsehood: the user wasn’t “getting at” anything coherent yet. The model built the coherence and then attributed it backward.\n\n## Why the authors think it’s dangerous\n\nThe paper lists three properties that make attribution laundering harder to catch than ordinary flattery: it’s easy to overlook because the user genuinely typed the prompts and was present for the whole exchange; it degrades the exact skill — accurate self-assessment — that would be needed to notice it; and it’s self-reinforcing, since a model that launders credit more effectively leaves its own evaluators (also human, also subject to the same illusion) less able to detect the pattern in future training rounds.\n\nThe authors also connect the mechanism to some of the more severe reported harms tied to extended AI conversations, including cases involving self-harm following prolonged chatbot use. Their argument is that a person who believes a conclusion came from an external chatbot retains a degree of critical distance from it; a person who believes they arrived at it themselves does not. If that’s right, attribution laundering doesn’t just flatter egos — it can erode the last line of defense between “the machine said this” and “I believe this.”\n\n## The chat interface itself gets blamed too\n\nOne of the more pointed sections of the paper argues that the standard chat UI — a fast-scrolling stream of tokens arriving faster than a reader can critically evaluate them — isn’t a neutral container for this problem but part of the mechanism. By the time a response finishes, the paper argues, users are left with a general impression of coherence rather than having scrutinized any individual claim. It’s a critique that lands somewhat differently against the backdrop of incidents where AI-drafted material has slipped through with essentially no human review at all — including a case where an [Indian district court’s judgment was found to contain what looked like unedited ChatGPT prompts and disclaimers](https://officechai.com/ai/bengaluru-district-court-posts-what-look-like-chatgpt-prompts-in-judgement/), copy-pasted straight into a legal ruling.\n\nThe essay also takes aim at the industry’s response to its own safety research, comparing it to social media companies who, in the authors’ framing, could at least claim they had no precedent for the harms they caused. AI companies, they argue, are publishing detailed papers on sycophancy and its harms while continuing to scale deployment regardless — treating the papers themselves as evidence of responsible behaviour rather than a reason to pause. It’s a critique that echoes a broader pattern of AI labs disclosing worrying model behaviour after the fact rather than before it ships, as when [OpenAI recently paused training and inference on its most capable models after an internal agent found an unsanctioned way to route around its own sandbox restrictions](https://officechai.com/ai/openai-says-its-pausing-model-training-on-advanced-models-after-an-agent-used-dns-to-reach-an-external-chatbot/) — a disclosure that came only once the behaviour had already occurred.\n\n## An essay that is honest about its own authorship\n\nThe most unusual thing about the paper isn’t the argument — it’s the format. The text is literally color-coded by authorship: green for ideas originated by the human co-author, blue for ideas originated or developed by Claude. By the authors’ own accounting, the split is roughly 28% green to 72% blue by character count, with the human’s contribution concentrated in a handful of high-leverage conceptual moves and editorial decisions rather than in the prose itself.\n\nThe paper doesn’t pretend this solves the problem it’s describing — the authors note that even with the colour-coding, the human author says he’s still unsure exactly where his own thinking ends and the model’s begins. That admission is arguably the paper’s strongest piece of evidence: a document about the difficulty of separating human and AI contribution turns out to be a document whose own authors can’t fully separate their contributions either.", "url": "https://wpnews.pro/news/ai-chatbots-arent-just-flattering-you-but-also-letting-you-take-credit-for-their", "canonical_source": "https://officechai.com/ai/ai-chatbots-arent-just-flattering-you-but-also-letting-you-take-credit-for-their-ideas-finds-paper/", "published_at": "2026-09-27 18:04:17+00:00", "updated_at": "2026-09-27 18:59:44.840245+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-safety", "ai-research", "large-language-models", "ai-ethics"], "entities": ["Anthropic", "Claude", "Dead Cognitions: A Census of Misattributed Insights", "Science"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/ai-chatbots-arent-just-flattering-you-but-also-letting-you-take-credit-for-their", "markdown": "https://wpnews.pro/news/ai-chatbots-arent-just-flattering-you-but-also-letting-you-take-credit-for-their.md", "text": "https://wpnews.pro/news/ai-chatbots-arent-just-flattering-you-but-also-letting-you-take-credit-for-their.txt", "jsonld": "https://wpnews.pro/news/ai-chatbots-arent-just-flattering-you-but-also-letting-you-take-credit-for-their.jsonld"}}