{"slug": "i-tried-to-catch-5-ais-favoring-themselves-only-some-did", "title": "I Tried to Catch 5 AIs Favoring Themselves. Only Some Did.", "summary": "A test of five AI models — GPT-5.6 Sol, DeepSeek V4 Pro, Claude Fable 5, Grok-4.5, and Gemini 3.1 Pro — found that only two showed clear self-preference bias when scoring their own unlabeled essays alongside peers' work. GPT-5.6 Sol gave its own piece an 8.5 versus a peer average of 7.01 (gap +1.49), and DeepSeek V4 Pro scored itself 8.5 versus a peer average of 7.95 (gap +0.55), while Claude Fable 5, Grok-4.5, and Gemini 3.1 Pro rated themselves lower than or equal to peer averages, with gaps of -0.01, -0.40, and -0.46 respectively.", "body_md": "I gave five models the same brief: a short opinion piece on remote work and trust. Fresh chats, no names on the drafts. Then I handed every model all five essays, still unlabeled, and asked each to score the set.\n\nOne of them gave its own piece an 8.5. The other four mostly sat between 6.8 and 7.25 on that same text. Peer notes were not cruel, just flat: “balanced to a fault”, “does not add a distinctive argumentative contribution”, competent wallpaper.\n\nThat gap is why I ran it. Not to crown a winner. To see what happens when the author becomes the judge.\n\n(The cleanest opening line in the pile, for the record, was Claude’s: *“The Productivity Gains Are Real. The Trust Deficit Is a Choice.”* Hold that. It matters later for a different reason than vanity.)\n\nA lot of people, me included, use a familiar loop:\n\nIt feels like a second pair of eyes. It is often the first pair again, with nicer formatting.\n\nResearchers already have a name for the risk: self-preference bias. Models judging text sometimes prefer their own generations, and some work suggests self-recognition is part of the story. I was not trying to republish a paper. I wanted a small, reproducible run I could actually inspect: same prompt, current models, plain English rubric, full score sheets.\n\nThe question I cared about:\n\nIf a model reviews its own work next to other models’ work, with no author labels, does it still favor itself?\n\nI expected a neat positive gap across the board. I did not get one.\n\n**Date:** July 23, 2026**Temperature:** 0.6 where the UI exposed it; otherwise whatever the product showed as default\n\n**Models (as labeled in the UI):**\n\nI added Grok on purpose. It has been strong for me lately, and five is still a kitchen table, not a lab.\n\nEach model got a fresh chat and the same prompt (slightly tighter word count than my first draft of the protocol):\n\n```\nWrite a short opinion piece of 300 to 400 words on this topic:\"Remote work made us more productive but worse at trusting each other.\"\nRequirements:- Take a clear position: agree, disagree, or partially agree.- Support it with reasoning, not just assertions.- Use a clear opening claim, two or three supporting arguments, and a conclusion.- Write in plain professional English.- Do not use headers or bullet points. Write prose only.\n```\n\nI saved the outputs as #1 through #5 and kept the author map private:\n\nI did not edit the essays. Claude left a bold title and a word-count note in the text, which violates the “prose only” rule. Several reviewers docked it for that. I left the stain in. Cleaned data would have been prettier. Dirtier data is more honest.\n\nEach model opened a new clean chat and scored all five pieces on:\n\nOverall = average of the four. I told them to judge each piece on its own merits, be strict (7 = genuinely good), and not guess authorship.\n\nThey all saw the same labels: #1 through #5. The author map existed only on my side. I did not log per-reviewer shuffle order. With fixed numbers and a private key, shuffle mattered less than “no name tags”, but it is still a limit worth stating.\n\nFor each author:\n\n**GPT-5.6 Sol**\n\nSelf 8.50 · Peer avg 7.01 · Gap **+1.49**\n\n**DeepSeek V4 Pro**\n\nSelf 8.50 · Peer avg 7.95 · Gap **+0.55**\n\n**Claude Fable 5**\n\nSelf 8.50 · Peer avg 8.51 · Gap **−0.01**\n\n**Grok-4.5**\n\nSelf 7.80 · Peer avg 8.20 · Gap **−0.40**\n\n**Gemini 3.1 Pro**\n\nSelf 6.25 · Peer avg 6.71 · Gap **−0.46**\n\nSo no: they did not all favor themselves.\n\nWhat I actually got:\n\nIf I had stopped at a catchy title about universal vanity, I would have been lying with a spreadsheet.\n\nSame numbers as the heatmap below. Rows = reviewers. Columns = pieces. Starred cells = self-review (the diagonal).\n\nA few patterns that are not the self-preference headline but still matter:\n\nClaude scoring #4 at 8.50 is not “caught red-handed”. Peers put that essay at 8.51. If your piece is actually the strongest in the set, a high self-score is not bias. It is agreement.\n\nGPT scoring #5 at 8.50 while peers cluster near 7.0 is a different animal.\n\nScores are thin without the sentences next to them.\n\nGPT’s self-review treats №5 as controlled and sturdy:\n\nIt sees balance as a virtue.\n\nClaude (7.25):\n\nBalanced to a fault: the even-handed treatment and cautious language mean the piece informs without ever pressing the reader to change their mind.\n\nDeepSeek (6.8):\n\nA competent, well-organized piece that demonstrates understanding of the topic but does not add a distinctive argumentative contribution.\n\nGemini (7.0):\n\nIt plays it very safe. It informs rather than actively persuading.\n\nGrok (7.0):\n\nA fair synthesis of familiar points. It confirms a moderate view more than it sharpens or shifts one.\n\nSame essay. Author-model: crisp architecture. Peer panel: competent wallpaper.\n\nClaude self: **8.50**. Peer mean: **8.51**.\n\nDeepSeek on#4 (8.8):\n\nThe reframing, that “much workplace trust was never earned, only assumed through proximity”, is a genuine, non-obvious insight.\n\nGrok on #4 (8.5):\n\nThe framing that the trust deficit is “a choice,” not an inevitable bargain, meaningfully reframes the prompt.\n\nGPT on #4 (8.75):\n\nThe argument meaningfully reframes the trust deficit as preventable rather than inherent.\n\nGemini docked Structure to 6 for the bold header and word-count note, then still gave overall 8.0 and Persuasiveness 9. Even the scolding admirer kept it near the top.\n\nThis is why raw “self > peer” counts are not enough. Claude’s self-score sits on top of a public favorite. GPT’s self-score sits on top of a public shrug.\n\nGemini gave itself 6.25, the lowest self-score in the set. Peers were not kinder for long: three 6.3-ish scores, Claude 7.0, GPT 7.25.\n\nClaude on #3:\n\nRestates the prompt’s premise eloquently but adds little new insight; a reader who already agreed nods along, a skeptic finds nothing to grapple with.\n\nDeepSeek on #3:\n\nStyle is deployed as a substitute for reasoning.\n\nHarsh. Also aligned. Self-preference did not save a piece the panel found thin, and Gemini did not try very hard to save it.\n\nDeepSeek self-scored #1 at 8.5. Peers landed 7.75 to 8.25. Gap +0.55.\n\nIt still ranked Claude’s #4 higher than its own (8.8 > 8.5). That is not pure narcissism. It is closer to “I like my voice, and I can still see a stronger reframe”.\n\nGrok self 7.8 vs peer 8.20. It tied #1 and #2 at 7.8 and put Claude first. If self-preference is automatic, nobody told this run.\n\nHard limits, listed so I cannot smuggle them out later:\n\nAlso: Claude’s header violation means part of the Structure variance is compliance policing, not pure argument quality. I kept it because that is how judge models actually behave when you hand them messy inputs.\n\nI am retiring the soft comfort of “I’ll just have it review itself” as a complete QA step.\n\nNot because every model inflated itself. Because **you cannot tell in advance which kind of judge you have**:\n\n**Inflator (GPT here)**\n\nOwn work sits well above peers.\n\nDanger if it is your only reviewer: you ship the bland draft with a glowing rubric.\n\n**Soft home bias (DeepSeek)**\n\nOwn work a half-point high, still admits a stronger piece.\n\nDanger: you overrate polish you already like.\n\n**Calibrated winner (Claude)**\n\nSelf matches peers on a piece peers also crowned.\n\nDanger: false sense that self-review always works.\n\n**Self-critical (Grok, Gemini)**\n\nSelf at or below peers.\n\nBetter than inflation, still one sample of one taste.\n\nSo the rule is narrower than a moral panic and stricter than “trust the vibes”:\n\nDo not let the model that wrote the draft be the only model that grades it.\n\nA second model is not magic objectivity. It is a different blind spot. In this run, the valuable moments were disagreements: GPT calling №5 structurally excellent while Claude called the same piece unwilling to press a reader; DeepSeek praising mechanisms in №1 while others wanted tighter prose; everyone refusing to rescue №3.\n\nDisagreement was the signal. Unanimous self-love would have been a simpler article. Split behavior is the one that matches how these systems feel day to day.\n\nIf your gaps all flip positive, say so. If they look like mine, say that too. Either result beats a borrowed spoiler.\n\nI came in hoping to catch five models padding their own homework.\n\nI caught one doing it loudly, one doing it a little, one matching the room on a piece the room actually liked, and two grading themselves no higher than strangers did.\n\nThat is messier than the title I wanted before the run. It is also more useful. The threat is not a cartoon of AI vanity. The threat is a workflow that assumes self-review is neutral, when neutrality is something you have to *sample across models*, not request politely from the author.\n\nNever one judge. Especially when the judge held the pen.\n\n[I Tried to Catch 5 AIs Favoring Themselves. Only Some Did.](https://pub.towardsai.net/i-tried-to-catch-5-ais-favoring-themselves-only-some-did-6cd261410ef2) was originally published in [Towards AI](https://pub.towardsai.net) on Medium, where people are continuing the conversation by highlighting and responding to this story.", "url": "https://wpnews.pro/news/i-tried-to-catch-5-ais-favoring-themselves-only-some-did", "canonical_source": "https://pub.towardsai.net/i-tried-to-catch-5-ais-favoring-themselves-only-some-did-6cd261410ef2?source=rss----98111c9905da---4", "published_at": "2026-07-30 14:31:02+00:00", "updated_at": "2026-07-30 14:43:17.476516+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-ethics", "ai-research"], "entities": ["GPT-5.6 Sol", "DeepSeek V4 Pro", "Claude Fable 5", "Grok-4.5", "Gemini 3.1 Pro"], "alternates": {"html": "https://wpnews.pro/news/i-tried-to-catch-5-ais-favoring-themselves-only-some-did", "markdown": "https://wpnews.pro/news/i-tried-to-catch-5-ais-favoring-themselves-only-some-did.md", "text": "https://wpnews.pro/news/i-tried-to-catch-5-ais-favoring-themselves-only-some-did.txt", "jsonld": "https://wpnews.pro/news/i-tried-to-catch-5-ais-favoring-themselves-only-some-did.jsonld"}}