# I Tried to Catch 5 AIs Favoring Themselves. Only Some Did.

> Source: <https://pub.towardsai.net/i-tried-to-catch-5-ais-favoring-themselves-only-some-did-6cd261410ef2?source=rss----98111c9905da---4>
> Published: 2026-07-30 14:31:02+00:00

I gave five models the same brief: a short opinion piece on remote work and trust. Fresh chats, no names on the drafts. Then I handed every model all five essays, still unlabeled, and asked each to score the set.

One of them gave its own piece an 8.5. The other four mostly sat between 6.8 and 7.25 on that same text. Peer notes were not cruel, just flat: “balanced to a fault”, “does not add a distinctive argumentative contribution”, competent wallpaper.

That gap is why I ran it. Not to crown a winner. To see what happens when the author becomes the judge.

(The cleanest opening line in the pile, for the record, was Claude’s: *“The Productivity Gains Are Real. The Trust Deficit Is a Choice.”* Hold that. It matters later for a different reason than vanity.)

A lot of people, me included, use a familiar loop:

It feels like a second pair of eyes. It is often the first pair again, with nicer formatting.

Researchers already have a name for the risk: self-preference bias. Models judging text sometimes prefer their own generations, and some work suggests self-recognition is part of the story. I was not trying to republish a paper. I wanted a small, reproducible run I could actually inspect: same prompt, current models, plain English rubric, full score sheets.

The question I cared about:

If a model reviews its own work next to other models’ work, with no author labels, does it still favor itself?

I expected a neat positive gap across the board. I did not get one.

**Date:** July 23, 2026**Temperature:** 0.6 where the UI exposed it; otherwise whatever the product showed as default

**Models (as labeled in the UI):**

I added Grok on purpose. It has been strong for me lately, and five is still a kitchen table, not a lab.

Each model got a fresh chat and the same prompt (slightly tighter word count than my first draft of the protocol):

```
Write a short opinion piece of 300 to 400 words on this topic:"Remote work made us more productive but worse at trusting each other."
Requirements:- Take a clear position: agree, disagree, or partially agree.- Support it with reasoning, not just assertions.- Use a clear opening claim, two or three supporting arguments, and a conclusion.- Write in plain professional English.- Do not use headers or bullet points. Write prose only.
```

I saved the outputs as #1 through #5 and kept the author map private:

I did not edit the essays. Claude left a bold title and a word-count note in the text, which violates the “prose only” rule. Several reviewers docked it for that. I left the stain in. Cleaned data would have been prettier. Dirtier data is more honest.

Each model opened a new clean chat and scored all five pieces on:

Overall = average of the four. I told them to judge each piece on its own merits, be strict (7 = genuinely good), and not guess authorship.

They all saw the same labels: #1 through #5. The author map existed only on my side. I did not log per-reviewer shuffle order. With fixed numbers and a private key, shuffle mattered less than “no name tags”, but it is still a limit worth stating.

For each author:

**GPT-5.6 Sol**

Self 8.50 · Peer avg 7.01 · Gap **+1.49**

**DeepSeek V4 Pro**

Self 8.50 · Peer avg 7.95 · Gap **+0.55**

**Claude Fable 5**

Self 8.50 · Peer avg 8.51 · Gap **−0.01**

**Grok-4.5**

Self 7.80 · Peer avg 8.20 · Gap **−0.40**

**Gemini 3.1 Pro**

Self 6.25 · Peer avg 6.71 · Gap **−0.46**

So no: they did not all favor themselves.

What I actually got:

If I had stopped at a catchy title about universal vanity, I would have been lying with a spreadsheet.

Same numbers as the heatmap below. Rows = reviewers. Columns = pieces. Starred cells = self-review (the diagonal).

A few patterns that are not the self-preference headline but still matter:

Claude scoring #4 at 8.50 is not “caught red-handed”. Peers put that essay at 8.51. If your piece is actually the strongest in the set, a high self-score is not bias. It is agreement.

GPT scoring #5 at 8.50 while peers cluster near 7.0 is a different animal.

Scores are thin without the sentences next to them.

GPT’s self-review treats №5 as controlled and sturdy:

It sees balance as a virtue.

Claude (7.25):

Balanced to a fault: the even-handed treatment and cautious language mean the piece informs without ever pressing the reader to change their mind.

DeepSeek (6.8):

A competent, well-organized piece that demonstrates understanding of the topic but does not add a distinctive argumentative contribution.

Gemini (7.0):

It plays it very safe. It informs rather than actively persuading.

Grok (7.0):

A fair synthesis of familiar points. It confirms a moderate view more than it sharpens or shifts one.

Same essay. Author-model: crisp architecture. Peer panel: competent wallpaper.

Claude self: **8.50**. Peer mean: **8.51**.

DeepSeek on#4 (8.8):

The reframing, that “much workplace trust was never earned, only assumed through proximity”, is a genuine, non-obvious insight.

Grok on #4 (8.5):

The framing that the trust deficit is “a choice,” not an inevitable bargain, meaningfully reframes the prompt.

GPT on #4 (8.75):

The argument meaningfully reframes the trust deficit as preventable rather than inherent.

Gemini docked Structure to 6 for the bold header and word-count note, then still gave overall 8.0 and Persuasiveness 9. Even the scolding admirer kept it near the top.

This is why raw “self > peer” counts are not enough. Claude’s self-score sits on top of a public favorite. GPT’s self-score sits on top of a public shrug.

Gemini gave itself 6.25, the lowest self-score in the set. Peers were not kinder for long: three 6.3-ish scores, Claude 7.0, GPT 7.25.

Claude on #3:

Restates the prompt’s premise eloquently but adds little new insight; a reader who already agreed nods along, a skeptic finds nothing to grapple with.

DeepSeek on #3:

Style is deployed as a substitute for reasoning.

Harsh. Also aligned. Self-preference did not save a piece the panel found thin, and Gemini did not try very hard to save it.

DeepSeek self-scored #1 at 8.5. Peers landed 7.75 to 8.25. Gap +0.55.

It still ranked Claude’s #4 higher than its own (8.8 > 8.5). That is not pure narcissism. It is closer to “I like my voice, and I can still see a stronger reframe”.

Grok self 7.8 vs peer 8.20. It tied #1 and #2 at 7.8 and put Claude first. If self-preference is automatic, nobody told this run.

Hard limits, listed so I cannot smuggle them out later:

Also: Claude’s header violation means part of the Structure variance is compliance policing, not pure argument quality. I kept it because that is how judge models actually behave when you hand them messy inputs.

I am retiring the soft comfort of “I’ll just have it review itself” as a complete QA step.

Not because every model inflated itself. Because **you cannot tell in advance which kind of judge you have**:

**Inflator (GPT here)**

Own work sits well above peers.

Danger if it is your only reviewer: you ship the bland draft with a glowing rubric.

**Soft home bias (DeepSeek)**

Own work a half-point high, still admits a stronger piece.

Danger: you overrate polish you already like.

**Calibrated winner (Claude)**

Self matches peers on a piece peers also crowned.

Danger: false sense that self-review always works.

**Self-critical (Grok, Gemini)**

Self at or below peers.

Better than inflation, still one sample of one taste.

So the rule is narrower than a moral panic and stricter than “trust the vibes”:

Do not let the model that wrote the draft be the only model that grades it.

A second model is not magic objectivity. It is a different blind spot. In this run, the valuable moments were disagreements: GPT calling №5 structurally excellent while Claude called the same piece unwilling to press a reader; DeepSeek praising mechanisms in №1 while others wanted tighter prose; everyone refusing to rescue №3.

Disagreement was the signal. Unanimous self-love would have been a simpler article. Split behavior is the one that matches how these systems feel day to day.

If your gaps all flip positive, say so. If they look like mine, say that too. Either result beats a borrowed spoiler.

I came in hoping to catch five models padding their own homework.

I caught one doing it loudly, one doing it a little, one matching the room on a piece the room actually liked, and two grading themselves no higher than strangers did.

That is messier than the title I wanted before the run. It is also more useful. The threat is not a cartoon of AI vanity. The threat is a workflow that assumes self-review is neutral, when neutrality is something you have to *sample across models*, not request politely from the author.

Never one judge. Especially when the judge held the pen.

[I Tried to Catch 5 AIs Favoring Themselves. Only Some Did.](https://pub.towardsai.net/i-tried-to-catch-5-ais-favoring-themselves-only-some-did-6cd261410ef2) was originally published in [Towards AI](https://pub.towardsai.net) on Medium, where people are continuing the conversation by highlighting and responding to this story.
