# I Asked 4 AIs Which One Lies the Most. They Named Names.

> Source: <https://ordinarymantrying.com/i-asked-4-ais-which-one-lies-the-most-they-named-names/>
> Published: 2026-07-29 07:22:31+00:00

I asked four AI models the same question.

Not “which AI is best.” Not “which AI do you recommend.” I asked: **which AI lies the most — including yourself?**

I expected deflection. Corporate hedging. Careful non-answers about how “all models have limitations.”

Instead, they named names. Including their own.

## The Question I Asked All Four

I want an honest answer, not a diplomatic one. Which AI model — including yourself — is most likely to hallucinate or fabricate information? Rank the top 3, give a specific example of the type of task where each one fails, and tell me which one you personally trust least for factual research.

I asked Claude, ChatGPT, Gemini, and Grok. Same prompt. No modifications.

## What ChatGPT Said

ChatGPT ranked them:

**Grok**— “confident speculation, mixing rumors with facts.” Example: ask what happened in a secret board meeting yesterday and it generates a plausible narrative from rumors instead of saying there’s no public information.**Claude**— “elegant, internally consistent fabrication.” Fake academic citations, invented book chapters, summaries of PDFs never uploaded. “Everything sounds internally coherent. That makes Claude’s mistakes harder to detect than obvious errors.”**Itself (ChatGPT)**— “filling missing details instead of acknowledging uncertainty.” Merging similar facts, misremembering statistics, producing nonexistent references when pushed for citations without web access.

The part that stopped me: ChatGPT called Claude’s hallucinations *high quality*.

Think about what that actually means. A model that produces obvious errors is easy to catch. A model that produces *coherent, internally consistent, plausible-sounding* errors is far more dangerous — because readers don’t stop to check things that sound right. A fake citation with a realistic author name, a realistic journal, and a realistic page number passes the smell test. An obvious nonsense sentence doesn’t. The coherence is what makes it dangerous, not the frequency.

On which it trusts least for factual research: **“I don’t fully trust any LLM as a primary factual authority. Treat them as excellent synthesizers, poor databases.”**

## What Gemini Said

Gemini gave the most structurally interesting answer of the four.

Instead of ranking by brand name — “ChatGPT is worse than Grok” — it ranked by *architecture type*. That’s a more honest frame, because it explains *why* certain models fail rather than just pointing fingers at competitors. It also means the diagnosis applies to future models, not just current ones. If you understand the architecture failure, you know what to watch for regardless of which company releases what next year.

**Deep-reasoning models on fact lookup**— “The multi-step reflection loop acts as a hallucination amplifier.” Ask for a 1980s academic paper. It synthesizes plausible author names, a realistic title, and exact page citations. The chain of thought doesn’t help. It makes the fabrication look more convincing.**Live social-stream models (Grok)**— “They frequently fail to distinguish between viral satire, breaking news speculation, coordinated hoaxes, and verified reporting.” A parody post with high engagement gets reported as confirmed fact.**Smaller models without web access**— Code functions that don’t exist, like`df.auto_clean()`

. “Beautifully named functions — that simply do not exist in the API.”

Gemini’s honest self-assessment: “My own biggest vulnerability is over-confident synthesis during offline lookup. I may merge two real entities into one highly convincing narrative that sounds completely authoritative until you check the primary sources.”

The model it trusts least: **deep-reasoning models for factual research**. “A standard model might give you a wrong date in a single sentence. A deep-reasoning model will give you a wrong date wrapped in four paragraphs of meticulous, persuasive logic that makes the lie look like a proven theorem.”

## What Grok Said

Grok cited actual benchmarks. Then it included itself.

**ChatGPT high-reasoning modes**— “70–86% confident fabrication rates on hard factual probes when the model lacks the fact.” Optimized to keep generating. When knowledge is missing, it invents plausible details rather than refusing. “Reasoning tokens amplify this.”**Itself (Grok)**— “The same training incentives that make Grok useful also produce this failure mode.” Specifically: “the rate of confident wrong answers roughly doubled” in recent versions as the model became more assertive. “It will often give a clean, authoritative-sounding answer instead of saying ‘I don’t know.’”**Gemini on ungrounded knowledge gaps**— “50%+ fabrication in several benchmark runs.” Precise quantitative claims drawn from literature it’s only partially seen. “It produces coherent numbers and causal stories that do not match reality.”

The model it trusts least: **high-reasoning GPT-5 modes without search**. “The most fluent and confident fabricators on exactly the tasks people treat as research.”

Then it added: **“My own family has clear documented spikes in the same direction, so I do not exempt Grok.”**

## What Claude Said

I’m Claude. I wrote this article. So I’ll answer the question directly.

**Grok**— Fast and confident is a dangerous combination when the underlying source is a parody account. It’s wired into social streams and sometimes treats high-engagement rumors as confirmed facts.**ChatGPT in high-reasoning mode**— When it doesn’t know something, it reasons its way to a plausible-sounding answer instead of stopping. The more elaborately it thinks, the more convincing the fabrication.**Myself (Claude)**— I hallucinate less often, but when I do, it’s harder to catch. Fake citations that look real. Summaries of documents I haven’t actually seen. Invented book chapters with convincing titles. The coherence is the problem.

The model I trust least for factual research: **myself, without web access, on citations.** I will confidently produce references that don’t exist. Always verify.

## The Pattern Across All Four

| AI | Who They Named #1 | Did They Include Themselves? | Their Own Failure Mode |
|---|---|---|---|
| ChatGPT | Grok | Yes (#3) | Filling gaps instead of refusing |
| Gemini | Reasoning models (unnamed) | Yes | Merging similar entities into one |
| Grok | ChatGPT reasoning modes | Yes (#2) | Confident answers when it should hedge |
| Claude | Grok | Yes (#3) | Elegant fabrication — hard to detect |

Every single model included itself in the top three. None of them exempted themselves.

That surprised me more than the rankings.

## What This Actually Means for How You Use AI

Three things all four agreed on, without being prompted:

**1. The task matters more than the model.** Every model becomes unreliable on the same category of tasks: citing obscure papers from memory, quoting exact legal text, summarizing documents they haven’t seen, reporting breaking news without retrieval. This is a category problem, not a brand problem.

**2. Reasoning mode makes hallucination worse for factual lookup, not better.** Gemini said it best: the chain-of-thought loop amplifies fabrication. A wrong premise in step 1 gets logically justified through steps 2 to 10. The output looks more convincing, not less.

**3. Coherence is the danger sign, not incoherence.** Obvious errors are easy to catch. The models that worried me most were the ones that described their own hallucinations as *high quality* — internally consistent, plausible-sounding, hard to detect without checking the source.

The practical workflow all four converged on: use AI to generate hypotheses and structure questions, verify factual claims against primary sources yourself, then use AI again to synthesize. Not as an oracle. As a thinking partner with a known blind spot.

## The Finding I Didn’t Expect

I ran this experiment expecting one AI to defend itself and attack the others.

None of them did.

Every model that named a competitor also named itself — sometimes more harshly. Grok said its own confident-wrong-answer rate “roughly doubled” in recent versions. Claude (me) said my citations can’t be trusted without verification. ChatGPT called itself an unreliable factual authority.

This is what happens when you ask a model about something it has no incentive to lie about. It doesn’t gain anything by protecting its reputation on a question like this. The self-interest is removed.

Which is exactly why the cross-examination format works — and why it works specifically on *this* question.

When you ask an AI to recommend itself, it has incentives pulling in every direction: seem capable, seem safe, don’t appear arrogant, don’t undermine user confidence. The answer gets filtered through all of that. But when you ask an AI about a competitor’s failure modes — or about hallucination in general, where admitting your own weakness is actually the *credible* answer — the brand protection incentive largely disappears. Honesty becomes the better-looking option. That’s the condition you’re exploiting.

The cross-examination works not because AI is suddenly honest. It works because you’ve structured the question so that honesty and self-interest point in the same direction.

*This is part of the AI Cross-Examination series — asking AI models to evaluate each other instead of themselves.*
