cd /news/artificial-intelligence/claude-and-chatgpt-vs-a-human-annota… · home topics artificial-intelligence article
[ARTICLE · art-73376] src=discuss.huggingface.co ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Claude and ChatGPT vs. a human annotator on "show, don't tell" features one definition, four very different thresholds

A new study testing Claude Fable 5 and ChatGPT 5.5 on the 'show, don't tell' narrative feature 'materialized metaphor' found that Claude identified 78 positive instances out of 100, ChatGPT identified 40, while earlier models Grok and Gemini found 0 and ~2 respectively, all with near-zero inter-rater agreement (κ ≈ 0), suggesting the definition itself is problematic. The study, conducted by researcher Levent Bulut and published on Hugging Face, also corrected an earlier claim that atmosphere contradiction was machine-proof, as Claude achieved κ 0.27 and ChatGPT 0.18 on that feature.

read1 min views1 publishedJul 25, 2026

Following up on an earlier held-out study (rule-based detector + Gemini + Grok vs. an independent human on 100 scenes), we ran the same test on Claude Fable 5 and ChatGPT 5.5 — same scenes, same locked human reference, same prompts.

Headline: on the hardest inferential feature (materialized metaphor; human found 9/100), the four LLMs’ positive counts were Grok 0, Gemini ~2, ChatGPT 40, Claude 78. Every κ ≈ 0 — the fifth machine rater in a row — but the spread itself suggests the definition, not only the models, is part of the problem. One earlier claim also gets corrected: atmosphere contradiction turned out not to be machine-proof (Claude κ 0.27, ChatGPT 0.18 vs. near-zero for the earlier raters).

Full write-up, both models’ complete label sets, and the deterministic scoring script: We Added Claude and ChatGPT to the "Show, Don't Tell" Detection Test. The Wall Held — But It Has Two Sides.

Conflict of interest is declared prominently in the post (one of the evaluated models is Claude; so is the analysis assistant). Everything is recomputable without any model in the loop. Dataset: leventbulut/objective-projection

(DOI 10.57967/hf/8960, CC BY-NC-ND). Official: https://leventbulut.com/ We are still looking for a second independent human rater — disagreement with the existing labels is the most useful possible contribution.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @claude fable 5 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/claude-and-chatgpt-v…] indexed:0 read:1min 2026-07-25 ·