{"slug": "claude-and-chatgpt-vs-a-human-annotator-on-show-don-t-tell-features-one-four", "title": "Claude and ChatGPT vs. a human annotator on \"show, don't tell\" features one definition, four very different thresholds", "summary": "A new study testing Claude Fable 5 and ChatGPT 5.5 on the 'show, don't tell' narrative feature 'materialized metaphor' found that Claude identified 78 positive instances out of 100, ChatGPT identified 40, while earlier models Grok and Gemini found 0 and ~2 respectively, all with near-zero inter-rater agreement (κ ≈ 0), suggesting the definition itself is problematic. The study, conducted by researcher Levent Bulut and published on Hugging Face, also corrected an earlier claim that atmosphere contradiction was machine-proof, as Claude achieved κ 0.27 and ChatGPT 0.18 on that feature.", "body_md": "Following up on an earlier held-out study (rule-based detector + Gemini + Grok vs. an independent human on 100 scenes), we ran the same test on Claude Fable 5 and ChatGPT 5.5 — same scenes, same locked human reference, same prompts.\n\nHeadline: on the hardest inferential feature (materialized metaphor; human found 9/100), the four LLMs’ positive counts were **Grok 0, Gemini ~2, ChatGPT 40, Claude 78**. Every κ ≈ 0 — the fifth machine rater in a row — but the spread itself suggests the definition, not only the models, is part of the problem. One earlier claim also gets corrected: atmosphere contradiction turned out *not* to be machine-proof (Claude κ 0.27, ChatGPT 0.18 vs. near-zero for the earlier raters).\n\nFull write-up, both models’ complete label sets, and the deterministic scoring script: [We Added Claude and ChatGPT to the \"Show, Don't Tell\" Detection Test. The Wall Held — But It Has Two Sides.](https://huggingface.co/blog/leventbulut/claude-vs-chatgpt-on-narrative-analysis)\n\nConflict of interest is declared prominently in the post (one of the evaluated models is Claude; so is the analysis assistant). Everything is recomputable without any model in the loop. Dataset: `leventbulut/objective-projection`\n\n(DOI 10.57967/hf/8960, CC BY-NC-ND). Official: [https://leventbulut.com/](https://leventbulut.com/)\n\nWe are still looking for a second independent human rater — disagreement with the existing labels is the most useful possible contribution.", "url": "https://wpnews.pro/news/claude-and-chatgpt-vs-a-human-annotator-on-show-don-t-tell-features-one-four", "canonical_source": "https://discuss.huggingface.co/t/claude-and-chatgpt-vs-a-human-annotator-on-show-dont-tell-features-one-definition-four-very-different-thresholds/178204#post_1", "published_at": "2026-07-25 13:54:52+00:00", "updated_at": "2026-07-25 14:03:13.565875+00:00", "lang": "en", "topics": ["artificial-intelligence", "natural-language-processing", "ai-research"], "entities": ["Claude Fable 5", "ChatGPT 5.5", "Grok", "Gemini", "Levent Bulut", "Hugging Face"], "alternates": {"html": "https://wpnews.pro/news/claude-and-chatgpt-vs-a-human-annotator-on-show-don-t-tell-features-one-four", "markdown": "https://wpnews.pro/news/claude-and-chatgpt-vs-a-human-annotator-on-show-don-t-tell-features-one-four.md", "text": "https://wpnews.pro/news/claude-and-chatgpt-vs-a-human-annotator-on-show-don-t-tell-features-one-four.txt", "jsonld": "https://wpnews.pro/news/claude-and-chatgpt-vs-a-human-annotator-on-show-don-t-tell-features-one-four.jsonld"}}