# Claude and ChatGPT vs. a human annotator on "show, don't tell" features one definition, four very different thresholds

> Source: <https://discuss.huggingface.co/t/claude-and-chatgpt-vs-a-human-annotator-on-show-dont-tell-features-one-definition-four-very-different-thresholds/178204#post_1>
> Published: 2026-07-25 13:54:52+00:00

Following up on an earlier held-out study (rule-based detector + Gemini + Grok vs. an independent human on 100 scenes), we ran the same test on Claude Fable 5 and ChatGPT 5.5 — same scenes, same locked human reference, same prompts.

Headline: on the hardest inferential feature (materialized metaphor; human found 9/100), the four LLMs’ positive counts were **Grok 0, Gemini ~2, ChatGPT 40, Claude 78**. Every κ ≈ 0 — the fifth machine rater in a row — but the spread itself suggests the definition, not only the models, is part of the problem. One earlier claim also gets corrected: atmosphere contradiction turned out *not* to be machine-proof (Claude κ 0.27, ChatGPT 0.18 vs. near-zero for the earlier raters).

Full write-up, both models’ complete label sets, and the deterministic scoring script: [We Added Claude and ChatGPT to the "Show, Don't Tell" Detection Test. The Wall Held — But It Has Two Sides.](https://huggingface.co/blog/leventbulut/claude-vs-chatgpt-on-narrative-analysis)

Conflict of interest is declared prominently in the post (one of the evaluated models is Claude; so is the analysis assistant). Everything is recomputable without any model in the loop. Dataset: `leventbulut/objective-projection`

(DOI 10.57967/hf/8960, CC BY-NC-ND). Official: [https://leventbulut.com/](https://leventbulut.com/)

We are still looking for a second independent human rater — disagreement with the existing labels is the most useful possible contribution.
