Following up on an earlier held-out study (rule-based detector + Gemini + Grok vs. an independent human on 100 scenes), we ran the same test on Claude Fable 5 and ChatGPT 5.5 — same scenes, same locked human reference, same prompts.
Headline: on the hardest inferential feature (materialized metaphor; human found 9/100), the four LLMs’ positive counts were Grok 0, Gemini ~2, ChatGPT 40, Claude 78. Every κ ≈ 0 — the fifth machine rater in a row — but the spread itself suggests the definition, not only the models, is part of the problem. One earlier claim also gets corrected: atmosphere contradiction turned out not to be machine-proof (Claude κ 0.27, ChatGPT 0.18 vs. near-zero for the earlier raters).
Full write-up, both models’ complete label sets, and the deterministic scoring script: We Added Claude and ChatGPT to the "Show, Don't Tell" Detection Test. The Wall Held — But It Has Two Sides.
Conflict of interest is declared prominently in the post (one of the evaluated models is Claude; so is the analysis assistant). Everything is recomputable without any model in the loop. Dataset: leventbulut/objective-projection
(DOI 10.57967/hf/8960, CC BY-NC-ND). Official: https://leventbulut.com/ We are still looking for a second independent human rater — disagreement with the existing labels is the most useful possible contribution.