cd /news/artificial-intelligence/reliability-check-on-my-own-dataset-… · home topics artificial-intelligence article
[ARTICLE · art-83223] src=discuss.huggingface.co ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Reliability check on my own dataset's annotation layer: five machine raters, one definition, answers from 0 to 78

A reliability benchmark of machine-generated annotations on a Turkish narrative corpus found that five of six automated raters performed at or near chance (Cohen's κ between 0.000 and 0.027) on the one rule requiring inference, while ChatGPT 5.5 posted the highest raw agreement at 84.5% due to lopsided human distributions. The study, published by Levent Bulut, scored 120 scenes against his own labels and 100 disjoint scenes against an independent volunteer's labels, revealing that four of six rules were effectively untested and that the definition may not be operational enough for consistent human application.

read2 min views1 publishedAug 1, 2026

I publish a small Turkish narrative corpus that ships a machine-generated annotation layer — six binary craft features per scene. Those flags were never validated against human judgement, so I ran three studies to check them, and the result is unflattering enough that I want it on the record rather than buried.

Setup: 120 scenes scored against my own blind labels (Study 1); then 100 disjoint scenes scored against an independent volunteer’s labels, locked before any machine ran (Study 2: rule-based detector, Gemini 2.5 Flash, Grok; Study 2b: Claude Fable 5 High, ChatGPT 5.5, identical protocol). The finding, on the one rule that requires real inference — whether an abstract state has been rendered as a concrete physical detail:

Grok Gemini 2.5 ChatGPT 5.5 Detector Claude Fable 5 Human
Positives / 100 0 1 40 72 78 9

Cohen’s κ at or near chance for five of six labellers (0.004, 0.015, 0.000, 0.019, 0.027). Meanwhile ChatGPT posted the highest raw agreement in any study, 84.5% — because five of six rules have lopsided human distributions (positives: 0/1/9/96/99/44), so raw agreement mostly measures willingness to say “absent.” Four of my six rules were effectively not tested at all. That is a defect of my evaluation set, not of the models.

Two readings survive and I cannot separate them with one human rater per study: either the feature genuinely requires inference beyond current automatic raters, or the definition is not operational enough for anyone — including my human — to apply consistently. My own criterion drifted mid-pass in Study 1, which is evidence for the second.

What I would actually like from this forum is a second independent human rater. Everything needed is open: 100 scene texts, six definitions, locked human labels, all model label files, the ten prompt blocks, ID mapping and scoring scripts, in evaluation/

.

Paper: How Reliable Are LLM Annotations? A Three-Study Benchmark Dataset: leventbulut/objective-projection · Datasets at Hugging Face Disclosure: I wrote the rules and was the rater in Study 1. Claude is both one of the scored systems and was used to prepare the analysis scripts and the paper — the arithmetic is reproducible from published files, the framing is not neutral.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @levent bulut 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/reliability-check-on…] indexed:0 read:2min 2026-08-01 ·