cd /news/natural-language-processing/what-counts-as-a-mistake-annotating-… · home topics natural-language-processing article
[ARTICLE · art-128722] src=arxiv.org ↗ pub= topic=natural-language-processing verified=true sentiment=· neutral

What Counts as a Mistake? Annotating Recitation Events in Quran Memorization Transcripts

A completed human annotation of 100 Quran recitation recording cases produced 348 scored units and 162 localized events across ten combined labels, according to an arXiv paper (2609.12085v1) on distinguishing unresolved mistakes from repetitions, repairs, opening formulas and accepted spelling differences in ASR transcripts. A plain diff reached label-aware F1 0.525 and localization F1 0.826, while adapted production cleaner/alignment components reached 0.518 and 0.786, with exact-span F1 0.505 for both. In a preliminary pilot of eight single 20-minute runs across three coding agents and eight models, label-aware F1 ranged from 0.143 to 0.892, with 970 of 972 gold-event instances drawing an overlapping prediction, leaving span extent and adjudication-stipulated label boundaries as the remaining convention problems.

by read1 min views1 publishedSep 14, 2026

arXiv:2609.12085v1 Announce Type: new Abstract: Checking Quran recitation from an ASR transcript requires distinguishing unresolved mistakes from repetitions, repairs, opening formulas and accepted spelling differences. We report a completed human annotation of 100 production recording cases: 348 scored units and 162 localized events across ten combined labels. An executable evaluator scores labels and word positions together. A plain diff reaches label-aware F1 0.525 and localization F1 0.826; adapted production cleaner/alignment components reach 0.518 and 0.786, with exact-span F1 0.505 for both. Correcting the adapter's word coordinates recovers all five annotated repetition events, showing why annotation interfaces must be checked before interpreting baseline failures. In a preliminary pilot, eight single 20-minute runs across three coding agents and eight models span label-aware F1 0.143 to 0.892: seven land far above every baseline, and one collapses below the naive diff from a missing normalization step. Across the six, 970 of 972 gold-event instances draw an overlapping prediction, so what remains is not detection but convention: span extent, and the labels whose boundary is stipulated by adjudication rather than visible in the text. Seven of 162 events defeat all six same-day runs, five of them one orthographic rule, and the strongest run still misses the same ones. No run annotated before building, so the pilot measures the algorithm half of the task only.

── more in #natural-language-processing 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/what-counts-as-a-mis…] indexed:0 read:1min 2026-09-14 ·