{"slug": "what-counts-as-a-mistake-annotating-recitation-events-in-quran-memorization", "title": "What Counts as a Mistake? Annotating Recitation Events in Quran Memorization Transcripts", "summary": "A completed human annotation of 100 Quran recitation recording cases produced 348 scored units and 162 localized events across ten combined labels, according to an arXiv paper (2609.12085v1) on distinguishing unresolved mistakes from repetitions, repairs, opening formulas and accepted spelling differences in ASR transcripts. A plain diff reached label-aware F1 0.525 and localization F1 0.826, while adapted production cleaner/alignment components reached 0.518 and 0.786, with exact-span F1 0.505 for both. In a preliminary pilot of eight single 20-minute runs across three coding agents and eight models, label-aware F1 ranged from 0.143 to 0.892, with 970 of 972 gold-event instances drawing an overlapping prediction, leaving span extent and adjudication-stipulated label boundaries as the remaining convention problems.", "body_md": "arXiv:2609.12085v1 Announce Type: new \nAbstract: Checking Quran recitation from an ASR transcript requires distinguishing unresolved mistakes from repetitions, repairs, opening formulas and accepted spelling differences. We report a completed human annotation of 100 production recording cases: 348 scored units and 162 localized events across ten combined labels. An executable evaluator scores labels and word positions together. A plain diff reaches label-aware F1 0.525 and localization F1 0.826; adapted production cleaner/alignment components reach 0.518 and 0.786, with exact-span F1 0.505 for both. Correcting the adapter's word coordinates recovers all five annotated repetition events, showing why annotation interfaces must be checked before interpreting baseline failures. In a preliminary pilot, eight single 20-minute runs across three coding agents and eight models span label-aware F1 0.143 to 0.892: seven land far above every baseline, and one collapses below the naive diff from a missing normalization step. Across the six, 970 of 972 gold-event instances draw an overlapping prediction, so what remains is not detection but convention: span extent, and the labels whose boundary is stipulated by adjudication rather than visible in the text. Seven of 162 events defeat all six same-day runs, five of them one orthographic rule, and the strongest run still misses the same ones. No run annotated before building, so the pilot measures the algorithm half of the task only.", "url": "https://wpnews.pro/news/what-counts-as-a-mistake-annotating-recitation-events-in-quran-memorization", "canonical_source": "https://arxiv.org/abs/2609.12085", "published_at": "2026-09-14 04:00:00+00:00", "updated_at": "2026-09-14 04:27:39.806039+00:00", "lang": "en", "topics": ["natural-language-processing", "machine-learning", "ai-research"], "entities": ["arXiv", "Quran"], "alternates": {"html": "https://wpnews.pro/news/what-counts-as-a-mistake-annotating-recitation-events-in-quran-memorization", "markdown": "https://wpnews.pro/news/what-counts-as-a-mistake-annotating-recitation-events-in-quran-memorization.md", "text": "https://wpnews.pro/news/what-counts-as-a-mistake-annotating-recitation-events-in-quran-memorization.txt", "jsonld": "https://wpnews.pro/news/what-counts-as-a-mistake-annotating-recitation-events-in-quran-memorization.jsonld"}}