{"slug": "interpretable-fairly-evaluated-automated-l2-speaking-assessment-that-beats-the", "title": "Interpretable, Fairly Evaluated Automated L2 Speaking Assessment that Beats the Single-Human Ceiling and Why Pause Encoding Does Not Change LLM Fluency Scores", "summary": "A new hybrid scoring system for automated second-language speaking assessment, combining interpretable speech-timing features with a text-LLM fluency judgment, achieved a Spearman rho of 0.818 against consensus human ratings on 130 L2 speeches from the ICNALE Global Rating Archive, outperforming 81% of 80 individual trained raters and beating the single-human ceiling. The study, posted on arXiv (2608.26137v1), also found that pause encoding in prompts does not significantly change LLM fluency scores, with effects bounded below about +/-0.1 rho.", "body_md": "arXiv:2608.26137v1 Announce Type: new\nAbstract: Second-language (L2) English learners can rarely rehearse speaking with a partner. Speaking is also the most anxiety-laden skill. These gaps drive a fast-growing market for automated speaking practice and scoring. But an automated score is trustworthy only if it is accurate, interpretable, fair, and benchmarked against the right human bar. We build an interpretable feature-plus-LLM hybrid for spontaneous L2 dialogue. We evaluate it without ever fitting to the human labels, against the ICNALE Global Rating Archive: 140 speeches rated by ~80 trained raters on 10 analytic criteria. We score the 130 L2 speeches with usable audio. A deterministic De-Jong speech-timing composite reaches rho=0.764. Blended with a single text-LLM fluency judgment, it reaches Spearman rho=0.818 against the consensus gold. This agrees with the consensus better than 81% of the 80 individual trained raters: above the median rater (rho=0.73) and near the best, and at ~83% of the reliability-corrected maximum (kappa_max=0.99). The blend improves on the composite alone by +0.054 (paired-bootstrap 95% CI [0.017, 0.108], excludes 0); the LLM adds a coarse fluency ranking that the continuous composite refines. We also report a controlled null on pause encoding, bounded to effects below about +/-0.1 rho at this sample size. Holding the LLM and learner words fixed and varying only how pauses are written into the prompt, inline pause locations do not beat aggregate pause statistics (-0.069, CI [-0.15, +0.08]), and a grounded mid-clause criterion gives no reliable gain. The fluency signal comes from the measured speech-timing features, not from how pauses are written for the LLM. We back every claim with two agreeing learner-isolation methods, paired-bootstrap CIs, a monologue negative control, per-feature reproduction of classical measurements, and a per-L1 fairness audit.", "url": "https://wpnews.pro/news/interpretable-fairly-evaluated-automated-l2-speaking-assessment-that-beats-the", "canonical_source": "https://arxiv.org/abs/2608.26137", "published_at": "2026-08-28 04:00:00+00:00", "updated_at": "2026-08-28 04:20:31.883628+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-products"], "entities": ["arXiv", "ICNALE Global Rating Archive", "De-Jong"], "alternates": {"html": "https://wpnews.pro/news/interpretable-fairly-evaluated-automated-l2-speaking-assessment-that-beats-the", "markdown": "https://wpnews.pro/news/interpretable-fairly-evaluated-automated-l2-speaking-assessment-that-beats-the.md", "text": "https://wpnews.pro/news/interpretable-fairly-evaluated-automated-l2-speaking-assessment-that-beats-the.txt", "jsonld": "https://wpnews.pro/news/interpretable-fairly-evaluated-automated-l2-speaking-assessment-that-beats-the.jsonld"}}