{"slug": "grammar-concept-annotation-at-scale-deployed-fine-tuned-small-language-models", "title": "Grammar Concept Annotation at Scale: Deployed Fine-Tuned Small Language Models Outperform Prompted Frontier Models", "summary": "A deployed 0.8B-parameter Qwen3.5 small language model fine-tuned on filtered, rebalanced teacher-generated supervision outperformed prompted GPT-5.4 and GPT-5.6 Sol in precision and recall on two human-curated grammar concept annotation benchmarks, according to arXiv paper 2610.10827v1. The 0.8B model, deployed in an end-to-end grammar mastery tracker for all English learners on the platform, cut serving cost by approximately 16x, and a feature-level online experiment showed gains of +15.8% in learner engagement, +2.1% in scheduled hours, and +13.2% in GMV from new lessons. A 4B reference comparator also beat the prompted frontier models under nested matching criteria of increasing strictness: concept, evidence span, and correctness.", "body_md": "arXiv:2610.10827v1 Announce Type: new \nAbstract: Corrective feedback is among the best-evidenced drivers of second-language acquisition, yet corrections delivered during lessons rarely accumulate into an actionable view of grammar mastery. Prompted frontier models can provide such a view from learner--tutor lesson transcripts, but they are costly at scale. We close this gap by fine-tuning Qwen3.5 small language models (SLMs) on filtered and rebalanced teacher-generated supervision, then deploying an efficient 0.8B model in an end-to-end grammar mastery tracker for all English learners on our platform. Internalizing the annotation contract into adapter weights enables pairing the 0.8B model with a compact matched prompt rather than verbose instructions. On two human-curated benchmarks, both the deployed 0.8B model and a 4B reference comparator outperform prompted GPT-5.4 and GPT-5.6 Sol in precision and recall under nested matching criteria of increasing strictness: concept, evidence span, and correctness. The deployed 0.8B SLM reduces serving cost by approximately 16$\\times$. A feature-level online experiment shows significant gains in learner engagement ($+15.8\\%$) and key business metrics, including scheduled hours ($+2.1\\%$) and GMV from new lessons ($+13.2\\%$).", "url": "https://wpnews.pro/news/grammar-concept-annotation-at-scale-deployed-fine-tuned-small-language-models", "canonical_source": "https://arxiv.org/abs/2610.10827", "published_at": "2026-10-09 04:00:00+00:00", "updated_at": "2026-10-09 04:17:52.964399+00:00", "lang": "en", "topics": ["natural-language-processing", "large-language-models", "machine-learning", "ai-research"], "entities": ["Qwen3.5", "GPT-5.4", "GPT-5.6 Sol", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/grammar-concept-annotation-at-scale-deployed-fine-tuned-small-language-models", "markdown": "https://wpnews.pro/news/grammar-concept-annotation-at-scale-deployed-fine-tuned-small-language-models.md", "text": "https://wpnews.pro/news/grammar-concept-annotation-at-scale-deployed-fine-tuned-small-language-models.txt", "jsonld": "https://wpnews.pro/news/grammar-concept-annotation-at-scale-deployed-fine-tuned-small-language-models.jsonld"}}