{"slug": "edu-qurating-multi-dimensional-educational-data-curation-with-distilled-pairwise", "title": "Edu-QuRating: Multi-Dimensional Educational Data Curation with Distilled Pairwise Judgements", "summary": "Researchers introduced Edu-QuRating, a pipeline that scores and curates educational text on multiple education-specific criteria by distilling LLM-judge pairwise preferences into reusable Edu-QuRaters, with the best scorer recovering held-out GPT-4.1-mini pairwise judgements at mean accuracy 0.917 across two sequence-classification base models and six criteria. The team scored 322.25M FineWeb-Edu-Fortified documents to build a filtered pre-training mixture, and in matched single-run pre-training comparisons models trained on Edu-QuRating-based mixtures reached higher observed aggregate accuracy across nine benchmarks than the FineWeb-Edu baseline. Edu-QuRater scores were also used as reward terms for GRPO post-training, where combining Edu-QuRater and answer-structure rewards produced responses preferred to the Qwen3-4B base model on both pedagogical quality and instruction following.", "body_md": "arXiv:2609.09425v1 Announce Type: new \nAbstract: Educational data filters have become a practical way to improve language-model pre-training, but most filters treat educational value as a single scalar property. This may be too broad for some applications, especially if the data set already features a high density of educational material. Useful learning material needs to be accurate, engaging, well structured, and appropriate for the intended audience and application (e.g. learner- vs teacher-facing). Following QuRating (Wettig et al. 2024), we introduce Edu-QuRating: a pipeline for multi-dimensional educational data scoring and curation. Edu-QuRating defines education-specific rubrics, uses an LLM judge to label sampled document pairs and distills those pairwise preferences into reusable Edu-QuRaters, which can score individual text chunks on a set of educational criteria. Across two sequence-classification base models and six educational criteria, the best Edu-QuRater recovers held-out GPT-4.1-mini pairwise judgements with mean accuracy 0.917. We then apply the resulting scorers in two applications. First, we investigate the potential of Edu-QuRaters for corpus filtering to improve pretraining of small language models. We scored 322.25M FineWeb-Edu-Fortified documents to obtain a filtered pre-training mixture. In matched single-run pre-training comparisons, models trained with Edu-QuRating-based mixtures reached higher observed aggregate accuracy across nine benchmarks than the FineWeb-Edu baseline, with gains concentrated in particular tasks. Second, we used Edu-QuRater scores as reward terms for GRPO post-training. In held-out pairwise judge evaluations, combining Edu-QuRater and answer-structure rewards produced responses preferred to the Qwen3-4B base model on both pedagogical quality and instruction following.", "url": "https://wpnews.pro/news/edu-qurating-multi-dimensional-educational-data-curation-with-distilled-pairwise", "canonical_source": "https://arxiv.org/abs/2609.09425", "published_at": "2026-09-10 04:00:00+00:00", "updated_at": "2026-09-10 04:19:34.180500+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "machine-learning", "natural-language-processing"], "entities": ["Edu-QuRating", "QuRating", "Wettig et al. 2024", "GPT-4.1-mini", "FineWeb-Edu-Fortified", "FineWeb-Edu", "Qwen3-4B", "GRPO"], "alternates": {"html": "https://wpnews.pro/news/edu-qurating-multi-dimensional-educational-data-curation-with-distilled-pairwise", "markdown": "https://wpnews.pro/news/edu-qurating-multi-dimensional-educational-data-curation-with-distilled-pairwise.md", "text": "https://wpnews.pro/news/edu-qurating-multi-dimensional-educational-data-curation-with-distilled-pairwise.txt", "jsonld": "https://wpnews.pro/news/edu-qurating-multi-dimensional-educational-data-curation-with-distilled-pairwise.jsonld"}}