{"slug": "beyond-score-prediction-llm-based-essay-scoring-and-feedback-generation-via-with", "title": "Beyond Score Prediction: LLM-Based Essay Scoring and Feedback Generation via Reinforcement Learning with Rubric Rewards", "summary": "Researchers propose RLAES, a unified LLM framework that jointly optimizes automated essay scoring and feedback generation through reinforcement learning, achieving the best scoring performance among LLM-based methods on the ASAP benchmark with a QWK of 0.803. The framework introduces Rubric-based Feedback Evaluation (RFE) with 166 fine-grained binary rubric items and an LLM-as-judge to make feedback quality measurable, and Adaptive Gated Feedback Optimization (AGFO) to reduce evaluation overhead while improving feedback quality. The study also presents Adjacent Contrastive Reasoning (ACR) to improve ordinal score calibration by contrasting adjacent score levels.", "body_md": "arXiv:2607.19219v1 Announce Type: new\nAbstract: Large language models (LLMs) have been widely applied to automated essay scoring (AES) and automated feedback generation (AFG). However, existing studies rely primarily on prompt engineering or supervised fine-tuning, while systematic research on reinforcement learning (RL) post-training and automated evaluation of feedback quality remains limited. We propose RLAES, a unified LLM framework that jointly optimizes essay scoring and feedback generation through RL. To make feedback quality measurable, interpretable, and usable for training, we introduce Rubric-based Feedback Evaluation (RFE), an essay-grounded feedback evaluation framework comprising 166 fine-grained binary rubric items and an LLM-as-judge. Building on RFE, we propose Adaptive Gated Feedback Optimization (AGFO), which activates rubric-based feedback rewards on demand during RL, reducing evaluation overhead while improving feedback quality. We also propose Adjacent Contrastive Reasoning (ACR) to improve ordinal score calibration by explicitly contrasting adjacent score levels. Experimental results show that the RFE framework captures essay-feedback consistency, exhibits strong pairwise discriminative power, and closely aligns with expert preferences. On the ASAP benchmark, RLAES-AGFO achieves the best scoring performance among LLM-based methods (QWK = 0.803), while maintaining feedback quality comparable to GPT-5.5 and avoiding the feedback degradation observed under score-only RL. Code and datasets are publicly available at https://github.com/hellomuyi/RLAES.", "url": "https://wpnews.pro/news/beyond-score-prediction-llm-based-essay-scoring-and-feedback-generation-via-with", "canonical_source": "https://www.machinebrief.com/news/beyond-score-prediction-llm-based-essay-scoring-and-feedback-2vns", "published_at": "2026-07-22 04:00:00+00:00", "updated_at": "2026-07-22 05:34:44.526716+00:00", "lang": "en", "topics": ["large-language-models", "artificial-intelligence", "natural-language-processing", "ai-research"], "entities": ["RLAES", "ASAP benchmark", "Rubric-based Feedback Evaluation", "Adaptive Gated Feedback Optimization", "Adjacent Contrastive Reasoning", "GPT-5.5"], "alternates": {"html": "https://wpnews.pro/news/beyond-score-prediction-llm-based-essay-scoring-and-feedback-generation-via-with", "markdown": "https://wpnews.pro/news/beyond-score-prediction-llm-based-essay-scoring-and-feedback-generation-via-with.md", "text": "https://wpnews.pro/news/beyond-score-prediction-llm-based-essay-scoring-and-feedback-generation-via-with.txt", "jsonld": "https://wpnews.pro/news/beyond-score-prediction-llm-based-essay-scoring-and-feedback-generation-via-with.jsonld"}}