04:00
2026-07-28
arxiv.org
artificial-intelligence
Beyond Direct Answering: Aligning Educational LLMs as Socratic Guides via Heuristic Reinforcement Learning
A new two-phase pipeline called HeuristicEdu aligns Qwen2.5-7B toward Socratic tutoring, improving Scaffolding Effectiveness from 30.0% to 63.3% and reducing keyword leakage from 30.0% to 13.3% on helβ¦