04:00
2026-09-18
arxiv.org
large-language-models
Reflective Recovery: A Self-Supervised Method for Reasoning by Learning from Mistakes
A self-supervised method called Reflective Recovery, which turns failed reasoning trajectories into recovery training data, raised accuracy on DeepSeek-R1-Distill-Qwen-7B from 30.0% to 37.5% on AIME 2…