{"slug": "rltl-dr-self-improvement-by-internalizing-self-generated-feedback", "title": "RLTL;DR: Self-Improvement by Internalizing Self-Generated Feedback", "summary": "A new method called RLTL;DR lets a Qwen 3.5 9B Thinking policy break through a learning barrier on tool-calling and coding datasets filtered to Pass@128 = 0, reaching a Pass@1 of 14–31% with insights in context during training and 12–13% with no insight in context at evaluation time, according to the paper. Standard GRPO training of the same policy stays flat at a Pass@1 of 0% to 1%. The approach has the policy write a single TL;DR insight from the verifier output after each failed attempt, conditions the next rollout on all previous insights, and backpropagates on the in-context insights to internalize a direct task → insight mapping; a reduced variant, SFTL;DR, trained only on (task, insight) tuples, recovers almost the full performance of RLTL;DR and classical SFT on full rollouts using just 4k tuples.", "body_md": "The common paradigm of reinforcement learning with verifiable rewards (RLVR) is to let agents make multiple attempts at a task, and optimize towards the successful ones. This becomes problematic in the realms of self-improvement, where tasks are so difficult that the agent has a low or even no chance of success, and where there are no teacher models or example solutions to distill from. In this paper, we introduce RLTL;DR. After each failed attempt, we show the policy the verifier outputs and let it write its own feedback, in the form of a single TL;DR insight. The next rollout is conditioned on all previous insights, and we sequentially sample rollouts until a solution is found. Moreover, we enable backpropagation on the in-context insights to internalize a direct task → insight mapping. On challenging tool-calling and coding datasets (filtered to Pass@128 = 0), standard GRPO training of a Qwen 3.5 9B Thinking policy stays flat at a Pass@1 of 0% to 1%. RLTL;DR breaks through this learning barrier, achieving a Pass@1 of 14–31% with insights in context during training and, crucially, 12–13% when no insight is in context at eval time. We identify that the key is the task → insight internalization. To study this further, we reduce our approach to SFTL;DR, training only on (task, insight) tuples, without showing or backpropagating on any rollouts. Training on only 4k of these tuples recovers almost the full performance of RLTL;DR and classical SFT on full rollouts. This demonstrates a promising compacted training paradigm of the form “on this sort of task, keep this sort of thing in mind”, which we hope to inspire future research on.", "url": "https://wpnews.pro/news/rltl-dr-self-improvement-by-internalizing-self-generated-feedback", "canonical_source": "https://machinelearning.apple.com/research/rltl-dr-self-improvement", "published_at": "2026-10-01 00:00:00+00:00", "updated_at": "2026-10-01 15:14:39.511379+00:00", "lang": "en", "topics": ["machine-learning", "ai-research", "large-language-models", "ai-agents"], "entities": ["RLTL;DR", "SFTL;DR", "GRPO", "Qwen 3.5 9B Thinking"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/rltl-dr-self-improvement-by-internalizing-self-generated-feedback", "markdown": "https://wpnews.pro/news/rltl-dr-self-improvement-by-internalizing-self-generated-feedback.md", "text": "https://wpnews.pro/news/rltl-dr-self-improvement-by-internalizing-self-generated-feedback.txt", "jsonld": "https://wpnews.pro/news/rltl-dr-self-improvement-by-internalizing-self-generated-feedback.jsonld"}}