# Paper Summary: The Matthew Effect in RL

> Source: <https://blog.lukesalamone.com/posts/matthew-effect/>
> Published: 2026-09-26 09:36:55+00:00

[Learning to Solve Hard Problems in RL for LLMs by Never Giving Up](https://arxiv.org/pdf/2609.13443v1) discusses an approach for countering what the authors call the *Matthew Effect*, the tendency for the LLM to improve much more on easy problems that it is already good than harder problems that have a lower solve rate. Their approach is to allocate training time dynamically based on the difficulty of the problem.

After an LLM has been pretrained on a large corpus of supervised fine-tuning (SFT) data, it is common to post-train [using reinforcement learning methods like GRPO](./posts/notes-on-deepseek-r1/). This allows the model to improve on tasks with verifiable rewards and even exceed the performance of the original SFT data.
