# On-Policy Parameter Update Direction Underlies Generalization in LLM Post-Training

> Source: <https://aiflash.com/news/131123/>
> Published: 2026-10-05 03:30:16+00:00

The strong generalization performance of on-policy post-training paradigms has motivated studies of their parameter update behaviors. However, these studies treat the observed behaviors only as byproducts in on-policy training, overlooking their potential to serve as optimization principles for impr
