On-Policy Parameter Update Direction Underlies Generalization in LLM Post-Training A study of on-policy post-training paradigms in large language models finds that the on-policy parameter update direction underlies generalization, arguing prior work treated these update behaviors only as byproducts rather than as optimization principles. The research examines parameter update behavior during on-policy post-training to explain the strong generalization these paradigms achieve. The strong generalization performance of on-policy post-training paradigms has motivated studies of their parameter update behaviors. However, these studies treat the observed behaviors only as byproducts in on-policy training, overlooking their potential to serve as optimization principles for impr