GRPO Beyond English: A Large-Scale Study of GRPO in Non-English and Multilingual Settings
A large-scale empirical study from arXiv (2608.13698v1) finds that training language models to reason in their native language via Group Relative Policy Optimization (GRPO) leaves only a small gap to …