TGRL: Temperature-Grouped Reinforcement Learning for Efficient Exploration in LLMs Researchers propose TGRL (Temperature-Grouped Reinforcement Learning), a method that groups rollouts by temperature to improve exploration efficiency in reinforcement learning with verifiable rewards (RLVR) for large language models. The approach targets a central bottleneck in RLVR, where temperature control and test-time scaling strategies either expand the sample budget at rollout time or leave exploration constrained. Efficient exploration often remains a central bottleneck in reinforcement learning with verifiable rewards RLVR . Although temperature control and test-time scaling strategies can increase rollout diversity of large language models LLMs , they either expand the sample budget at rollout time or lea