# TGRL: Temperature-Grouped Reinforcement Learning for Efficient Exploration in LLMs

> Source: <https://aiflash.com/news/128960/>
> Published: 2026-09-30 02:00:12+00:00

Efficient exploration often remains a central bottleneck in reinforcement learning with verifiable rewards (RLVR). Although temperature control and test-time scaling strategies can increase rollout diversity of large language models (LLMs), they either expand the sample budget at rollout time or lea
