{"slug": "tgrl-temperature-grouped-reinforcement-learning-for-efficient-exploration-in", "title": "TGRL: Temperature-Grouped Reinforcement Learning for Efficient Exploration in LLMs", "summary": "Researchers propose TGRL (Temperature-Grouped Reinforcement Learning), a method that groups rollouts by temperature to improve exploration efficiency in reinforcement learning with verifiable rewards (RLVR) for large language models. The approach targets a central bottleneck in RLVR, where temperature control and test-time scaling strategies either expand the sample budget at rollout time or leave exploration constrained.", "body_md": "Efficient exploration often remains a central bottleneck in reinforcement learning with verifiable rewards (RLVR). Although temperature control and test-time scaling strategies can increase rollout diversity of large language models (LLMs), they either expand the sample budget at rollout time or lea", "url": "https://wpnews.pro/news/tgrl-temperature-grouped-reinforcement-learning-for-efficient-exploration-in", "canonical_source": "https://aiflash.com/news/128960/", "published_at": "2026-09-30 02:00:12+00:00", "updated_at": "2026-09-30 02:17:54.966300+00:00", "lang": "en", "topics": ["machine-learning", "large-language-models", "ai-research", "artificial-intelligence"], "entities": ["TGRL", "Temperature-Grouped Reinforcement Learning", "RLVR"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/tgrl-temperature-grouped-reinforcement-learning-for-efficient-exploration-in", "markdown": "https://wpnews.pro/news/tgrl-temperature-grouped-reinforcement-learning-for-efficient-exploration-in.md", "text": "https://wpnews.pro/news/tgrl-temperature-grouped-reinforcement-learning-for-efficient-exploration-in.txt", "jsonld": "https://wpnews.pro/news/tgrl-temperature-grouped-reinforcement-learning-for-efficient-exploration-in.jsonld"}}