cd/entity/GRPO· home entities GRPO
grep -l @grpo /news/*.json | wc -l → 87

GRPO

mentions 87 type Organization page 2/5 feed RSS

// recent coverage 87 mentions

15:12
2026-07-21
developers.googleblog.com
artificial-intelligence

Scaling Agentic RL: High-Throughput Agentic Training with Tunix

Google's Tunix post-training library introduces a high-throughput framework for agentic reinforcement learning that keeps TPUs fully utilized during multi-turn training. Tunix uses an asynchronous tra…

04:00
2026-07-21
arxiv.org
artificial-intelligence

Group Entropy-Controlled Policy Optimization

Researchers propose Group Entropy-Controlled Policy Optimization (GEPO), a lightweight extension to GRPO that uses group entropy to perform entropy-conditioned asymmetric advantage shaping, addressing…

04:00
2026-07-20
arxiv.org
artificial-intelligence

Process Reward Informed Tree Rollout for Effective Multi-Turn RL

Researchers propose Process-Scorer Guided Adaptive Tree Rollout (PATR), a quality-aware rollout framework for multi-turn reinforcement learning (RL) that uses process feedback to selectively branch fr…

07:52
2026-07-16
machinebrief.com
artificial-intelligence

Exploration in AI: Unearthing New Behaviors

New research shows that incentivizing diversity in AI exploration through a representation-based bonus derived from hidden states of pre-trained models boosts verifier efficiency by over 50% for the Q…

23:09
2026-07-14
arxiv.org
large-language-models

LLM-as-a-Verifier: A General-Purpose Verification Framework

Researchers introduced LLM-as-a-Verifier, a general-purpose verification framework that computes continuous scores from scoring token logits to determine solution correctness without additional traini…

16:05
2026-07-14
sourcefeed.dev
artificial-intelligence

Nested RL Agents That Write Real Training Jobs

An open-source pipeline called ai-trains-ai uses nested reinforcement learning loops where an outer agent is RL-trained to write inner training jobs for small models, with the entire system run for ro…

05:52
2026-07-14
machinebrief.com
artificial-intelligence

STAMP's New Approach: Fixing the Reward-Credit Mismatch in AI

Researchers have introduced STAMP (Step-wise Attribution of Modulated Potential), a new reinforcement learning approach that addresses the reward-credit mismatch by linking actions to rewards more dir…

05:39
2026-07-14
machinebrief.com
artificial-intelligence

Self-Verified Reasoner: The AI That Talks to Itself

A new framework called Self-Verified Reasoner (SVR-R1) allows AI models to self-verify their answers by asking themselves 'Yes or No' before finalizing a response, boosting accuracy on vision-language…

← prev page 2 / 5 next →
// co-occurs with top 8 entities
// topics top 6 topics