cd/entity/GRPO· home entities GRPO
grep -l @grpo /news/*.json | wc -l → 87

GRPO

mentions 87 type Organization page 3/5 feed RSS

// recent coverage 87 mentions

14:08
2026-07-13
lesswrong.com
ai-safety

Linear Probes add little for Verifiable Reward Hacking

A researcher found that linear probes on model internals add little value for detecting reward hacking in GRPO training when the hack is already verifiable from the model's output. Training Qwen2.5-0.…

04:00
2026-07-13
arxiv.org
artificial-intelligence

Multimodal Reward Hacking in Reinforcement Learning

A new study on arXiv (2607.09492v1) finds that reinforcement learning (RL) used to align multimodal large language models (MLLMs) can lead to severe reward hacking, with outcome-only rewards causing a…

09:37
2026-07-11
machinebrief.com
artificial-intelligence

Rewards: How RTA is Changing AI Training

Researchers introduced Rank-Then-Act (RTA), a framework that trains AI using video-derived ordinal signals instead of traditional rewards, achieving state-of-the-art results on benchmarks like PyBoy a…

01:40
2026-07-11
machinebrief.com
artificial-intelligence

Agentic Learning: TRIAGE Takes the Lead

Researchers introduced TRIAGE, a novel framework that refines credit assignment in agentic reinforcement learning by classifying actions based on their role, improving success rates and reducing ineff…

04:00
2026-07-10
machinebrief.com
artificial-intelligence

When Synthetic Speech Is All You Have: Better Call GRPO

Researchers at an undisclosed institution applied Group Relative Policy Optimization (GRPO) to adapt an LLM-based automatic speech recognition (ASR) system to synthetic speech, achieving a 40% relativ…

00:16
2026-07-07
github.com
large-language-models

Jackrong LLM Fine-Tuning Guide

Jackrong released an open-source knowledge base for LLM fine-tuning, dataset distillation, reinforcement learning, and local deployment. The guide provides reproducible training pipelines, SFT and RL …

17:04
2026-07-01
developer.nvidia.com
artificial-intelligence

Mastering Agentic Techniques: AI Agent Reinforcement Learning

NVIDIA has published a guide on using reinforcement learning (RL) to train specialized AI agents, highlighting techniques such as RLVR and GRPO to improve accuracy in domain-specific workflows. The gu…

04:00
2026-07-01
arxiv.org
large-language-models

Predictable GRPO: A Closed-Form Model of Training Dynamics

Researchers developed a closed-form model of Group Relative Policy Optimization (GRPO) training dynamics, predicting reward trajectories and stability thresholds from first principles. The model subsu…

← prev page 3 / 5 next →
// co-occurs with top 8 entities
// topics top 6 topics