cd/entity/GRPO· home entities GRPO
grep -l @grpo /news/*.json | wc -l → 87

GRPO

mentions 87 type Organization page 4/5 feed RSS

// recent coverage 87 mentions

04:00
2026-06-29
arxiv.org
large-language-models

Tandem Reinforcement Learning with Verifiable Rewards

Researchers propose Tandem Reinforcement Learning (TRL), extending the tandem training paradigm to reinforcement learning with verifiable rewards (RLVR). Training Qwen3-4B-Instruct on competition math…

06:04
2026-06-23
devclubhouse.com
large-language-models

When a 3B Model Out-Reasons Opus 4.5, Read the Fine Print

Weibo's AI group released VibeThinker-3B, a 3-billion-parameter model that achieves 94.3 on AIME and outperforms Claude Opus 4.5 on competition reasoning, but collapses on general knowledge tasks. The…

13:39
2026-06-20
byungkwanlee.github.io
machine-learning

Nvidia-ZPPO: Zone of Proximal Policy Optimization

Nvidia researchers introduced Zone of Proximal Policy Optimization (ZPPO), a method that uses a replay buffer to repeatedly expose student models to hard questions, improving rollout accuracy without …

21:45
2026-06-19
dwarkesh.com
artificial-intelligence

The data black hole at the center of AI

AI progress is driven primarily by massive amounts of domain-specific human expert data and reinforcement learning, not by architectural breakthroughs, creating a data black hole that powers model cap…

19:07
2026-06-19
dev.to
artificial-intelligence

Self-Evolving AI Agents: The Optimizer Is the Easy Part

Decagon deployed self-evolving AI agents that automatically improve their own prompts without human intervention, using a feedback loop that tests and promotes winning versions. The team found that th…

18:30
2026-06-17
castform.com
large-language-models

I post-trained a model to reliably roll a die

A developer post-trained a language model to roll a die, revealing that models default to outputting 4 due to training data bias. Standard reinforcement learning fails to encourage exploration, but ad…

00:00
2026-06-13
research.rudrite.com
artificial-intelligence

Comparisons — AI & ML approaches side by side | Rudrite Research

Rudrite Research published a comprehensive comparison of AI and ML approaches, covering 14 side-by-side analyses of techniques such as Transformers vs Mamba, FlashAttention vs PagedAttention, and PPO …

04:00
2026-06-05
arxiv.org
large-language-models

Self-supervised User Profile Generation for Personalization

Researchers have developed BUMP, a self-supervised framework that trains large language models to generate personalized user profiles without requiring labeled data from downstream tasks. The system u…

04:00
2026-06-04
arxiv.org
large-language-models

POLARIS: Guiding Small Models to Write Long Stories

Researchers have developed POLARIS-9B, a 9-billion-parameter language model that generates long-form stories with improved quality and length adherence, using a novel training recipe that combines a f…

21:32
2026-06-02
github.com
machine-learning

FeynRL- Don't let systems swallow the algorithm

FeynRL, an algorithm-first framework for post-training and fine-tuning large models, has been released as an open-source tool supporting supervised fine-tuning, preference learning, and reinforcement …

← prev page 4 / 5 next →
// co-occurs with top 8 entities
// topics top 6 topics