cd/entity/GRPO· home entities GRPO
grep -l @grpo /news/*.json | wc -l → 87

GRPO

mentions 87 type Organization page 5/5 feed RSS

// recent coverage 87 mentions

04:00
2026-05-28
arxiv.org
artificial-intelligence

Cross-Entropy Games and Frost Training

Researchers introduced Frost Training, a method that improves Monte Carlo-based policy optimization for Cross-Entropy Games by exploiting the gradient of the reward function in embedding space. The te…

18:22
2026-05-16
research.nvidia.com
large-language-models

iGRPO: Self-Feedback-Driven LLM Reasoning

Researchers introduced Iterative Group Relative Policy Optimization (iGRPO), a two-stage reinforcement learning method that improves large language model reasoning by having the model generate and ref…

19:06
2026-05-06
huggingface.co
large-language-models

vLLM V0 to V1: Correctness Before Corrections in RL

Here is a 2-3 sentence factual summary of the article: The article describes the process of migrating an online reinforcement learning (RL) training system from the vLLM V0 engine to the V1 rewrite, …

15:01
2026-04-29
huggingface.co
large-language-models

Granite 4.1 LLMs: How They’re Built

The Granite 4.1 family consists of dense, decoder-only LLMs (3B, 8B, and 30B parameters) trained from scratch on approximately 15 trillion tokens through a five-phase pre-training pipeline that progre…

← prev page 5 / 5
// co-occurs with top 8 entities
// topics top 6 topics