cd/entity/GRPO· home entities GRPO
grep -l @grpo /news/*.json | wc -l → 87

GRPO

mentions 87 type Organization page 1/5 feed RSS

// recent coverage 87 mentions

04:00
2026-08-03
machinebrief.com
machine-learning

SAF-OPD: Stable Advantage Fusion for On-Policy Distillation

Researchers propose SAF, a Stable Advantage Fusion framework that combines reinforcement learning with verifiable rewards (RLVR) and on-policy distillation (OPD) for training language models, addressi…

16:10
2026-08-01
lesswrong.com
artificial-intelligence

Do your capabilities homework

A technical AI safety researcher argues that safety-focused researchers should engage with capabilities research, highlighting On-Policy Self-Distillation (OPSD) as a promising alternative to GRPO for…

10:22
2026-07-30
pub.towardsai.net
large-language-models

What is Parameter Lower Bound in Efficient LLM Adaptation

A practical analysis of full fine-tuning, LoRA, QLoRA, and TinyLoRA shows that TinyLoRA improved mathematical reasoning in a frozen Qwen2.5-7B-Instruct model with only 13 trainable parameters under GR…

12:13
2026-07-28
byteiota.com
artificial-intelligence

$500 RL Fine-Tune Beats Claude Opus 4.6 on Real Task

Ramp and Prime Intellect published a case study showing a small RL-trained model, FastAsk, outperformed Claude Opus 4.6 by 4 percentage points on exact-match accuracy for financial spreadsheet retriev…

06:02
2026-07-25
promptcube3.com
machine-learning

REINFORCE vs DQN: Learning Policies Directly

A technical comparison of REINFORCE and DQN reinforcement learning algorithms shows that REINFORCE learns policies directly by outputting action probabilities and sampling, eliminating the need for re…

page 1 / 5 next →
// co-occurs with top 8 entities
// topics top 6 topics