cd/entity/RLVR· home› entities› RLVR
grep -l @rlvr /news/*.json | wc -l → 26

RLVR

mentions 26 type Organization page 2/2 feed RSS

// recent coverage 26 mentions

04:00
2026-06-29
arxiv.org
large-language-models

Tandem Reinforcement Learning with Verifiable Rewards

Researchers propose Tandem Reinforcement Learning (TRL), extending the tandem training paradigm to reinforcement learning with verifiable rewards (RLVR). Training Qwen3-4B-Instruct on competition math…

19:40
2026-06-27
lesswrong.com
large-language-models

Neuralese is Actually Probably Good for Alignment

Reinforcement Learning with Verifiable Rewards (RLVR) allows language models to bootstrap beyond human-level capabilities on exactly graded problems like coding and formal proofs, but alignment-flavor…

04:00
2026-06-04
arxiv.org
machine-learning

Self-Distilled Policy Gradient

Researchers introduced SDPG, a self-distilled policy-gradient framework that combines group-relative verifier advantages with normalized standard deviation and full-vocabulary on-policy self-distillat…

21:11
2026-05-20
vmax.ai
artificial-intelligence

PopuLoRA: Co-Evolving LLM Populations for Reasoning Self- Play

PopuLoRA is a method for training large language models (LLMs) that uses co-evolving populations of teacher and student adapters to generate and solve verifiable reasoning tasks, such as code and math…

← prev page 2 / 2
// co-occurs with top 8 entities
// topics top 6 topics