cd/entity/RLVR· home entities RLVR
grep -l @rlvr /news/*.json | wc -l → 15

RLVR

mentions 15 type Organization feed RSS

// recent coverage 15 mentions

20:06
2026-08-15
lesswrong.com
artificial-intelligence

What if Parameter Updates were Text?

A new fine-tuning method called 'Advice String Distillation' is proposed as a safer alternative to RLVR for training AI models, using context distillation to update weights with text-associated change…

04:00
2026-08-05
arxiv.org
artificial-intelligence

BODHI: Do LLMs Branch Out and Discover Heterogeneous Inferences?

A new arXiv preprint (2608.02867v1) from researchers studying reinforcement learning with verifiable rewards (RLVR) finds that RLVR-trained large language models (LLMs) exhibit reduced semantic branch…

12:01
2026-08-03
dev.to
artificial-intelligence

Top AI Papers on Hugging Face - 2026-08-03

A developer highlights the top 10 papers on Hugging Face, covering GUI agents, LLM self-improvement via RL, parameterized long-term memory, RAG, multimodal robots, 3D generation, and role-play benchma…

04:00
2026-08-03
machinebrief.com
machine-learning

SAF-OPD: Stable Advantage Fusion for On-Policy Distillation

Researchers propose SAF, a Stable Advantage Fusion framework that combines reinforcement learning with verifiable rewards (RLVR) and on-policy distillation (OPD) for training language models, addressi…

04:00
2026-07-07
arxiv.org
machine-learning

Reinforcement Learning for Data-Efficient Code-Switched ASR

Researchers propose a reinforcement learning with verifiable rewards (RLVR) method for data-efficient adaptation of audio-language models to code-switched automatic speech recognition (ASR). Using Qwe…

04:00
2026-06-29
arxiv.org
large-language-models

Tandem Reinforcement Learning with Verifiable Rewards

Researchers propose Tandem Reinforcement Learning (TRL), extending the tandem training paradigm to reinforcement learning with verifiable rewards (RLVR). Training Qwen3-4B-Instruct on competition math…

19:40
2026-06-27
lesswrong.com
large-language-models

Neuralese is Actually Probably Good for Alignment

Reinforcement Learning with Verifiable Rewards (RLVR) allows language models to bootstrap beyond human-level capabilities on exactly graded problems like coding and formal proofs, but alignment-flavor…

04:00
2026-06-04
arxiv.org
machine-learning

Self-Distilled Policy Gradient

Researchers introduced SDPG, a self-distilled policy-gradient framework that combines group-relative verifier advantages with normalized standard deviation and full-vocabulary on-policy self-distillat…

21:11
2026-05-20
vmax.ai
artificial-intelligence

PopuLoRA: Co-Evolving LLM Populations for Reasoning Self- Play

PopuLoRA is a method for training large language models (LLMs) that uses co-evolving populations of teacher and student adapters to generate and solve verifiable reasoning tasks, such as code and math…

// co-occurs with top 8 entities
// topics top 6 topics