cd/entity/RLVR· home› entities› RLVR
grep -l @rlvr /news/*.json | wc -l → 26

RLVR

mentions 26 type Organization page 1/2 feed RSS

// recent coverage 26 mentions

23:31
2026-09-23
pub.towardsai.net
artificial-intelligence

Jev: The Model That Doesn’t Write Answers

TypeSafe shipped a model called Jev on 15 September 2026 that returns typed answers with probabilities instead of generating text, a category TypeSafe calls a System One model. Jev accepts a state and…

10:40
2026-09-18
docs.typesafe.ai
artificial-intelligence

TypeSafe AI's "Meaningful Intelligence"

TypeSafe AI introduced a training approach it calls RLCD (reinforcement learning from calibrated decisions), which it says makes models return decisions and probabilities rather than generated text, w…

12:30
2026-09-14
aiflash.com
machine-learning

DataFlex-RL: An Evaluation Platform for RLVR Data Policies

Researchers introduced DataFlex-RL, an evaluation platform for comparing data policies in reinforcement learning with verifiable rewards (RLVR) under a common GRPO recipe. The platform targets how RLV…

20:06
2026-08-15
lesswrong.com
artificial-intelligence

What if Parameter Updates were Text?

A new fine-tuning method called 'Advice String Distillation' is proposed as a safer alternative to RLVR for training AI models, using context distillation to update weights with text-associated change…

04:00
2026-08-05
arxiv.org
artificial-intelligence

BODHI: Do LLMs Branch Out and Discover Heterogeneous Inferences?

A new arXiv preprint (2608.02867v1) from researchers studying reinforcement learning with verifiable rewards (RLVR) finds that RLVR-trained large language models (LLMs) exhibit reduced semantic branch…

04:00
2026-08-03
machinebrief.com
machine-learning

SAF-OPD: Stable Advantage Fusion for On-Policy Distillation

Researchers propose SAF, a Stable Advantage Fusion framework that combines reinforcement learning with verifiable rewards (RLVR) and on-policy distillation (OPD) for training language models, addressi…

04:00
2026-07-07
arxiv.org
machine-learning

Reinforcement Learning for Data-Efficient Code-Switched ASR

Researchers propose a reinforcement learning with verifiable rewards (RLVR) method for data-efficient adaptation of audio-language models to code-switched automatic speech recognition (ASR). Using Qwe…

page 1 / 2 next →
// co-occurs with top 8 entities
// topics top 6 topics