{"slug": "openjev-rlcd-a-working-rlcd-implementation", "title": "OpenJev-RLCD: A Working RLCD Implementation", "summary": "A working implementation of reinforcement learning for calibrated decisions (RLCD) for reasoning models, OpenJev-RLCD, matches or beats supervised fine-tuning, RFT/STaR and GRPO in accuracy and beats all of them in selective prediction, according to an arXiv paper (2609.38850v1) whose code is posted at github.com/ZimmyGao/openjev-rlcd. With Qwen3-1.7B across two reasoning tasks and 3 seeds with paired tests, a single query on GSM8K answer verification decides 25% of items at 5% error or less, versus 5% for GRPO. The paper's two-stage recipe — calibrate, then reinforce — scores the answer distribution a model commits to after sampling a rationale with a strictly proper scoring rule, and shows RLVR is that mixture objective without its diversity term.", "body_md": "arXiv:2609.38850v1 Announce Type: new \nAbstract: Decision models such as Jev answer questions with probabilities, which are only useful if they are calibrated. Open-source reproductions rely on supervised fine-tuning plus temperature scaling, while reinforcement learning from verifiable rewards (RLVR) makes reasoning models overconfident. We present a working implementation of reinforcement learning for calibrated decisions (RLCD) for reasoning models: the model samples a rationale, and we score the answer distribution it commits to afterwards with a strictly proper scoring rule. A variance identity shows that scoring the mixture of several samples rewards disagreeing rationales, and that RLVR is exactly this mixture objective without its diversity term. Optimized naively, the per-rationale objective either switches reasoning off or is drowned out by policy-gradient noise, which leads to a two-stage recipe: calibrate, then reinforce. With Qwen3-1.7B on two reasoning tasks (3 seeds, paired tests), RLCD matches or beats SFT, RFT/STaR and GRPO (each temperature-scaled) in accuracy and beats all of them in selective prediction; on GSM8K answer verification a single query decides \\gvTwoCovFive\\% of the items at $\\le$5\\% error, versus \\gvGrpoCovFive\\% for GRPO. When uncertainty comes from annotator disagreement, RLCD provably cannot beat cross-entropy. Code and results: https://github.com/ZimmyGao/openjev-rlcd.", "url": "https://wpnews.pro/news/openjev-rlcd-a-working-rlcd-implementation", "canonical_source": "https://www.machinebrief.com/news/openjev-rlcd-a-working-rlcd-implementation-cjsf", "published_at": "2026-10-01 04:00:00+00:00", "updated_at": "2026-10-01 04:47:48.511557+00:00", "lang": "en", "topics": ["machine-learning", "ai-research", "large-language-models", "ai-safety"], "entities": ["OpenJev-RLCD", "Qwen3-1.7B", "GRPO", "GSM8K", "arXiv", "ZimmyGao"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/openjev-rlcd-a-working-rlcd-implementation", "markdown": "https://wpnews.pro/news/openjev-rlcd-a-working-rlcd-implementation.md", "text": "https://wpnews.pro/news/openjev-rlcd-a-working-rlcd-implementation.txt", "jsonld": "https://wpnews.pro/news/openjev-rlcd-a-working-rlcd-implementation.jsonld"}}