cd/entity/RLHF· home› entities› RLHF
grep -l @rlhf /news/*.json | wc -l → 37

RLHF

mentions 37 type Organization page 1/2 feed RSS

// recent coverage 37 mentions

23:31
2026-09-23
pub.towardsai.net
artificial-intelligence

Jev: The Model That Doesn’t Write Answers

TypeSafe shipped a model called Jev on 15 September 2026 that returns typed answers with probabilities instead of generating text, a category TypeSafe calls a System One model. Jev accepts a state and…

09:30
2026-09-22
machinebrief.com
machine-learning

The most Viral new AI model isn’t an LLM at all

Former OpenAI employee Diogo Almeida announced Jev on September 15, a probabilistic classifier built by his startup TypeSafe AI that returns calibrated decisions in 70–500 ms at $0.042 per million inp…

10:40
2026-09-18
docs.typesafe.ai
artificial-intelligence

TypeSafe AI's "Meaningful Intelligence"

TypeSafe AI introduced a training approach it calls RLCD (reinforcement learning from calibrated decisions), which it says makes models return decisions and probabilities rather than generated text, w…

03:16
2026-09-18
typesafe.ai
artificial-intelligence

The Bitterest Lesson

OpenAI's InstructGPT/RLHF work showed that GPT-2-sized models (over 100x smaller than GPT-3) trained on the right task beat GPT-3, according to an essay by the author of The Bitterest Lesson. The essa…

00:00
2026-09-16
mindstudio.ai
artificial-intelligence

RLCD vs RLHF: What Is Typesafe's Jeff Model Actually Claiming?

Typesafe's Diogo Almeida is publicly arguing that RLHF has structural flaws and is promoting RLCD (reinforcement learning for calibrated decisions), a training method that rewards outcome accuracy and…

21:03
2026-09-15
typesafe.ai
artificial-intelligence

Typesafe AI

TypeSafe AI announced Jev, a new "System One Model" built on a new architecture, sampler, and training algorithm it calls Reinforcement Learning for Calibrated Decisions (RLCD), which returns typed de…

15:27
2026-09-14
sungeuns.github.io
large-language-models

Foundation Model Engineering: From Theory to Production

A technical textbook titled "Foundation Model Engineering: From Theory to Production" is being released for AI engineers and research-oriented readers, covering architectures, training pipelines, infe…

03:21
2026-09-12
frontierroles.com
ai-safety

Researcher, Safety Training, National Security — OpenAI

OpenAI posted an on-site San Francisco job listing for a Researcher, Safety Training, National Security role paying $380,000 to $500,000 per year, a range the listing says sits 87% above the $236,000 …

19:01
2026-08-29
pub.towardsai.net
machine-learning

SFT, RL and DPO: The Other Stack

Post-training methods such as supervised fine-tuning (SFT), direct preference optimization (DPO), and reinforcement learning (RL) shape a model's behavior after pre-training, with SFT remaining the mo…

18:01
2026-08-21
promptcube3.com
artificial-intelligence

COPA treats prompt injection as lifelong learning not one-time

A new method called COPA reduces attack success rates by 6.3× versus the best static baseline and 4.4× on average across lifelong attack streams, while retaining 92% defense on month-old attacks compa…

17:11
2026-08-12
dev.to
artificial-intelligence

RLAIF: The Model as Preference Labeller

RLAIF replaces human preference labeling with a model that chooses between two responses, keeping the downstream RLHF pipeline unchanged. The method, exemplified by Constitutional AI, offers scalabili…

04:00
2026-08-12
arxiv.org
artificial-intelligence

Procedural Fairness Failures in RLHF from Preference Averaging

Researchers from an unnamed institution introduced Preference-Aware RLHF (PA-RLHF), a method that separates optimization across preference modes during reward learning, to address procedural fairness …

04:00
2026-08-11
arxiv.org
artificial-intelligence

Contextual Value Alignment via Multilayer Combinatorial Fusion

Researchers propose MCF-CVA, a multilayer combinatorial fusion framework for contextual value alignment in large language models, which instantiates multiple moral agents and combines their outputs ac…

04:55
2026-07-24
completeskeptic.com
artificial-intelligence

Is it even possible for the Chinese Labs to distill US models?

Ex-OpenAI AI researcher and co-author of the original RLHF paper argues that Chinese labs cannot effectively distill US frontier reasoning models because these models do not return full reasoning trac…

page 1 / 2 next →
// co-occurs with top 8 entities
// topics top 6 topics