Jev: The Model That Doesn’t Write Answers
TypeSafe shipped a model called Jev on 15 September 2026 that returns typed answers with probabilities instead of generating text, a category TypeSafe calls a System One model. Jev accepts a state and…
TypeSafe shipped a model called Jev on 15 September 2026 that returns typed answers with probabilities instead of generating text, a category TypeSafe calls a System One model. Jev accepts a state and…
Former OpenAI employee Diogo Almeida announced Jev on September 15, a probabilistic classifier built by his startup TypeSafe AI that returns calibrated decisions in 70–500 ms at $0.042 per million inp…
Rijul, a developer building the AI code review tool LiveReview, published an explainer comparing two methods for aligning language models with human preferences: Direct Preference Optimization (DPO) a…
TypeSafe AI introduced a training approach it calls RLCD (reinforcement learning from calibrated decisions), which it says makes models return decisions and probabilities rather than generated text, w…
OpenAI's InstructGPT/RLHF work showed that GPT-2-sized models (over 100x smaller than GPT-3) trained on the right task beat GPT-3, according to an essay by the author of The Bitterest Lesson. The essa…
Former OpenAI researcher Diogo Almeida, who describes himself as a co-inventor of RLHF and ChatGPT, has launched TypeSafe AI and its Jev model, which abandons text generation entirely in favor of retu…
Typesafe's Diogo Almeida is publicly arguing that RLHF has structural flaws and is promoting RLCD (reinforcement learning for calibrated decisions), a training method that rewards outcome accuracy and…
TypeSafe AI announced Jev, a new "System One Model" built on a new architecture, sampler, and training algorithm it calls Reinforcement Learning for Calibrated Decisions (RLCD), which returns typed de…
A technical textbook titled "Foundation Model Engineering: From Theory to Production" is being released for AI engineers and research-oriented readers, covering architectures, training pipelines, infe…
OpenAI posted an on-site San Francisco job listing for a Researcher, Safety Training, National Security role paying $380,000 to $500,000 per year, a range the listing says sits 87% above the $236,000 …
DeepSeek-R1's use of Reinforcement Learning with Verifiable Rewards (RLVR) marks a shift from RLHF's subjective human feedback to deterministic correctness checks, enabling models to reason more relia…
Reinforcement learning for language models has moved from Reinforcement Learning from Human Feedback (RLHF) through LLM-as-a-Judge grading to Reinforcement Learning with Verifiable Rewards (RLVR), whe…
Post-training methods such as supervised fine-tuning (SFT), direct preference optimization (DPO), and reinforcement learning (RL) shape a model's behavior after pre-training, with SFT remaining the mo…
A new method called COPA reduces attack success rates by 6.3× versus the best static baseline and 4.4× on average across lifelong attack streams, while retaining 92% defense on month-old attacks compa…
Researchers propose inference-time mitigation strategies using Chain of Thought prompting and Direct Preference Optimization to shield large language models from adversarial political bias, raising Po…
A new analysis argues that reinforcement learning from human feedback (RLHF) is insufficient to ensure the safety of autonomous AI agents, advocating instead for runtime contracts that enforce hard bo…
RLAIF replaces human preference labeling with a model that chooses between two responses, keeping the downstream RLHF pipeline unchanged. The method, exemplified by Constitutional AI, offers scalabili…
Researchers from an unnamed institution introduced Preference-Aware RLHF (PA-RLHF), a method that separates optimization across preference modes during reward learning, to address procedural fairness …
Researchers propose MCF-CVA, a multilayer combinatorial fusion framework for contextual value alignment in large language models, which instantiates multiple moral agents and combines their outputs ac…
Ex-OpenAI AI researcher and co-author of the original RLHF paper argues that Chinese labs cannot effectively distill US frontier reasoning models because these models do not return full reasoning trac…