TypeSafe AI's "Meaningful Intelligence" TypeSafe AI introduced a training approach it calls RLCD (reinforcement learning from calibrated decisions), which it says makes models return decisions and probabilities rather than generated text, with outcomes assigned a probability of 0.2 occurring about 20% of the time and outcomes assigned 0.8 occurring about 80% of the time. The company, whose cofounder Diogo Almeida co-invented RLHF used to train InstructGPT and ChatGPT, frames the method as a third post-training path alongside RLHF and RLVR and targets production systems it expects to run about 99% machine-to-machine interactions and 1% human interaction. TypeSafe argues RLHF can reward sycophancy and confident-sounding hallucinations and causes mode dropping, and it published a manifesto describing the approach as Machine Native Intelligence. We call this Machine Native Intelligence: AI with software-like properties such as structure, reliability, observability, testability, speed, consistency, and low cost. Building prod, not God TypeSafe is not trying to build a model that does everything. It is designed for production systems where code needs a narrow decision it can inspect and act on. Our expectation is that large-scale AI automation will be closer to 99% machine-to-machine interactions and 1% human interaction. That shifts the design target from responses that feel good to read toward outputs that behave predictably inside software. Read the TypeSafe manifesto https://typesafe.ai/manifesto . Three post-training approaches Pretrained language models have been adapted in two major ways. TypeSafe adds a third. RLHF and RLVR are shown here for context; TypeSafe’s training path is RLCD. RLHF was used to train InstructGPT and ChatGPT and was co-invented by Diogo Almeida https://scholar.google.com/citations?user=0T4y07QAAAAJ&hl=en , cofounder of TypeSafe. RLCD and calibrated decisions RLCD optimizes for a different output contract: - The model does not generate text. - It returns decisions and probabilities. - Higher probability should correspond to a greater chance that the answer is correct. - Outcomes assigned a probability of 0.2 should occur about 20% of the time. - Outcomes assigned a probability of 0.8 should occur about 80% of the time. - Outcomes assigned a probability of 1.0 should occur 100% of the time. Confidence https://docs.typesafe.ai/confidence for guidance on deciding when software should act or escalate. The problems with RLHF RLHF teaches a model to say things that people prefer. That objective works well for chatbots, but it can also reward sycophancy and confident-sounding hallucinations. Preference optimization also causes mode dropping : the model learns to favor a particular style, such as instruction following, while reducing the probability of other possible outputs. mode collapse . In the classic generative-adversarial-network failure mode, a generator learns to produce the same kind of output repeatedly because that output continues to fool the discriminator. Mode collapse analogy Mode collapse analogy