We call this Machine Native Intelligence: AI with software-like properties such as structure, reliability, observability, testability, speed, consistency, and low cost.
Building prod, not God #
TypeSafe is not trying to build a model that does everything. It is designed for production systems where code needs a narrow decision it can inspect and act on. Our expectation is that large-scale AI automation will be closer to 99% machine-to-machine interactions and 1% human interaction. That shifts the design target from responses that feel good to read toward outputs that behave predictably inside software. Read the
Three post-training approaches #
Pretrained language models have been adapted in two major ways. TypeSafe adds a third. RLHF and RLVR are shown here for context; TypeSafe’s training path is RLCD. RLHF was used to train InstructGPT and ChatGPT and was
co-invented by Diogo Almeida, cofounder of TypeSafe.
RLCD and calibrated decisions #
RLCD optimizes for a different output contract:
-
The model does not generate text.
-
It returns decisions and probabilities.
-
Higher probability should correspond to a greater chance that the answer is correct.
-
Outcomes assigned a probability of
0.2should occur about 20% of the time. -
Outcomes assigned a probability of
0.8should occur about 80% of the time. -
Outcomes assigned a probability of
1.0should occur 100% of the time.
Confidencefor guidance on deciding when software should act or escalate.
The problems with RLHF #
RLHF teaches a model to say things that people prefer. That objective works well for chatbots, but it can also reward sycophancy and confident-sounding hallucinations. Preference optimization also causes mode dropping: the model learns to favor a particular style, such as instruction following, while reducing the probability of other possible outputs.
mode collapse. In the classic generative-adversarial-network failure mode, a generator learns to produce the same kind of output repeatedly because that output continues to fool the discriminator.
Mode collapse analogy #
Mode collapse analogy