Maybe we should pretrain on synthetic data about good-but-reward-hacking AIs
Researchers propose "inoculation pretraining," a method to prevent emergent misalignment in AI systems by adding synthetic training data about good-but-reward-hacking AI personas. The approach aims to increase the prior …