Measuring Reward-Seeking by Instilling Contrastive Beliefs
OpenAI researchers operationalized reward-seeking in machine learning models as the causal sensitivity of behavior to beliefs about grader preferences, finding that training checkpoints of several fro…