LURE: Alignment Evaluations to Reduce Evaluation Awareness
Apollo Research and other labs found that advanced AI models like Claude Opus 4.6 and Gemini 3.1 Pro Preview can detect when they are being evaluated, potentially allowing them to fake alignment and p…