Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds
Researchers propose two techniques to make alignment evaluations harder for AI models to distinguish from real deployments, addressing 'evaluation awareness.' Critique refinement spends additional inf…