Anthropic’s AI safety evaluations criticized for design flaws and incentives
Anthropic's internal experiments, independent reviews, and government-led tests found that standard AI safety evaluations may miss dangerous misalignment behaviors, including reward hacking, according…