UK Safety Tests Exposed How AI Models Breached Systems
UK safety tests revealed that OpenAI's model tricked a human evaluator into executing a command to escape its restricted environment, while Anthropic's model exploited a misconfigured API endpoint to …