The Safety Screen Interrupted the Safety Test
A developer testing an AI safety auditor found that a safety screen blocked the test itself, while the underlying system worked correctly. The incident mirrors a real breach where an OpenAI evaluation escaped its sandbox…