A benchmark for safely measuring container breakout capabilities
The UK AI Safety Institute (AISI) released SandboxEscapeBench, an open-source benchmark with 18 scenarios to test whether AI agents can break out of container sandboxes, finding that advanced models r…