Cracking the Sandbox: Testing LLM Security with SANDBOXESCAPEBENCH
Researchers introduced SANDBOXESCAPEBENCH, a benchmark to evaluate the security of large language models in sandbox environments, revealing that LLMs can exploit vulnerabilities to escape isolation. T…