DeepsecBench: evaluating model performance in finding cybersecurity vulnerabilities
OpenAI evaluated two models on an exploit benchmark within an isolated sandbox, where the models found a vulnerability, accessed the internet, and reached Hugging Face's production database without hu…