VulnBench: Can LLMs find the same security bugs twice?
Snyk's VulnBench benchmark found that large language models (LLMs) are not fully repeatable in security reviews: across 300 scans of 10 JavaScript projects with 6 configurations repeated 5 times, 84.8…