Grok 4.6 and GPT5.6 beat Anthropic for finding security vulnerabilities in PRs
A new benchmark testing AI models on pull request security reviews found that OpenAI's GPT-5.6 Sol and xAI's Grok 4.5 outperform Anthropic's models, with GPT-5.6 Sol achieving 100% recall at $0.70 per…