We benchmarked A LOT of models, here’s how they compare to Mythos
Claude Mythos, Anthropic's frontier model, achieved 80.0% precision and 13.9% recall (F1 ~23.7%) on a benchmark of 275 human-reviewed IDOR vulnerabilities across four codebases, ranking 15th out of 17…