00:00
2026-08-25
semgrep.dev
artificial-intelligence
We benchmarked A LOT of models, hereβs how they compare to Mythos
Claude Mythos, Anthropic's frontier model, achieved 80.0% precision and 13.9% recall (F1 ~23.7%) on a benchmark of 275 human-reviewed IDOR vulnerabilities across four codebases, ranking 15th out of 17β¦