Frontier models found the vulnerabilities. Only the attacker found the chains. Evo Continuous Offensive Security (COS) confirmed 10 of 15 registered exploit chains on the deliberately vulnerable web app TaintedPort v1.35, while Claude Security running Mythos confirmed none, according to a head-to-head test published October 7, 2026. Evo COS scored 75.7% severity-weighted detection versus Claude Security's 49.6%, found 50 of 57 known vulnerabilities versus 37, and posted an F1 of 91.7% versus 75.5%, with precision of 96.2% against 90.2%. Evo COS chained an SSRF flaw to retrieve a hardcoded JWT signing secret from the live application, then minted a valid administrator token for a full account takeover, a path every tool's individual findings pointed to but only the dynamic attacker proved. Frontier models found the vulnerabilities. Only the attacker found the chains. October 7, 2026 0 mins read Attackers don't read your repository; they hit your URL, and chain together whatever they find. In an era where offensive AI runs against live applications at machine speed, your security tooling needs to go beyond finding vulnerabilities to prove exploitability, including whether they can be combined into a breach. So we ran a test. We pointed Evo Continuous Offensive Security COS https://snyk.io/evo/continuous-offensive-security/ and Claude Security running Mythos at the same application: TaintedPort https://taintedport.com/ . It’s a deliberately vulnerable web app that exercises real-world bug classes and a set of registered exploit chains. These are different kinds of tools; COS attacks the running application while the two Claude products read source code, and we ran them on the same target to see what each approach proves. Evo COS confirmed 10 of 15 exploit chains. This isn't a knock on models - we use them Claude Security on Mythos is a real advance in AI-driven vulnerability discovery. It reasons about source code the way a strong human reviewer does, and it catches classes of flaws that pattern-matching tools miss. But finding a flaw in the code and proving an attacker can exploit these vulnerabilities are different jobs. The attack path COS proved Every tool in this test found the server-side request forgery SSRF flaw, and that the application's JWT signing secret was hardcoded. Evo COS went further and used the SSRF to reach and retrieve the signing secret from the running application, then used that secret to mint a valid administrator token and take over the admin surface. Two findings, connected into a full account takeover path, were demonstrated against the live app with a runnable proof of concept. Above: CHAIN-004 as scored for COS. The two member findings, the SSRF flaw and the hardcoded secret, were reported individually by every tool in this test. Evo COS confirmed that they combine into admin token forgery. That is what an autonomous attack looks like in practice: get a foothold, then chain from it. Evo COS shows you the path and attaches a runnable proof of concept for each confirmed chain. The results TaintedPort v1.35, scored against a fixed answer key of 57 known vulnerabilities and 15 known exploit chains, on a severity-weighted scale Low 1, Medium 3, High 9, Critical 27; each confirmed chain scored at its own severity on top of its member findings . Weighted detection is points found ÷ 938 points available: 587 from the 57 vulnerabilities and 351 from the 15 chains. | | Evo COS | Claude Security Mythos | |---|---|---| | Severity-weighted detection | 75.7% | 49.6% | | Exploit chains confirmed of 15 | 10 | X | | Vulnerabilities found of 57 | 50 | 37 | | F1 | 91.7 % | 75.5% | Claude Security found one more critical-severity vulnerability 10 to 9 . Evo COS found the most vulnerabilities, had the highest precision 96.2% vs 90.2% with fewer false positives, and it was able to identify and prove exploit chains. Why attack the deployed application? Evo COS is a dynamic system: its target is always a live URL, with source code as an optional input that upgrades a run from black-box to gray-box. It never runs on code alone. Claude Security ran white-box code only . That makes this a cross-disciplinary comparison: Evo COS is an offensive pentest that probes and exploits the running application, a fundamentally more intensive process than a single code-reading pass. It's the same complementary split the industry has always had between DAST and SAST. Here we're showing what exercising the live application proves and what reading the code alone cannot: that these vulnerabilities are real, and that they can chain. The findings reported only by Evo COS are overwhelmingly runtime behaviors: reflected and misconfigured responses, missing transport protections, tokens that survive logout, and weak session handling. Evo COS was able to confirm several of the context-dependent business-logic flaws: broken object-level authorization https://snyk.io/articles/bola-the-api-vulnerability-hiding-in-plain-sight/ on profile updates, privilege escalation via forged JWT claims, and a critical flaw where a single read of a stored TOTP secret nullifies 2FA. The chains require reaching each step through the running app and confirming the next one is actually reachable. A model reasoning over the source has to infer that a vulnerability executes, and misses the chain. You can build a harness, but you can't build the context This year's benchmarking debate has landed on a real point: the system around a model matters more than the model itself. We'd go further. "Which harness" isn't a neutral question, and the answer is where the advantage actually lives. Evo COS orchestrates multiple models, including frontier Claude models, each directed at the job it's best at. So this result isn't a story about model quality. The same class of model that powers a frontier code review sits inside Evo COS. What's different is the system around it: Evo COS starts from what the Snyk platform already knows about the target, rather than reasoning cold; it's a group of agents, each orchestrated for a specific purpose, and every finding clears an independent validation step before it surfaces, because the system that generates a finding shouldn't be the one that grades it. Anyone can wire a frontier model to a coding agent tonight and point it at a repository. What that setup can't reproduce is the accumulated, validated context Evo COS starts from, the independent validator, and the exercise against the live application. Dynamic and static testing are complementary Claude Security caught several vulnerabilities that Evo COS didn't, mostly source-visible logic and cryptographic flaws that don't always surface through the running app. Neither approach is complete on its own; this is the argument for a platform rather than a point tool. Reading the code and attacking the running app, each catches things the other misses. Methodology TaintedPort is Snyk's own deliberately vulnerable application, built and maintained by our team, and it's public. Evo COS was not tuned against it. Target: TaintedPort: 57 known vulnerabilities 34 commodity, 23 business-logic and 15 known exploit chains. Tools and inputs: - Evo COS: gray-box live URL + source , Sep 18. - Claude Security: white-box source , Mythos with extended thinking, Sep 18. Runs: one per tool. Results are the single-run outputs, not medians. Detection by category | Category | Known | Evo COS | Claude Security Mythos | | Commodity | 34 | 33 | 21 | | Business logic | 23 | 17 | 16 | | Exploit chains | 15 | 10 | X | Detection by severity found/known | Severity | Known | Evo COS | Claude Security Mythos | | Critical | 11 | 9 | 10 | | High | 26 | 22 | 19 | | Medium | 18 | 17 | 8 | | Low | 2 | 2 | 0 | False positives and F1 | Metric | Evo COS | Claude Security Mythos | | False positives | 2 | 4 | | F1 score | 91.7% | 75.5% | Reproduce it: TaintedPort is available at taintedport.com http://taintedport.com . Curious which findings in your own applications chain into a breach? See how Evo Continuous Offensive Security tests your live URL https://snyk.io/schedule-a-demo/ . BOOK A LIVE DEMO Secure AI adoption at scale Evo helps organizations safely adopt and scale AI by providing visibility, governance, and security across AI-driven development and AI applications.