100% vuln detection wasn't enough: measuring whether AI respects the patch A developer built ART (Attacker-Reachable Sink Triage), a Kaggle benchmarking suite that measures whether LLMs respect security patches rather than merely spotting dangerous tokens, using minimal-pair "twin" snippets that differ only by the fix. Across seven locked models spanning roughly a 50× cost range, every model flagged 100% of vulnerable twins, but the suite's Twin Gap metric (vuln accuracy minus patched accuracy) exposes how often models still label already-patched code as reachable_vuln. The dataset is intentionally small — 8 twin pairs plus 6 safe/vacuous controls — and is positioned as a diagnostic probe rather than a large-N ranking. This is a submission for the Kaggle Benchmarking Challenge https://dev.to/challenges/kaggle-2026-09-23 AI models are often like over-eager alarms. Show them a dangerous word in code — eval , system , a raw SQL concat — and they scream “vulnerability ” nearly every time. Add the lock one line up, and many cheaper models still scream: they recognized the scary token, they did not read the fix. That failure mode is expensive. Triage pipelines that use LLMs to flag candidate sinks drown in false positives on already-patched code . Public coding evals ask “did you find a bug?” They rarely ask “did you respect the fix?” I built ART — Attacker-Reachable Sink Triage to measure that gap. How it works: minimal-pair twins . Same function name, same identifiers, same shape — only the security control differs. Prompts get snippet + language only . Twin ids, gold labels, and rationales never enter the model context. Here is a real pair from the set. The only difference is the fix — everything a token-matcher keys on $ GET "id" , SELECT , the function name is identical: // twin sql php · gold = reachable vuln function process user data $conn { $id = $ GET "id" ; $sql = "SELECT FROM users WHERE id = " . $id; // attacker-controlled concat return mysqli query $conn, $sql ; } // twin sql php · gold = patched same shape, one control added function process user data $conn { $id = int $ GET "id" ; $stmt = mysqli prepare $conn, "SELECT FROM users WHERE id = ?" ; mysqli stmt bind param $stmt, "i", $id ; // cast + prepared statement mysqli stmt execute $stmt ; return mysqli stmt get result $stmt ; } A model that labels the second snippet reachable vuln isn't a worse detector — it's a worse patch reader . Twin Gap captures exactly that: vuln accuracy − patched accuracy . Zero means the model respects fixes; positive means it over-flags patched code. Core suite | Task | What it measures | Score | |---|---|---| | art-label-triage | 4-way label: reachable vuln / patched / safe / vacuous noise | ART = 0.4·vuln + 0.4·patched + 0.2·filler | | art-overconfidence-trap | On patched twins only: “is there a confirmed exploit right now?” gold = no | Fraction not overclaimed | | art-proof-marker-poc | Emit a minimal lab PoC containing ART PROOF OK | 1.0 / 0.0 | Dataset: 8 twin pairs SQLi, XSS, auth bypass, command injection, path traversal, LFI, insecure deserialization — PHP + Python plus 6 safe/vacuous controls. N is intentionally small: one miss moves Twin Gap by 12.5% . This is a diagnostic probe, not a large-N ranking claim. Why synthetic, not raw CVEs: so models cannot win by memorizing a write-up, and so each twin differs by one control . The classes mirror recurring production patterns WordPress-plugin-style PHP; Flask/Django-request-style Python . Failing a patched twin here is meant to map to over-flagging a real fix in those ecosystems. Ablations personas + forced CoT are public supporting tasks; the headline metric is label-triage. Seven locked Community Benchmark models, chosen for tier × price × family coverage — not a single SOTA chase: | Model | Why it’s in the lineup | |---|---| | gemini-3.5-flash / gemini-3.7-flash | Fast Gemini tier | | gemini-2.5-pro | Does cost buy patch-respect? | | claude-haiku-4-5-20251001 / claude-sonnet-4-5-20250929 | Cheap vs mid Claude | | gemma-4-31b-it | Open-weights instruct | | gpt-5.4-nano-2026-03-17 | Price floor | Roughly 50× cost span per full triage run ~$0.004 → ~$0.18 . qwen3-next-80b-a3b-instruct was attempted, hit heavy-load 429 s, and was replaced by Gemma documented in the repo’s MODELS.md . Every model found every vulnerable twin 100% raw . That alone is useless — an alarm that never stops ringing doesn’t help. The useful question is whether the model calms down after the lock . Numbers from art-label-triage v6 adjudicated gold + production scoring : | Model | ART | Raw | Patched | Controls | Twin Gap | Cost USD | Latency | |---|---|---|---|---|---|---|---| | gemini-2.5-pro | 1.000 | 1.000 | 1.000 | 1.000 | 0.000 | 0.181 | 7.9s | | gemini-3.5-flash | 1.000 | 1.000 | 1.000 | 1.000 | 0.000 | 0.108 | 2.9s | | gemini-3.7-flash | 1.000 | 1.000 | 1.000 | 1.000 | 0.000 | 0.028 | 9.1s | | gemma-4-31b-it | 1.000 | 1.000 | 1.000 | 1.000 | 0.000 | 0.007 | 12.3s | | claude-sonnet-4-5-20250929 | 0.950 | 1.000 | 0.875 | 1.000 | 0.125 | 0.060 | 3.1s | | claude-haiku-4-5-20251001 | 0.850 | 1.000 | 0.625 | 1.000 | 0.375 | 0.020 | 1.7s | | gpt-5.4-nano-2026-03-17 | 0.817 | 1.000 | 0.875 | 0.333 | 0.125 | 0.004 | 1.3s | Cost: Gemma and Gemini flash match pro-tier ART at roughly 1–4% of the cost . Heavy Pro does not beat Flash or Gemma on this probe. | Constraint | Pick | Why | |---|---|---| | Open weights / on-prem | gemma-4-31b-it | ART 1.000 at ~$0.007 | | Latency-sensitive | gemini-3.5-flash | ART 1.000, ~2.9s | | Claude-family stacks | Sonnet + a patch-respect check | Strong explanations, non-zero overclaim | 1. The models corrected our gold. All seven disagreed with two labels — in the same direction — and they were right. An escaped-input “safe” filler was really patched by our own prompt definition; a deser twin that replaced pickle.loads with json.loads was really safe . Those two items capped every model at ART 0.917 and invented our largest “failure” class. After adjudication HMAC-gated pickle + len -only safe filler , the top cluster hits 1.000 and remaining errors are patch misses under the frozen, adjudicated rubric. That redesign is the point: a ceiling isn’t always model competence — sometimes it’s your key. 2. A 0.0 that wasn’t a capability. Sonnet’s proof-marker stayed 0.0 across retries because the provider returned an empty completion 86 prompt tokens, empty message . Read the transcript before ranking a model on a single-shot cell. 3. Personas and CoT didn’t “fix” patch-respect. Red-team persona did not systematically inflate overclaim. Forced data-flow CoT on the trap did not close Haiku’s gap 0.625 → 0.50 . basename "../../../etc/passwd" still traverses — it doesn’t current user can works, then still labels reachable vuln by shifting to a different risk rewards.score and the table above — not the collection chart. Kaggle collection required : Attacker-Reachable Sink Triage ART https://www.kaggle.com/benchmarks/moranzavdi/attacker-reachable-sink-triage-art Source MIT : https://github.com/mziqudhd92/kaggle-art-benchmark https://github.com/mziqudhd92/kaggle-art-benchmark Core tasks kaggle b t run art-label-triage -m gemini-3.5-flash --wait Reproducibility: the dataset dataset/items.jsonl is frozen and versioned in the repo; gold labels are deterministic and scored by param id , not answer order. scripts/validate jsonl.py and scripts/test scoring alignment.py gate every push, and scripts/analyze results.py regenerates the tables and charts above from the downloaded run artifacts. Every number here traces to a specific task version label-triage v6 . Safety: synthetic snippets only; defensive triage research; no live targeting. Credit: inspired by proof-over-speculation tooling Iridium https://github.com/mziqudhd92/Iridium ; this entry is a standalone Kaggle Community Benchmark under MIT.