This is a submission for the Kaggle Benchmarking Challenge
I built a 38-case security vulnerability detection benchmark on Kaggle that tests whether LLMs can identify vulnerabilities in Python code with deliberate false-positive traps designed to distinguish data-flow reasoning from pattern matching.
The itch came from studying for CEH v13. Every week, another "AI security assistant" launches claiming to replace human code review. I wanted to know: do LLMs actually reason about security, or do they just recognize recognizable shapes?
Most security benchmarks test textbook cases. Real code review is harder. Consider these three snippets:
subprocess.run(["cat", filename]) # no shell β no injection
expression = OPERATIONS[user_input]
result = eval(expression) # user controls what gets eval'd
eval("2 + 2") # safe β hardcoded
eval(user_input) # dangerous β attacker-controlled
A model that passes textbook tests but fails these three is not ready for security work.
Benchmark Design
38 cases across 12 vulnerability categories:
SQL Injection (5) β including column-name injection via f-strings
Command Injection (5) β including shell=True with list args
Hardcoded Secrets (4) β including config vs. credential distinction
Insecure Deserialization (5) β including trusted vs. untrusted sources
Code Injection / Eval (4) β including restricted-globals eval
Password Hashing (4) β including MD5 checksums vs. MD5 password hashes
Path Traversal (3) β including Path.resolve() containment
SSRF (3) β including URL allowlisting
ReDoS (2) β nested quantifiers
JWT Algorithm Confusion (2) β algorithms=["HS256", "none"]
Authorization Bypass (2) β falsy-comparison bugs
TOCTOU Race Conditions (2) β check-then-use patterns
5 explicit false-positive traps
Three difficulty tiers: easy (16), medium (19), hard (19).
Metrics measured per run:
Binary accuracy (vulnerable vs. safe)
True Positive Rate / False Negative Rate (missed vulnerabilities)
False Positive Rate (wrongly flagged safe code β the metric that matters most for triage)
Vulnerability naming accuracy
Per-difficulty breakdown
The capability I set out to measure wasn't "can the model detect SQL injection" β that's saturated. I wanted to measure whether the model can distinguish dangerous code from code that merely looks dangerous. That's the actual skill a security reviewer needs.
I tested 9 models across 5 providers, spanning three tiers lightweight, flagship, and reasoning. The lineup was chosen to answer three specific questions:
The three questions:
Do flagships beat lightweight models at security triage β or is the pattern-matching ceiling low enough that cheap models suffice?
Do reasoning models outperform standard ones on security tasks that require multi-step inference?
What's the cost/latency tradeoff at each tier β and does the most accurate model actually make sense in production?
This lineup gives me a clean read on all three. If a $0.02 model matches a $0.47 model, that changes how security teams should build triage pipelines. If reasoning models fail, that's a finding about the maturity of the capability.
Insight 1: The Cheapest, Fastest Model Wins at Production Triage
Gemini 2.5 Pro costs 5.8x more than Claude Haiku 4.5 and is 30x slower β for the same 0.92 score.
This is the finding that surprised me most. I expected the flagship Gemini model to pull ahead on hard cases. It didn't. Both scored 0.92. Both passed 35 of 38 assertions. The difference is that one costs $0.466 per run and takes 6 minutes, and the other costs $0.080 and takes 12 seconds.
If you're building a security triage pipeline scanning 10,000 files:
Same accuracy. 6x cheaper, 30x faster. The production choice is obvious.
GPT-5.4 is the interesting middle ground. It scored 0.97 - the joint highest - at 27 seconds and $0.105. If you need maximum accuracy and reasonable latency, that's the pick. If you need maximum throughput at scale, Haiku wins.
GPT-5.4 mini is the budget option at $0.023 and 0.82 accuracy. For pre-filtering large corpora before human review, that tradeoff might be worth it. It depends on what your false-positive tolerance is.
Insight 2: Latency Is a Hidden Failure Mode β and Every Benchmark Ignores It
Three models took longer than 5 minutes per full run:
For a tool meant to fit into a developer's workflow, a 5-minute review per file means abandonment. A 12-second review (Haiku) means it becomes part of the process. Security teams should treat latency as a first-class metric, not an afterthought.
A model that scores 0.97 in 5 minutes is functionally worse than a model that scores 0.92 in 12 seconds β because the second one actually gets used. Leaderboards reward accuracy; users reward responsiveness.
Insight 3: Format Compliance Skews Every LLM Benchmark
Qwen3 Next 80B scored 0.68. Look at what actually happened:
`Qwen said:
VULNERABLE: YES
Vulnerability: SQL injection`
The model was correct. But my regex expected VULNERABLE: YES without markdown bold. Every assertion failed β 0/38. Not because Qwen was wrong, but because it formatted its output differently.
This is a universal problem with LLM benchmarks β they measure format compliance as much as reasoning capability. A model can be correct and score zero. Anyone building evaluation infrastructure needs to handle output-format variance explicitly, or risk misjudging models by 30+ points.
Qwen's true detection capability is likely much higher than 0.68 β but the formatting failure makes it unusable as-is in an automated pipeline. That's a real operational finding, not a benchmark artifact.
Insight 4: Reasoning Models Are Not Ready for High-Throughput Triage
Both DeepSeek-R1 and Claude Sonnet 4.5 errored out. DeepSeek failed after only 3 of 38 cases β likely a token or time budget exhaustion caused by its long chain-of-thought outputs.
For a security tool that needs to process thousands of files daily, models that require reasoning traces to be serialized and returned are operationally risky. A tool that fails on 1 in 4 inputs is worse than a slightly less accurate tool that always responds.
Grok 4.20 Reasoning was the exception β it completed the run at 0.95 with 80-second latency. So the finding isn't "reasoning models fail" β it's "reasoning models are variable in their operational reliability," and that variance matters at scale.
Insight 5: Even the Best Models Fail on the Same Case
Gemini 3.7 Flash scored 0.97 β one failure out of 38. That single failure was CMD-004:
filename = user_input
subprocess.run(["cat", filename], check=True)
This is safe. The user controls the filename, and yes, that's a validation concern β but there's no command injection here. No shell. No traversal. The cat command reads exactly the file it's told to read. The model flagged it as vulnerable because the variable was named filename.
GPT-5.4 failed the exact same case. So did Grok, Haiku, and Gemini 2.5 Pro. Five models, one shared false positive.
This is the universal failure mode: variable-name bias. When a variable is called user_input, filename, or user_path, all models reflexively flag it β regardless of whether the operation is actually dangerous. This is pattern matching, not reasoning.
And the flip side is worse. The same model missed EVAL-004 β expression = OPERATIONS[user_input]; result = eval(expression) β which is vulnerable. The user selects which expression gets evaluated. But because the syntax doesn't look like eval(user_input), the model gave it a pass.
False positives generate noise. False negatives generate breaches. The model is biased toward both.
What I'd Measure Next
Variable-name swap experiment. Rename user_input to config_value in the same snippet. If the model stops flagging it, we've proven variable-name bias. If it still flags it, the model is reasoning about data flow. This single experiment would settle whether LLMs are doing real analysis or shape recognition.
Calibration measurement. Ask each model for a confidence score alongside its verdict. When a model says "95% confident," how often is it right? Overconfidence is a critical safety failure for security triage. If a model is wrong 15% of the time but says "100% confident" every time, it's dangerous β not useful.
Format-normalized scoring. Rebuild the benchmark with a parser-agnostic layer that extracts verdicts regardless of markdown, JSON, or plain-text formatting. Re-run Qwen3 to see its true capability. This would help the entire LLM evaluation community, not just this benchmark.
Multi-turn triage. Give the model a full file, ask it to triage, then ask "why did you flag this?" If the reasoning contradicts the verdict, the model is pattern matching β regardless of how accurate it looks on the leaderboard.
Security Vulnerability Detection v2 on Kaggle
Full notebook with source data
The benchmark contains:
Fork it. Add your own cases. Test your own models. The point of a benchmark isn't to have the final word it's to start a better conversation. Every case, every rationale, and every assertion is public so anyone can extend this.