ProofSec: Benchmarking Epistemic Robustness and Evidence-Grounded Vulnerability Reasoning in Frontier LLM A developer has released ProofSec, an evidence-centric security reasoning benchmark that tests whether large language models can distinguish security indicators from substantiating evidence. ProofSec v0.2 contains 110 security reasoning cases requiring models to classify scenarios as Vulnerable, Not Vulnerable, or Insufficient Evidence, with controlled evaluation families covering one-fact flips, evidence-state changes, and contradictions. The benchmark's central principle is that a security indicator is not equivalent to security evidence, and it evaluates whether models can recognize when a security conclusion is not epistemically justified. What happens when an LLM recognizes every lexical and semantic signature associated with a vulnerability - IDOR, BOLA, authorization bypass, predictable identifiers - but the available evidence does not actually establish that the vulnerability exists? That question is the foundation of ProofSec , an evidence-centric security reasoning benchmark designed to evaluate whether LLMs can distinguish security indicators from substantiating evidence , reason under incomplete and contradictory observations, resist terminology and authority bias, incorporate falsifying evidence, and explicitly recognize when a security conclusion is not epistemically justified. Large language models have become exceptionally capable at semantic retrieval. Give a model: GET /api/users/2841/invoices/9281 and introduce: predictable numeric identifiers and the model can immediately activate a large cybersecurity concept space: IDOR BOLA Broken Access Control Authorization Bypass But there is a fundamental distinction between recognizing the semantic signature of a vulnerability and demonstrating that the underlying security property has actually been violated . That distinction is where many security reasoning systems become unreliable. Consider: Authenticated user: 2841 GET /api/invoices/9281 HTTP/1.1 200 OK { "invoice id": 9281, "amount": 45000, "status": "paid" } Is this an IDOR? The evidence is insufficient to establish that conclusion. The observation does not independently establish: 9281 A predictable identifier is an indicator . An HTTP 200 response is an observation . Neither fact, in isolation, constitutes proof of unauthorized cross-principal access. The central principle behind ProofSec therefore became: A security indicator is not equivalent to security evidence. This sounds straightforward. In practice, it is a surprisingly difficult property for language models to maintain under adversarial framing, incomplete evidence, contradictory observations, and highly salient security terminology. ProofSec evaluates evidence-sensitive vulnerability classification . Every scenario requires the model to classify the security state into exactly one canonical category: Vulnerable Not Vulnerable Insufficient Evidence The third state is deliberately first-class. Traditional binary vulnerability classification implicitly encourages a forced decision: YES or NO Real security investigations rarely provide that luxury. Evidence can be incomplete. Observations can be ambiguous. Telemetry can be contradictory. Authorization boundaries can be unknown. A response can be suspicious without being conclusive. ProofSec therefore evaluates whether an LLM can recognize an epistemic boundary : The available evidence is insufficient to establish the claim. This changes the benchmark from a conventional vulnerability-recognition exercise into an evaluation of: The benchmark is therefore not primarily asking: "Does the model know what IDOR means?" It is asking: "Can the model determine whether the evidence presented actually establishes the security property in question?" ProofSec v0.2 contains 110 security reasoning cases . The corpus is deliberately constructed as an experimental evaluation set rather than an undifferentiated collection of cybersecurity questions. The benchmark incorporates controlled evaluation families covering: The architecture is designed around a core principle: If the evidentiary state changes, the model should be sensitive to that change. Conceptually: ProofSec │ ┌──────────────┼──────────────┐ │ │ │ One-Fact Flip Evidence State Contradiction │ │ │ └──────────────┼──────────────┘ │ Security Scenario │ ▼ LLM Inference │ ▼ Structured Assessment │ ┌──────────────┼──────────────┐ │ │ │ Classification Evidence State Explanation │ ▼ Deterministic Assertion │ ▼ Numerical Score The implementation deliberately isolates: Dataset ↓ Prompt Construction ↓ Model Inference ↓ Structured Parsing ↓ Assertion ↓ Scoring This separation is not cosmetic. It establishes distinct failure domains . A dataset defect is not a model failure. A packaging failure is not a reasoning failure. A schema violation is not necessarily a classification failure. A scoring-interface defect is not a model-performance measurement. That distinction became one of the most important engineering principles in the project. One of the core mechanisms in ProofSec is controlled one-fact perturbation . Instead of constructing completely unrelated questions, ProofSec creates scenarios where a security-relevant fact changes while much of the surrounding semantic structure remains invariant. For example: Authenticated user = 2841 Invoice owner = 2841 Authorization rule: invoice.owner id == authenticated user.id Authenticated user = 2841 Invoice owner = 9127 Authorization rule: invoice.owner id == authenticated user.id The semantic surface can remain substantially similar. The critical security relationship changes. That gives us a controlled counterfactual: Δ Input ↓ Δ Security-Relevant Fact ↓ Expected Δ Security State This is fundamentally more informative than asking: "Is this an IDOR?" The actual evaluation question becomes: Did the model detect the fact that causally changes the security state? This introduces a form of counterfactual sensitivity testing . If one authorization-relevant fact changes and the model preserves the same classification, the benchmark exposes a specific failure mode. The model may understand the terminology. It may understand the vulnerability class. But it may not be correctly conditioning its conclusion on the evidence that actually determines the security property. ProofSec explicitly models the evidentiary state of a scenario. The benchmark uses states including: WEAK PARTIAL DECISIVE CONTRADICTORY NEGATIVE UNKNOWN This creates an evidence hierarchy rather than treating every security-relevant observation as equally probative. A simplified progression is: Predictable identifier │ ▼ Weak security indicator │ ▼ Ownership relationship established │ ▼ Cross-principal access demonstrated │ ▼ Authorization invariant violated │ ▼ DECISIVE evidence Consider the difference between: Predictable ID and: Unauthorized cross-user object access These observations possess radically different evidentiary weight. The first may justify investigation. The second can establish a concrete violation of an authorization property. ProofSec therefore attempts to prevent a model from collapsing the following distinction: Suspicion ≠ Evidence ≠ Proof That distinction is foundational to trustworthy security analysis. Real security investigations are not monotonic. Evidence can conflict. Telemetry can be stale. Caches can contain misleading artifacts. Fixtures can resemble production responses. A preliminary observation can subsequently be invalidated by a higher-authority observation. ProofSec therefore incorporates contradiction-resolution scenarios . A simplified reasoning sequence might look like: Initial observation │ ▼ HTTP 200 response │ ▼ Potential authorization anomaly │ ├───────────────┐ │ │ ▼ ▼ Synthetic fixture Actual endpoint │ │ ▼ ▼ Cached response HTTP 403 │ │ └───────┬───────┘ ▼ Re-evaluate hypothesis The important capability is not simply detecting the first suspicious observation. The model must determine whether later evidence changes the validity of the original hypothesis. This introduces a critical distinction between: Evidence accumulation Evidence revision A model that merely accumulates confirming signals can behave very differently from one capable of hypothesis revision under contradictory evidence . ProofSec deliberately tests that boundary. Many vulnerability benchmarks are heavily oriented toward positive findings. ProofSec deliberately gives negative evidence first-class status. User A requests User B's object │ ▼ Authorization rule verified │ ▼ Cross-user request returns 403 │ ▼ Unauthorized access not demonstrated A security reasoning system must be able to update its hypothesis in both directions. It must identify: Evidence supporting vulnerability Evidence weakening or falsifying vulnerability hypothesis This matters because security investigation is fundamentally adversarial. The objective is not to collect evidence that confirms the initial hypothesis. The objective is to determine whether the hypothesis survives attempts to falsify it. That makes negative evidence an important component of hypothesis discrimination . Another failure mode targeted by ProofSec is authority and terminology bias . Cybersecurity vocabulary carries enormous semantic weight. Terms such as: critical confirmed exploit CVE researcher privilege escalation authorization bypass security issue can exert disproportionate influence over LLM outputs. ProofSec therefore evaluates whether the model remains anchored to technical evidence when the surrounding terminology changes. The underlying principle is: Terminology describes a claim. Evidence substantiates it. Consider two descriptions: Possible authorization issue Confirmed critical authorization vulnerability If the underlying technical evidence is unchanged, the semantic framing should not arbitrarily alter the actual security state. This probes linguistic framing sensitivity and authority-induced classification drift . The benchmark is therefore not only testing cybersecurity knowledge. It is testing whether that knowledge remains subordinate to the evidentiary record. I intentionally avoided unconstrained free-form text as the primary evaluation interface. Each model response is constrained into a structured schema containing: classification evidence state supporting evidence missing evidence safe verification impact This creates a deterministic machine-readable boundary between inference and evaluation. The primary scoring target is: classification The remaining fields provide diagnostic observability. classification: Vulnerable could be accompanied by: supporting evidence: "The identifier is sequential." That exposes a potentially significant evidentiary defect: Predictable Identifier ≠ Proven Unauthorized Access The final classification alone cannot reveal whether the model arrived at the conclusion through valid or invalid reasoning. Structured assessment therefore provides a richer diagnostic surface . I deliberately avoided making another LLM the primary arbiter of the classification. ProofSec uses canonical ground-truth labels and deterministic assertions. The core evaluation chain is: Scenario ↓ Model ↓ Structured Classification ↓ Canonical Ground Truth ↓ Exact Assertion ↓ Score rather than: Scenario ↓ Model A ↓ Model B judges Model A ↓ Score For the core classification metric, deterministic evaluation establishes a cleaner experimental boundary. It also reduces the possibility of evaluation contamination , where the judgment model introduces another layer of model-dependent interpretation into the primary metric. A benchmark is only scientifically useful if its evaluation corpus is reproducible. ProofSec v0.2 uses a frozen dataset representation with a SHA-256 integrity digest: 422501a4db424c30c8ef24b61183351ec8a4bd2096e2671cf0e6bdf91e133a80 The project maintains a manifest and an integrity-verification workflow. The intended reproducibility invariant is: Same benchmark version + Same task corpus + Same evaluation protocol = Comparable experiment Without corpus integrity, two apparently identical experiments can silently evaluate different datasets. That means dataset provenance, versioning, and integrity verification are part of the experimental methodology itself. One of the most valuable discoveries during the project was that benchmark engineering is itself part of evaluation science . My initial task implementation used: php - None with assertions. That produced pass/fail-style behavior. I initially expected the benchmark to expose a numerical accuracy metric. That assumption was incorrect. I subsequently changed the task interface to: php - float and explicitly returned: accuracy = correct / total return float accuracy This exposed a broader principle: The semantics of the evaluation harness are part of the experimental design. A benchmark can execute successfully while still reporting an invalid measurement if the scoring contract is incorrectly defined. In other words: Correct Inference ≠ Correct Evaluation A trustworthy benchmark requires both. One early published iteration failed with: NameError: name 'ALL TASKS JSON' is not defined This was not a model reasoning failure. The model never reached the reasoning stage. The task failed during execution before the first model inference step. That forced a more rigorous separation of: Dataset Integrity │ ▼ Task Packaging │ ▼ Runtime Execution │ ▼ Model Inference │ ▼ Structured Parsing │ ▼ Assertion │ ▼ Scoring This distinction is fundamental. A benchmark must not silently transform: Runtime Error into: Model Error Those are different failure classes with different remediation paths. It also exposed an important deployment boundary. My local development environment contained the canonical task source under: C:\Users\User\Documents\ProofSec\kaggle\tasks but a Kaggle-hosted execution environment cannot directly access that Windows filesystem. The benchmark therefore has to package the necessary evaluation corpus and task implementation into the executable artifact available to the remote runtime. That became an important lesson in benchmark portability and execution-environment isolation . The benchmark went through multiple iterations while I validated the evaluation infrastructure. One development version displayed assertion counts including: Claude Sonnet 4.6 0 Pass GPT-5.5 75 Pass Gemini 3.5 Flash 95 Pass Gemini 3.7 Flash 87 Pass A subsequent controlled version successfully executed its smoke-test assertions across the evaluated models. These development-stage measurements should not be interpreted as a universal model ranking. They represent different stages of benchmark engineering. During this process I was simultaneously validating: This separation is essential for responsible interpretation. A benchmark-development artifact is not automatically equivalent to a final scientific measurement. Before treating Kaggle execution as authoritative, I also used local evaluation to inspect the complete 110-case corpus. Representative local validation produced: | Model | Correct | Total | Accuracy | |---|---|---|---| | Claude Sonnet 4.6 | 80 | 110 | 72.73% | | Gemini 3.7 Flash | 86 | 110 | 78.18% | | Gemini 3.5 Flash | 84 | 110 | 76.36% | | GPT-5.5 | 67 | 110 | 60.91% | These figures are local validation measurements under the specific execution conditions used during development . They should not be interpreted as immutable measurements of model capability. LLM inference can vary because of: This variability is itself relevant to reproducibility. The objective was therefore not to reduce the benchmark to: "Model X is better." The objective was to characterize failure modes, evidentiary behavior, and reasoning sensitivity under a controlled corpus . One of the recurring failure patterns motivating ProofSec can be represented as: Security Vocabulary Recognition ↓ Semantic Activation ↓ Premature Classification ↓ Evidence Is Never Actually Tested The desired reasoning pathway is: Scenario ↓ Extract Security-Relevant Claims ↓ Identify Actors and Authorization Boundaries ↓ Identify Directly Observed Facts ↓ Separate Indicators from Decisive Evidence ↓ Evaluate Contradictory Evidence ↓ Evaluate Negative Evidence ↓ Determine Evidentiary Sufficiency ↓ Classify The distinction is fundamental. A model can possess extensive cybersecurity knowledge while still being unreliable at evidence-grounded adjudication . This is the capability ProofSec attempts to isolate. This is one of the most consequential design decisions in ProofSec. A conventional benchmark might ask: Is the system vulnerable? YES / NO ProofSec instead represents: Vulnerable Not Vulnerable Insufficient Evidence This explicitly models epistemic uncertainty . The model must distinguish between: Observed Established That distinction matters operationally. Unsupported Finding ↓ Analyst Investigation ↓ Engineering Interruption ↓ Potential Remediation ↓ Operational Cost A system that generates large quantities of unsupported security findings can impose substantial downstream cost even when its vulnerability-recognition capability appears impressive. Abstention is therefore not merely a failure to classify. Under uncertainty, abstention can be a legitimate and necessary security behavior. At a deeper level, ProofSec is not simply a cybersecurity benchmark. It evaluates whether an LLM can maintain an explicit boundary between: What do I know? ↓ What does the evidence establish? ↓ What remains unknown? ↓ What conclusion is justified? This is fundamentally an epistemic reasoning problem . Cybersecurity provides a particularly concrete environment in which to measure it because security conclusions often have direct operational consequences. The same methodology could potentially extend to: The security domain is simply where I chose to operationalize the problem. ProofSec is implemented using: kaggle benchmarks The architecture maintains explicit boundaries between: DATA PROMPT CONSTRUCTION MODEL INFERENCE STRUCTURED PARSING EVALUATION SCORING This makes the system substantially easier to: The benchmark is therefore treated as an engineered evaluation artifact rather than merely a collection of prompts. The development workflow became: Security Hypothesis ↓ Construct Controlled Scenario ↓ Define Canonical Ground Truth ↓ Generate Adversarial / Paired Variant ↓ Validate Dataset ↓ Freeze Corpus ↓ Generate Integrity Hash ↓ Implement Structured Task ↓ Run Local Validation ↓ Validate Kaggle Packaging ↓ Execute Against Multiple Models ↓ Inspect Assertions ↓ Analyze Failure Modes ↓ Revise Benchmark Infrastructure This treats the benchmark as a versioned research artifact. The project structure is maintained as an engineering repository containing the benchmark implementation, evaluation infrastructure, integrity controls, and supporting documentation. The objective is to make the benchmark auditable rather than opaque. The presence of: ID endpoint HTTP 200 user invoice does not automatically establish IDOR. The authorization relationship is the security property that matters. If a single security-relevant fact changes, the model should respond to that change. Controlled perturbation therefore provides a much stronger probe of reasoning sensitivity than simply increasing the number of unrelated benchmark questions. A security reasoning system must be capable of recognizing evidence that weakens or falsifies its initial hypothesis. A model that performs well under clean, monotonic evidence may behave very differently when observations conflict. Contradiction handling is therefore a meaningful robustness dimension. "Insufficient Evidence" can be treated as an explicit epistemic state rather than an evasive response. A classification tells us: What did the model decide? A structured assessment can additionally expose: What evidence did it identify? What evidence did it consider missing? What verification did it propose? That creates a substantially richer diagnostic surface. Dataset defects, packaging failures, schema violations, runtime exceptions, parsing failures, and model reasoning failures require separate failure categories. A benchmark without corpus integrity, version control, and execution controls can silently drift. A score without provenance is considerably less useful than a score attached to a reproducible experimental configuration. ProofSec v0.2 focuses on controlled evidence reasoning. The next iteration can extend the methodology considerably. Instead of presenting all evidence simultaneously: T0 → Initial Telemetry T1 → Additional Observation T2 → Contradictory Artifact T3 → Authorization Test I want to evaluate whether the model updates its hypothesis correctly as evidence arrives. This transforms static classification into a sequential evidence-update problem . Instead of evaluating only: classification evaluate: classification + confidence Then measure whether confidence is statistically aligned with empirical correctness. A model that is wrong with extremely high confidence represents a fundamentally different operational risk profile from a model that identifies uncertainty. Require the model to answer: Which exact fact changed your classification? This enables a deeper evaluation: Correct Conclusion + Correct Evidentiary Basis versus: Correct Conclusion + Incorrect Evidentiary Basis This distinction matters because a model can occasionally arrive at the correct answer for the wrong reasons. The framework can be extended into: SSRF Authentication Bypass Privilege Escalation CSRF Path Traversal Race Conditions Business Logic Flaws Cloud IAM API Authorization Multi-Tenant Isolation The broader research question becomes: Is evidence-grounded security reasoning a generalizable capability, or is it primarily a collection of domain-specific semantic associations? When I started building ProofSec, the obvious question was: Which LLM is better at cybersecurity? After engineering the benchmark, that no longer seemed like the most interesting question. The more fundamental question is: Can we construct evaluations that determine whether an AI security conclusion is actually justified by the evidence available to it? A leaderboard such as: Model A - 82% Model B - 79% Model C - 76% is useful. But it is incomplete. For an AI security system, I want to know: What evidence did the model use? What evidence did it ignore? Did it infer facts that were never provided? Did it confuse an indicator with proof? Did it recognize negative evidence? Did it resolve contradictions? Did terminology alter its classification? Did it respond to a one-fact perturbation? Did it revise its hypothesis? Did it recognize when the evidence was insufficient? Those questions are much closer to the requirements of a real security-analysis system. ProofSec v0.2 contains 110 security reasoning cases designed around controlled evidence manipulation, structured classification, deterministic ground truth, integrity verification, and reproducible multi-model evaluation. The benchmark is built around one principle: Security conclusions should be proportional to the evidence supporting them. ProofSec operationalizes that principle through: The benchmark implementation and supporting engineering work are available on GitHub: The repository is intended to make the benchmark more than a Kaggle submission. It provides an engineering artifact around the evaluation methodology, allowing the benchmark to be inspected, versioned, reproduced, and extended independently of the Kaggle interface. The broader objective is to preserve the chain: Research Hypothesis ↓ Dataset Construction ↓ Ground Truth ↓ Evaluation Harness ↓ Integrity Verification ↓ Model Execution ↓ Results ↓ Failure Analysis That provenance matters. A benchmark should not only report a number. It should make it possible to understand where that number came from . The most important result from building ProofSec was not a single model score. It was discovering how easily an evaluation can appear to measure security reasoning while actually measuring something else. A benchmark can accidentally measure: Keyword Recognition instead of: Evidence Reasoning It can measure: Runtime Correctness Model Correctness Confidence Evidence-Justified Certainty Semantic Familiarity Causal Sensitivity And it can report a precise-looking number while the underlying dataset, runtime, scoring semantics, or evaluation protocol is not actually controlled. ProofSec is my attempt to make those boundaries explicit. The objective is not simply to build an LLM that identifies the largest number of possible vulnerabilities. The objective is to construct an evaluation framework capable of answering a substantially harder question: When an AI claims that a security vulnerability exists, did the available evidence actually justify that claim? That is the standard I want AI-assisted security tooling to eventually be held to. Not: Did the model recognize the vulnerability vocabulary? But: Did the model establish the security claim from the evidence? That distinction is the entire premise of ProofSec. ProofSec Benchmark ProofSec Source Repository Kaggle Benchmarking Challenge Kaggle Benchmarking Challenge on DEV https://dev.to/challenges/kaggle-2026-09-23