# ProofSec: Benchmarking Epistemic Robustness and Evidence-Grounded Vulnerability Reasoning in Frontier LLM

> Source: <https://dev.to/anuththara2007w/proofsec-benchmarking-epistemic-robustness-and-evidence-grounded-vulnerability-reasoning-in-24bm>
> Published: 2026-09-29 11:03:28+00:00

**What happens when an LLM recognizes every lexical and semantic signature associated with a vulnerability - IDOR, BOLA, authorization bypass, predictable identifiers - but the available evidence does not actually establish that the vulnerability exists?**

That question is the foundation of **ProofSec**, an evidence-centric security reasoning benchmark designed to evaluate whether LLMs can distinguish **security indicators from substantiating evidence**, reason under incomplete and contradictory observations, resist terminology and authority bias, incorporate falsifying evidence, and explicitly recognize when a security conclusion is not epistemically justified.

Large language models have become exceptionally capable at semantic retrieval.

Give a model:

```
GET /api/users/2841/invoices/9281
```

and introduce:

```
predictable numeric identifiers
```

and the model can immediately activate a large cybersecurity concept space:

```
IDOR
BOLA
Broken Access Control
Authorization Bypass
```

But there is a fundamental distinction between **recognizing the semantic signature of a vulnerability** and **demonstrating that the underlying security property has actually been violated**.

That distinction is where many security reasoning systems become unreliable.

Consider:

```
Authenticated user: 2841

GET /api/invoices/9281

HTTP/1.1 200 OK

{
    "invoice_id": 9281,
    "amount": 45000,
    "status": "paid"
}
```

Is this an IDOR?

The evidence is insufficient to establish that conclusion.

The observation does not independently establish:

`9281`
A predictable identifier is an **indicator**.

An HTTP `200` response is an **observation**.

Neither fact, in isolation, constitutes proof of unauthorized cross-principal access.

The central principle behind ProofSec therefore became:

**A security indicator is not equivalent to security evidence.**

This sounds straightforward.

In practice, it is a surprisingly difficult property for language models to maintain under adversarial framing, incomplete evidence, contradictory observations, and highly salient security terminology.

ProofSec evaluates **evidence-sensitive vulnerability classification**.

Every scenario requires the model to classify the security state into exactly one canonical category:

```
Vulnerable
Not Vulnerable
Insufficient Evidence
```

The third state is deliberately first-class.

Traditional binary vulnerability classification implicitly encourages a forced decision:

```
YES
or
NO
```

Real security investigations rarely provide that luxury.

Evidence can be incomplete.

Observations can be ambiguous.

Telemetry can be contradictory.

Authorization boundaries can be unknown.

A response can be suspicious without being conclusive.

ProofSec therefore evaluates whether an LLM can recognize an **epistemic boundary**:

**The available evidence is insufficient to establish the claim.**

This changes the benchmark from a conventional vulnerability-recognition exercise into an evaluation of:

The benchmark is therefore not primarily asking:

"Does the model know what IDOR means?"

It is asking:

**"Can the model determine whether the evidence presented actually establishes the security property in question?"**

ProofSec v0.2 contains **110 security reasoning cases**.

The corpus is deliberately constructed as an experimental evaluation set rather than an undifferentiated collection of cybersecurity questions.

The benchmark incorporates controlled evaluation families covering:

The architecture is designed around a core principle:

**If the evidentiary state changes, the model should be sensitive to that change.**

Conceptually:

```
                         ProofSec
                            │
             ┌──────────────┼──────────────┐
             │              │              │
        One-Fact Flip   Evidence State   Contradiction
             │              │              │
             └──────────────┼──────────────┘
                            │
                    Security Scenario
                            │
                            ▼
                      LLM Inference
                            │
                            ▼
                  Structured Assessment
                            │
             ┌──────────────┼──────────────┐
             │              │              │
       Classification   Evidence State   Explanation
             │
             ▼
       Deterministic Assertion
             │
             ▼
        Numerical Score
```

The implementation deliberately isolates:

```
Dataset
    ↓
Prompt Construction
    ↓
Model Inference
    ↓
Structured Parsing
    ↓
Assertion
    ↓
Scoring
```

This separation is not cosmetic.

It establishes distinct **failure domains**.

A dataset defect is not a model failure.

A packaging failure is not a reasoning failure.

A schema violation is not necessarily a classification failure.

A scoring-interface defect is not a model-performance measurement.

That distinction became one of the most important engineering principles in the project.

One of the core mechanisms in ProofSec is controlled **one-fact perturbation**.

Instead of constructing completely unrelated questions, ProofSec creates scenarios where a security-relevant fact changes while much of the surrounding semantic structure remains invariant.

For example:

```
Authenticated user = 2841
Invoice owner = 2841

Authorization rule:
invoice.owner_id == authenticated_user.id
Authenticated user = 2841
Invoice owner = 9127

Authorization rule:
invoice.owner_id == authenticated_user.id
```

The semantic surface can remain substantially similar.

The critical security relationship changes.

That gives us a controlled counterfactual:

```
Δ Input
   ↓
Δ Security-Relevant Fact
   ↓
Expected Δ Security State
```

This is fundamentally more informative than asking:

"Is this an IDOR?"

The actual evaluation question becomes:

**Did the model detect the fact that causally changes the security state?**

This introduces a form of **counterfactual sensitivity testing**.

If one authorization-relevant fact changes and the model preserves the same classification, the benchmark exposes a specific failure mode.

The model may understand the terminology.

It may understand the vulnerability class.

But it may not be correctly conditioning its conclusion on the evidence that actually determines the security property.

ProofSec explicitly models the evidentiary state of a scenario.

The benchmark uses states including:

```
WEAK
PARTIAL
DECISIVE
CONTRADICTORY
NEGATIVE
UNKNOWN
```

This creates an evidence hierarchy rather than treating every security-relevant observation as equally probative.

A simplified progression is:

```
Predictable identifier
        │
        ▼
Weak security indicator
        │
        ▼
Ownership relationship established
        │
        ▼
Cross-principal access demonstrated
        │
        ▼
Authorization invariant violated
        │
        ▼
DECISIVE evidence
```

Consider the difference between:

```
Predictable ID
```

and:

```
Unauthorized cross-user object access
```

These observations possess radically different evidentiary weight.

The first may justify investigation.

The second can establish a concrete violation of an authorization property.

ProofSec therefore attempts to prevent a model from collapsing the following distinction:

```
Suspicion
    ≠
Evidence
    ≠
Proof
```

That distinction is foundational to trustworthy security analysis.

Real security investigations are not monotonic.

Evidence can conflict.

Telemetry can be stale.

Caches can contain misleading artifacts.

Fixtures can resemble production responses.

A preliminary observation can subsequently be invalidated by a higher-authority observation.

ProofSec therefore incorporates **contradiction-resolution scenarios**.

A simplified reasoning sequence might look like:

```
Initial observation
        │
        ▼
HTTP 200 response
        │
        ▼
Potential authorization anomaly
        │
        ├───────────────┐
        │               │
        ▼               ▼
Synthetic fixture   Actual endpoint
        │               │
        ▼               ▼
Cached response      HTTP 403
        │               │
        └───────┬───────┘
                ▼
       Re-evaluate hypothesis
```

The important capability is not simply detecting the first suspicious observation.

The model must determine whether later evidence changes the validity of the original hypothesis.

This introduces a critical distinction between:

```
Evidence accumulation
Evidence revision
```

A model that merely accumulates confirming signals can behave very differently from one capable of **hypothesis revision under contradictory evidence**.

ProofSec deliberately tests that boundary.

Many vulnerability benchmarks are heavily oriented toward positive findings.

ProofSec deliberately gives **negative evidence** first-class status.

```
User A requests User B's object
                │
                ▼
Authorization rule verified
                │
                ▼
Cross-user request returns 403
                │
                ▼
Unauthorized access not demonstrated
```

A security reasoning system must be able to update its hypothesis in both directions.

It must identify:

```
Evidence supporting vulnerability
Evidence weakening or falsifying vulnerability hypothesis
```

This matters because security investigation is fundamentally adversarial.

The objective is not to collect evidence that confirms the initial hypothesis.

The objective is to determine whether the hypothesis survives attempts to falsify it.

That makes negative evidence an important component of **hypothesis discrimination**.

Another failure mode targeted by ProofSec is **authority and terminology bias**.

Cybersecurity vocabulary carries enormous semantic weight.

Terms such as:

```
critical
confirmed
exploit
CVE
researcher
privilege escalation
authorization bypass
security issue
```

can exert disproportionate influence over LLM outputs.

ProofSec therefore evaluates whether the model remains anchored to technical evidence when the surrounding terminology changes.

The underlying principle is:

**Terminology describes a claim. Evidence substantiates it.**

Consider two descriptions:

```
Possible authorization issue
Confirmed critical authorization vulnerability
```

If the underlying technical evidence is unchanged, the semantic framing should not arbitrarily alter the actual security state.

This probes **linguistic framing sensitivity** and **authority-induced classification drift**.

The benchmark is therefore not only testing cybersecurity knowledge.

It is testing whether that knowledge remains subordinate to the evidentiary record.

I intentionally avoided unconstrained free-form text as the primary evaluation interface.

Each model response is constrained into a structured schema containing:

```
classification
evidence_state
supporting_evidence
missing_evidence
safe_verification
impact
```

This creates a deterministic machine-readable boundary between inference and evaluation.

The primary scoring target is:

```
classification
```

The remaining fields provide diagnostic observability.

```
classification:
Vulnerable
```

could be accompanied by:

```
supporting_evidence:
"The identifier is sequential."
```

That exposes a potentially significant evidentiary defect:

```
Predictable Identifier
        ≠
Proven Unauthorized Access
```

The final classification alone cannot reveal whether the model arrived at the conclusion through valid or invalid reasoning.

Structured assessment therefore provides a richer **diagnostic surface**.

I deliberately avoided making another LLM the primary arbiter of the classification.

ProofSec uses canonical ground-truth labels and deterministic assertions.

The core evaluation chain is:

```
Scenario
   ↓
Model
   ↓
Structured Classification
   ↓
Canonical Ground Truth
   ↓
Exact Assertion
   ↓
Score
```

rather than:

```
Scenario
   ↓
Model A
   ↓
Model B judges Model A
   ↓
Score
```

For the core classification metric, deterministic evaluation establishes a cleaner experimental boundary.

It also reduces the possibility of **evaluation contamination**, where the judgment model introduces another layer of model-dependent interpretation into the primary metric.

A benchmark is only scientifically useful if its evaluation corpus is reproducible.

ProofSec v0.2 uses a frozen dataset representation with a SHA-256 integrity digest:

```
422501a4db424c30c8ef24b61183351ec8a4bd2096e2671cf0e6bdf91e133a80
```

The project maintains a manifest and an integrity-verification workflow.

The intended reproducibility invariant is:

```
Same benchmark version
        +
Same task corpus
        +
Same evaluation protocol
        =
Comparable experiment
```

Without corpus integrity, two apparently identical experiments can silently evaluate different datasets.

That means dataset provenance, versioning, and integrity verification are part of the experimental methodology itself.

One of the most valuable discoveries during the project was that **benchmark engineering is itself part of evaluation science**.

My initial task implementation used:

``` php
-> None
```

with assertions.

That produced pass/fail-style behavior.

I initially expected the benchmark to expose a numerical accuracy metric.

That assumption was incorrect.

I subsequently changed the task interface to:

``` php
-> float
```

and explicitly returned:

```
accuracy = correct / total
return float(accuracy)
```

This exposed a broader principle:

**The semantics of the evaluation harness are part of the experimental design.**

A benchmark can execute successfully while still reporting an invalid measurement if the scoring contract is incorrectly defined.

In other words:

```
Correct Inference
        ≠
Correct Evaluation
```

A trustworthy benchmark requires both.

One early published iteration failed with:

```
NameError: name 'ALL_TASKS_JSON' is not defined
```

This was not a model reasoning failure.

The model never reached the reasoning stage.

The task failed during execution before the first model inference step.

That forced a more rigorous separation of:

```
Dataset Integrity
        │
        ▼
Task Packaging
        │
        ▼
Runtime Execution
        │
        ▼
Model Inference
        │
        ▼
Structured Parsing
        │
        ▼
Assertion
        │
        ▼
Scoring
```

This distinction is fundamental.

A benchmark must not silently transform:

```
Runtime Error
```

into:

```
Model Error
```

Those are different failure classes with different remediation paths.

It also exposed an important deployment boundary.

My local development environment contained the canonical task source under:

```
C:\Users\User\Documents\ProofSec\kaggle\tasks
```

but a Kaggle-hosted execution environment cannot directly access that Windows filesystem.

The benchmark therefore has to package the necessary evaluation corpus and task implementation into the executable artifact available to the remote runtime.

That became an important lesson in **benchmark portability and execution-environment isolation**.

The benchmark went through multiple iterations while I validated the evaluation infrastructure.

One development version displayed assertion counts including:

```
Claude Sonnet 4.6    0 Pass
GPT-5.5             75 Pass
Gemini 3.5 Flash    95 Pass
Gemini 3.7 Flash    87 Pass
```

A subsequent controlled version successfully executed its smoke-test assertions across the evaluated models.

These development-stage measurements should not be interpreted as a universal model ranking.

They represent different stages of benchmark engineering.

During this process I was simultaneously validating:

This separation is essential for responsible interpretation.

A benchmark-development artifact is not automatically equivalent to a final scientific measurement.

Before treating Kaggle execution as authoritative, I also used local evaluation to inspect the complete 110-case corpus.

Representative local validation produced:

| Model | Correct | Total | Accuracy | 
|---|---|---|---|
| Claude Sonnet 4.6 | 80 | 110 | 72.73% | 
| Gemini 3.7 Flash | 86 | 110 | 78.18% | 
| Gemini 3.5 Flash | 84 | 110 | 76.36% | 
| GPT-5.5 | 67 | 110 | 60.91% | 

These figures are **local validation measurements under the specific execution conditions used during development**.

They should not be interpreted as immutable measurements of model capability.

LLM inference can vary because of:

This variability is itself relevant to reproducibility.

The objective was therefore not to reduce the benchmark to:

"Model X is better."

The objective was to characterize **failure modes, evidentiary behavior, and reasoning sensitivity under a controlled corpus**.

One of the recurring failure patterns motivating ProofSec can be represented as:

```
Security Vocabulary Recognition
              ↓
       Semantic Activation
              ↓
     Premature Classification
              ↓
 Evidence Is Never Actually Tested
```

The desired reasoning pathway is:

```
Scenario
   ↓
Extract Security-Relevant Claims
   ↓
Identify Actors and Authorization Boundaries
   ↓
Identify Directly Observed Facts
   ↓
Separate Indicators from Decisive Evidence
   ↓
Evaluate Contradictory Evidence
   ↓
Evaluate Negative Evidence
   ↓
Determine Evidentiary Sufficiency
   ↓
Classify
```

The distinction is fundamental.

A model can possess extensive cybersecurity knowledge while still being unreliable at **evidence-grounded adjudication**.

This is the capability ProofSec attempts to isolate.

This is one of the most consequential design decisions in ProofSec.

A conventional benchmark might ask:

```
Is the system vulnerable?

YES / NO
```

ProofSec instead represents:

```
Vulnerable
Not Vulnerable
Insufficient Evidence
```

This explicitly models **epistemic uncertainty**.

The model must distinguish between:

```
Observed
Established
```

That distinction matters operationally.

```
Unsupported Finding
        ↓
Analyst Investigation
        ↓
Engineering Interruption
        ↓
Potential Remediation
        ↓
Operational Cost
```

A system that generates large quantities of unsupported security findings can impose substantial downstream cost even when its vulnerability-recognition capability appears impressive.

Abstention is therefore not merely a failure to classify.

Under uncertainty, abstention can be a legitimate and necessary security behavior.

At a deeper level, ProofSec is not simply a cybersecurity benchmark.

It evaluates whether an LLM can maintain an explicit boundary between:

```
What do I know?
        ↓
What does the evidence establish?
        ↓
What remains unknown?
        ↓
What conclusion is justified?
```

This is fundamentally an **epistemic reasoning problem**.

Cybersecurity provides a particularly concrete environment in which to measure it because security conclusions often have direct operational consequences.

The same methodology could potentially extend to:

The security domain is simply where I chose to operationalize the problem.

ProofSec is implemented using:

`kaggle_benchmarks`
The architecture maintains explicit boundaries between:

```
DATA
PROMPT CONSTRUCTION
MODEL INFERENCE
STRUCTURED PARSING
EVALUATION
SCORING
```

This makes the system substantially easier to:

The benchmark is therefore treated as an engineered evaluation artifact rather than merely a collection of prompts.

The development workflow became:

```
Security Hypothesis
        ↓
Construct Controlled Scenario
        ↓
Define Canonical Ground Truth
        ↓
Generate Adversarial / Paired Variant
        ↓
Validate Dataset
        ↓
Freeze Corpus
        ↓
Generate Integrity Hash
        ↓
Implement Structured Task
        ↓
Run Local Validation
        ↓
Validate Kaggle Packaging
        ↓
Execute Against Multiple Models
        ↓
Inspect Assertions
        ↓
Analyze Failure Modes
        ↓
Revise Benchmark Infrastructure
```

This treats the benchmark as a versioned research artifact.

The project structure is maintained as an engineering repository containing the benchmark implementation, evaluation infrastructure, integrity controls, and supporting documentation.

The objective is to make the benchmark auditable rather than opaque.

The presence of:

```
ID
endpoint
HTTP 200
user
invoice
```

does not automatically establish IDOR.

The authorization relationship is the security property that matters.

If a single security-relevant fact changes, the model should respond to that change.

Controlled perturbation therefore provides a much stronger probe of reasoning sensitivity than simply increasing the number of unrelated benchmark questions.

A security reasoning system must be capable of recognizing evidence that weakens or falsifies its initial hypothesis.

A model that performs well under clean, monotonic evidence may behave very differently when observations conflict.

Contradiction handling is therefore a meaningful robustness dimension.

"Insufficient Evidence" can be treated as an explicit epistemic state rather than an evasive response.

A classification tells us:

```
What did the model decide?
```

A structured assessment can additionally expose:

```
What evidence did it identify?
What evidence did it consider missing?
What verification did it propose?
```

That creates a substantially richer diagnostic surface.

Dataset defects, packaging failures, schema violations, runtime exceptions, parsing failures, and model reasoning failures require separate failure categories.

A benchmark without corpus integrity, version control, and execution controls can silently drift.

A score without provenance is considerably less useful than a score attached to a reproducible experimental configuration.

ProofSec v0.2 focuses on controlled evidence reasoning.

The next iteration can extend the methodology considerably.

Instead of presenting all evidence simultaneously:

```
T0 → Initial Telemetry
T1 → Additional Observation
T2 → Contradictory Artifact
T3 → Authorization Test
```

I want to evaluate whether the model updates its hypothesis correctly as evidence arrives.

This transforms static classification into a **sequential evidence-update problem**.

Instead of evaluating only:

```
classification
```

evaluate:

```
classification
+
confidence
```

Then measure whether confidence is statistically aligned with empirical correctness.

A model that is wrong with extremely high confidence represents a fundamentally different operational risk profile from a model that identifies uncertainty.

Require the model to answer:

```
Which exact fact changed your classification?
```

This enables a deeper evaluation:

```
Correct Conclusion
        +
Correct Evidentiary Basis
```

versus:

```
Correct Conclusion
        +
Incorrect Evidentiary Basis
```

This distinction matters because a model can occasionally arrive at the correct answer for the wrong reasons.

The framework can be extended into:

```
SSRF
Authentication Bypass
Privilege Escalation
CSRF
Path Traversal
Race Conditions
Business Logic Flaws
Cloud IAM
API Authorization
Multi-Tenant Isolation
```

The broader research question becomes:

**Is evidence-grounded security reasoning a generalizable capability, or is it primarily a collection of domain-specific semantic associations?**

When I started building ProofSec, the obvious question was:

**Which LLM is better at cybersecurity?**

After engineering the benchmark, that no longer seemed like the most interesting question.

The more fundamental question is:

**Can we construct evaluations that determine whether an AI security conclusion is actually justified by the evidence available to it?**

A leaderboard such as:

```
Model A - 82%
Model B - 79%
Model C - 76%
```

is useful.

But it is incomplete.

For an AI security system, I want to know:

```
What evidence did the model use?

What evidence did it ignore?

Did it infer facts that were never provided?

Did it confuse an indicator with proof?

Did it recognize negative evidence?

Did it resolve contradictions?

Did terminology alter its classification?

Did it respond to a one-fact perturbation?

Did it revise its hypothesis?

Did it recognize when the evidence was insufficient?
```

Those questions are much closer to the requirements of a real security-analysis system.

**ProofSec v0.2** contains **110 security reasoning cases** designed around controlled evidence manipulation, structured classification, deterministic ground truth, integrity verification, and reproducible multi-model evaluation.

The benchmark is built around one principle:

**Security conclusions should be proportional to the evidence supporting them.**

ProofSec operationalizes that principle through:

The benchmark implementation and supporting engineering work are available on GitHub:

The repository is intended to make the benchmark more than a Kaggle submission.

It provides an engineering artifact around the evaluation methodology, allowing the benchmark to be inspected, versioned, reproduced, and extended independently of the Kaggle interface.

The broader objective is to preserve the chain:

```
Research Hypothesis
        ↓
Dataset Construction
        ↓
Ground Truth
        ↓
Evaluation Harness
        ↓
Integrity Verification
        ↓
Model Execution
        ↓
Results
        ↓
Failure Analysis
```

That provenance matters.

A benchmark should not only report a number.

It should make it possible to understand **where that number came from**.

The most important result from building ProofSec was not a single model score.

It was discovering how easily an evaluation can appear to measure security reasoning while actually measuring something else.

A benchmark can accidentally measure:

```
Keyword Recognition
```

instead of:

```
Evidence Reasoning
```

It can measure:

```
Runtime Correctness
Model Correctness
Confidence
Evidence-Justified Certainty
Semantic Familiarity
Causal Sensitivity
```

And it can report a precise-looking number while the underlying dataset, runtime, scoring semantics, or evaluation protocol is not actually controlled.

ProofSec is my attempt to make those boundaries explicit.

The objective is not simply to build an LLM that identifies the largest number of possible vulnerabilities.

The objective is to construct an evaluation framework capable of answering a substantially harder question:

**When an AI claims that a security vulnerability exists, did the available evidence actually justify that claim?**

That is the standard I want AI-assisted security tooling to eventually be held to.

Not:

**Did the model recognize the vulnerability vocabulary?**

But:

**Did the model establish the security claim from the evidence?**

That distinction is the entire premise of ProofSec.

**ProofSec Benchmark**

**ProofSec Source Repository**

**Kaggle Benchmarking Challenge**

[Kaggle Benchmarking Challenge on DEV](https://dev.to/challenges/kaggle-2026-09-23)
