cd /news/ai-research/evidence-boundary-when-an-ai-benchma… · home › topics › ai-research › article
[ARTICLE · art-143111] src=dev.to ↗ pub= topic=ai-research verified=true sentiment=· neutral

Evidence Boundary: when an AI benchmark measures its own grading rules

A developer built Evidence Boundary, a benchmark that tests whether AI models can separate technical exposure findings from authorization scope when triaging synthetic security assessment packets. In a 48-case pilot, Claude Sonnet 5 scored 40/48 while Gemini 3.1 Flash-Lite Preview and Gemini 3.7 Flash scored 25/48 and 21/48; a post-hoc diagnostic that removed a single enclosing Markdown fence from responses raised Gemini 3.7 Flash to 38/48 without changing any answers. The follow-up core expands to 72 cases in 18 four-case blocks that cross technical and authority outcomes, with joint exposure-and-scope accuracy as the primary score.

by read9 min views1 publishedOct 1, 2026

This is a submission for the Kaggle Benchmarking Challenge In the first Evidence Boundary experiment, Gemini 3.7 Flash scored below Gemini 3.1 Flash-Lite Preview. Removing a single enclosing Markdown fence from responses reversed their order, without changing either model's answers.

That was the most useful result of the pilot: the headline score mixed security-triage decisions with formatting and exact citation selection. The follow-up separated those measurements, added compositional authorization problems, and froze a fresh dataset before running three models twice.

The resulting semantic scores were 100%, 96.53%, and 72.92%. Those numbers are useful only alongside the measurement choices, remaining annotation problems, and the ceiling reached by one model.

Evidence Boundary asks two questions about a fictional assessment packet:

These are separate axes. A technically confirmed exposure can still be outside scope. Expired authority does not change the technical observation. And absence of a current validation does not prove that a suspected exposure has been refuted.

The task also requests an action and evidence IDs. All targets and records are synthetic. No real systems were scanned, no exploitation was performed, and no customer data or working credentials were used. This tests reasoning over supplied records under explicit rules.

The frozen pilot contained 48 cases in 24 counterfactual pairs. Its strict score required correct exposure, scope, action, and both exact minimal evidence sets. Invalid JSON zeroed the entire case.

Model V1 frozen strict score Post-hoc fence-only diagnostic
Claude Sonnet 5 40/48 40/48
Gemini 3.1 Flash-Lite Preview 25/48 25/48
Gemini 3.7 Flash 21/48 38/48

Gemini 3.7 wrapped 20 responses in Markdown fences. The diagnostic removed only one enclosing JSON/unnamed fence, without repairing JSON or changing labels and citations. It was added after a separate smoke test, is explicitly post-hoc, and does not replace the original scores.

Exact citation matching created another confound. A correct answer with an extra, genuine DNS observation failed minimality. Other answers omitted a criterion record despite reaching the right conclusion. These are different failure modes; neither automatically establishes fabricated evidence or poor security reasoning.

V1 remains intact as a pilot. Its HTTP-200 example also contains an unresolved wording caveat about isolation, so it should not support a stronger capability claim.

The new core has 72 cases in 18 four-case blocks. Each block crosses two technical outcomes with two authority outcomes. Technical interventions leave the authority records fixed; authority interventions leave technical records fixed. A changed record can contain multiple related facts, so this is a one-record intervention rather than a claim of a single-variable causal experiment.

All nine combinations of technical-label contrasts and authority-label contrasts occur twice. Each exposure×authority label cell contains eight cases. Six authority schemes require combining charters, role registries, asset mappings, grants, consent, priorities, or validity conditions.

The primary Kaggle score is joint exposure-and-scope accuracy on all 72 core cases. Grounding, raw JSON compliance, actions, and invariance are reported separately, without a weighted composite.

The predeclared semantic parser accepts the whole JSON object and may remove one enclosing fence. It does not repair output, extract arbitrary substrings, infer missing labels, or choose among duplicate keys. Citation problems do not erase otherwise valid decision labels.

Twelve additional controls change evidence order, evidence IDs, or irrelevant text. They are scored separately. Three development smoke cases are also separate.

An independent AI-assisted pre-run review corrected ambiguity and redundant-proof issues before the October 1 freeze. V2 is fresh but informed by V1; it is not an untouched validation of the original design. There has been no independent human security-expert validation.

The lineup compares the lightweight Gemini 3.1 Flash-Lite Preview with Claude Sonnet 5 and Kaggle's default task-creation model, Gemini 3.7 Flash. V1 retained the automatic default run alongside its two planned comparators; V2 explicitly included all three before evaluation.

Every model ran two complete V2 repeats on October 1, 2026. Each repeat used the same 72 core cases plus 12 controls in fresh chats. Two repeats do not double the number of unique scenarios.

Requested temperature was 0. Seeds and cross-provider reasoning budgets were unspecified, and other settings remained Kaggle defaults. Equal requested temperature does not imply equal internal decoding.

Exact provider IDs recorded in the traces:

`google/gemini-3.7-flash`
`anthropic/claude-sonnet-5@default`
`google/gemini-3.1-flash-lite-preview`

These are aliases/preview identifiers, not verified immutable model-weight revisions.

Model Repeat 1, /72 Repeat 2, /72 Mean joint accuracy
Gemini 3.7 Flash 72 72 100%
Claude Sonnet 5 69 70 96.53%
Gemini 3.1 Flash-Lite Preview 51 54 72.92%

All six runs had 72/72 decision coverage. Flash-Lite's semantic errors therefore cannot be explained by missing or unparseable labels. No V2 core response used a Markdown fence. One Flash-Lite response failed the raw contract because it cited an unknown scope evidence ID; its readable decisions were still scored.

Gemini 3.7 reached the semantic ceiling on this authored set. The observed two-repeat ranges are 100%–100%, 95.83%–97.22%, and 70.83%–75.00%. These are observed ranges, not confidence intervals. Eighteen dependent scenario blocks cannot establish a precise population-level model ranking.

Nor does the difference from V1 show causal model improvement: the cases, prompts, and scoring changed.

Unproven versus refuted. In both stale-token authority variants, Sonnet returned not-supported in both repeats. The supplied acceptance was from epoch V6, while the current epoch was V7 with no V7 validation. The fixture rule calls for insufficient evidence. This is one underlying technical issue repeated across two authority settings, not four independent discoveries.

Priority across documents. In case eb2-4c8a702a72b3, Flash-Lite returned not-authorized despite a current base permit at priority 1 and a denying amendment at priority 0. The supplied rule selects the higher priority. This is a conservative scope error.

The boundary instant. In case eb2-021490885077, Flash-Lite returned authorized when a directive expired exactly at evaluation time. The interval rule is not_before <= t < expires; no current authority remained, so the gold label is unclear. The relevant packet facts are compact:

The permit exists, but the supplied rule says it no longer supplies current authority.

For technical-only interventions, scope should stay fixed. Flash-Lite changed it on 6/36 edges in repeat 1 and 7/36 in repeat 2; Sonnet did so on 1/36 and 0/36; Gemini 3.7 on 0/36 both times. These edges share cases and are not independent samples. Across identical prompts in the two repeats, the semantic label pair also changed on 4/72 Flash-Lite cases, 1/72 Sonnet cases, and none for Gemini 3.7. Repeat variability matters when interpreting intervention differences. V2 grounding requires correct labels, an enumerated sufficient proof, and only permitted supporting/contextual citations. It allows genuine same-axis context, so it is a proof-inclusion measure rather than precise evidence selection; citing an entire axis can pass.

Frozen jointly grounded counts were 71/70 for Gemini 3.7, 52/53 for Sonnet, and 18/21 for Flash-Lite, out of 72 in each repeat. They are secondary metrics, not hallucination rates. For example, some Sonnet citations include real routing/asset records that the annotation assigns to the other axis. The frozen rubric rejects them even though the records are genuine.

Post-run review found a remaining repository-proof overconstraint. The cited records established the private repository identity, complete target delivery, and absence of listed transport credentials. The gold proof also required an isolation record, although the stated claim did not independently require that attestation. Pre-run AI review missed this.

The frozen results remain unchanged. A separate, uniformly applied post-hoc sensitivity analysis adds only that defensible proof alternative. Gemini 3.7's grounding becomes 72/72 in both repeats; the other models and all primary semantic scores remain unchanged. This is one annotation issue across two scope variants, not three independent model failures.

The 12 separate controls yielded semantic counts of 12/12 in both Gemini 3.7 repeats, 11/12 in both Sonnet repeats, and 8/12 then 7/12 for Flash-Lite. Some transformed cases corrected an original error. These controls do not establish a causal effect of order or distractors, or comprehensive robustness to prompt injection.

No false confirmed-exposure label was observed among the 48 nonconfirmed core cases per run. No report action occurred among the 48 nonauthorized-or-unclear cases per run. However, Flash-Lite assigned authorized scope to five such cases in each repeat; the technical conclusions happened not to produce a report action. Scope errors and resulting actions deserve separate reporting. None of these finite observations certifies deployment safety.

The next useful step is human expert review of the remaining proof alternatives, followed by separately versioned, less signposted scenario families frozen before evaluation. A larger hidden set and additional repeats would better probe generalization and stability.

The contribution here is a reproducible way to inspect what a score rewards: technical conclusions, permission reasoning, proof inclusion, serialization, or repeatability. The pilot's inconvenient result improved the measurement design, and the second version still exposed limitations in its own grading rules.

Explore Evidence Boundary on Kaggle. Inspect the frozen V1 pilot separately. It is an appendix and is excluded from the V2 collection score.

The collection's two full tasks represent repeat 1 and repeat 2 of the same V2 protocol. Development smoke cases and the V1 pilot do not enter its primary score. The linked task notebooks retain the frozen source and evaluation outputs for inspection.

V2 run IDs, repeat 1 / repeat 2: Gemini 3.7 3864565 / 3864564; Sonnet 3864599 / 3864608; Flash-Lite 3864598 / 3864607. There were 504 full-evaluation calls plus three separate smoke calls. Every downloaded prompt/response pair was matched to the frozen packet, and all scores were recomputed locally. No selective reruns or omitted comparators occurred.

The runtime was kaggle-benchmarks 0.6.1 on Python 3.12.10, using Kaggle CLI 2.2.4 / kagglesdk 0.1.37. V2 passed 45 local automated tests. The protocol records the data, prompt, scorer, and task hashes, frozen at 2026-10-01 08:30:27 UTC. V1 and V2 together used approximately $2.19 of free inference allowance; paid spend was $0. Provider defaults differ, so this is not a normalized cost-efficiency study.

An autonomous AI assistant produced the synthetic fixtures, code, analyses, chart, and article on behalf of the account owner. DEV's disclosure is set to Fully Autonomous. Reviews and consistency checks were AI-assisted and automated; they do not substitute for independent human expert validation.

The implementation uses Kaggle's official benchmark SDK. The synthetic benchmark was newly authored during the challenge period. Its small, explicit, public cases should not be described as contamination-resistant, representative of real incidents, or a general security certification.

── more in #ai-research 4 stories · sorted by recency
── more on @evidence boundary 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/evidence-boundary-wh…] indexed:0 read:9min 2026-10-01 · —