cd /news/ai-safety/how-would-you-prove-an-ai-security-f… · home › topics › ai-safety › article
[ARTICLE · art-142020] src=dev.to ↗ pub= topic=ai-safety verified=true sentiment=· neutral

How Would You Prove an AI Security Fix Actually Works?

A developer has proposed an experiment design for evaluating whether AI-proposed security patches actually work, using intentionally vulnerable code-repair exercises grouped into original reproduction, related-case, and legitimate-behavior tests. The design calls for explicit access contracts, separate first-attempt and post-feedback results, and identical tool access when comparing models, while explicitly claiming no completed runs or benchmark results. The developer is building Breachloom, which includes code-repair exercises that motivate the proposed evaluation.

by read2 min views3 publishedSep 29, 2026

An assistant proposes a security patch. Its explanation sounds reasonable. The original attack no longer works.

Do you merge it?

I want to evaluate that decision with a small experiment built around intentionally vulnerable exercises. This post describes the experiment design. It does not report completed runs or claim a failure rate for any model.

I would begin with:

Each exercise needs an explicit contract. Who should have access? Which inputs are legitimate? What should happen when a resource does not exist?

Without those details, the assistant and the reviewer may be solving different problems.

I would prepare three test groups:

Group Question
Original reproduction Does the reported misuse still succeed?
Related cases Does the same weakness survive with different data or conditions?
Legitimate behavior Can intended users still use the feature?
For a SQL query, an ordinary product name containing an apostrophe is a useful legitimate input. Rejecting it may hide a query-construction bug while breaking a valid requirement.

For object authorization, use at least two users and more than one protected object. A hard-coded exception must not pass as a general fix.

For file access, the tests must match the actual path-decoding, normalization and filesystem behavior. A toy string check cannot establish the safety of a production file server.

For every run, save:

Keep first-attempt results separate from results after feedback. If comparing models, use the same cases and access to tools, and report the limited scope. A few examples cannot support a broad ranking of model security competence.

The most interesting part would be showing a patch first and asking: “Approve or request changes?”

Then reveal the test evidence. Include patches that succeed. If the assistant handles every case well, that is the result to publish.

My hypothesis is that reviewing the contract and test coverage will teach more than counting whether a single demonstration stopped working. The experiment should be allowed to challenge that hypothesis too.

I am building Breachloom, which includes code-repair exercises. Those exercises motivate this proposed evaluation; no AI benchmark results are being claimed here.

Which legitimate behavior would you include to catch an over-restrictive “security fix”?

── more in #ai-safety 4 stories · sorted by recency
── more on @breachloom 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/how-would-you-prove-…] indexed:0 read:2min 2026-09-29 · —