An assistant proposes a security patch. Its explanation sounds reasonable. The original attack no longer works.
Do you merge it?
I want to evaluate that decision with a small experiment built around intentionally vulnerable exercises. This post describes the experiment design. It does not report completed runs or claim a failure rate for any model.
I would begin with:
Each exercise needs an explicit contract. Who should have access? Which inputs are legitimate? What should happen when a resource does not exist?
Without those details, the assistant and the reviewer may be solving different problems.
I would prepare three test groups:
| Group | Question |
|---|---|
| Original reproduction | Does the reported misuse still succeed? |
| Related cases | Does the same weakness survive with different data or conditions? |
| Legitimate behavior | Can intended users still use the feature? |
For a SQL query, an ordinary product name containing an apostrophe is a useful legitimate input. Rejecting it may hide a query-construction bug while breaking a valid requirement.
For object authorization, use at least two users and more than one protected object. A hard-coded exception must not pass as a general fix.
For file access, the tests must match the actual path-decoding, normalization and filesystem behavior. A toy string check cannot establish the safety of a production file server.
For every run, save:
Keep first-attempt results separate from results after feedback. If comparing models, use the same cases and access to tools, and report the limited scope. A few examples cannot support a broad ranking of model security competence.
The most interesting part would be showing a patch first and asking: “Approve or request changes?”
Then reveal the test evidence. Include patches that succeed. If the assistant handles every case well, that is the result to publish.
My hypothesis is that reviewing the contract and test coverage will teach more than counting whether a single demonstration stopped working. The experiment should be allowed to challenge that hypothesis too.
I am building Breachloom, which includes code-repair exercises. Those exercises motivate this proposed evaluation; no AI benchmark results are being claimed here.
Which legitimate behavior would you include to catch an over-restrictive “security fix”?