{"slug": "how-would-you-prove-an-ai-security-fix-actually-works", "title": "How Would You Prove an AI Security Fix Actually Works?", "summary": "A developer has proposed an experiment design for evaluating whether AI-proposed security patches actually work, using intentionally vulnerable code-repair exercises grouped into original reproduction, related-case, and legitimate-behavior tests. The design calls for explicit access contracts, separate first-attempt and post-feedback results, and identical tool access when comparing models, while explicitly claiming no completed runs or benchmark results. The developer is building Breachloom, which includes code-repair exercises that motivate the proposed evaluation.", "body_md": "An assistant proposes a security patch. Its explanation sounds reasonable. The original attack no longer works.\n\nDo you merge it?\n\nI want to evaluate that decision with a small experiment built around intentionally vulnerable exercises. This post describes the experiment design. It does not report completed runs or claim a failure rate for any model.\n\nI would begin with:\n\nEach exercise needs an explicit contract. Who should have access? Which inputs are legitimate? What should happen when a resource does not exist?\n\nWithout those details, the assistant and the reviewer may be solving different problems.\n\nI would prepare three test groups:\n\n| Group | Question | \n|---|---|\n| Original reproduction | Does the reported misuse still succeed? | \n| Related cases | Does the same weakness survive with different data or conditions? | \n| Legitimate behavior | Can intended users still use the feature? | \n\nFor a SQL query, an ordinary product name containing an apostrophe is a useful legitimate input. Rejecting it may hide a query-construction bug while breaking a valid requirement.\n\nFor object authorization, use at least two users and more than one protected object. A hard-coded exception must not pass as a general fix.\n\nFor file access, the tests must match the actual path-decoding, normalization and filesystem behavior. A toy string check cannot establish the safety of a production file server.\n\nFor every run, save:\n\nKeep first-attempt results separate from results after feedback. If comparing models, use the same cases and access to tools, and report the limited scope. A few examples cannot support a broad ranking of model security competence.\n\nThe most interesting part would be showing a patch first and asking: “Approve or request changes?”\n\nThen reveal the test evidence. Include patches that succeed. If the assistant handles every case well, that is the result to publish.\n\nMy hypothesis is that reviewing the contract and test coverage will teach more than counting whether a single demonstration stopped working. The experiment should be allowed to challenge that hypothesis too.\n\nI am building [Breachloom](https://breachloom.com/?lang=en&utm_source=devto&utm_medium=article&utm_campaign=ai_patch_review), which includes code-repair exercises. Those exercises motivate this proposed evaluation; no AI benchmark results are being claimed here.\n\n**Which legitimate behavior would you include to catch an over-restrictive “security fix”?**", "url": "https://wpnews.pro/news/how-would-you-prove-an-ai-security-fix-actually-works", "canonical_source": "https://dev.to/ookeolioli222/how-would-you-prove-an-ai-security-fix-actually-works-1fm2", "published_at": "2026-09-29 20:12:14+00:00", "updated_at": "2026-09-29 20:17:03.706443+00:00", "lang": "en", "topics": ["ai-safety", "ai-agents", "developer-tools", "ai-tools"], "entities": ["Breachloom"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/how-would-you-prove-an-ai-security-fix-actually-works", "markdown": "https://wpnews.pro/news/how-would-you-prove-an-ai-security-fix-actually-works.md", "text": "https://wpnews.pro/news/how-would-you-prove-an-ai-security-fix-actually-works.txt", "jsonld": "https://wpnews.pro/news/how-would-you-prove-an-ai-security-fix-actually-works.jsonld"}}