An AI coding agent proposes a small refactor. The tests pass. Before approving it, a reviewer still needs to know which behavior was compared, which inputs were covered, and what happened when the change exceeded the evaluator's scope.
This is the review problem behind NAIF Agent Mutation Firewall (AMF), the second component of NAIF Enterprise Assurance. RDR provides the analysis foundation; AMF provides the mutation review gate; the Pilot Kit packages the GitHub integration and evidence.
Enterprise Access — Coming Soon. This article explains the documented approach for institutional enquiries; it does not announce general production availability.
| Decision | Review meaning |
|---|---|
| ALLOW | No observed behavior change within the admitted scope and configured input domain. |
| QUARANTINE | Review the change and any reported counterexample before proceeding. An intentional behavior-changing repair can also be quarantined under a preserve-behavior policy. |
| UNSUPPORTED | The evaluator cannot assess the change within its supported scope. Route it to another review method. |
An unsupported result must never become ALLOW just because the agent produced plausible code or the surrounding workflow finished successfully.
Consider this illustrative edit:
def classify(value):
if value >= 0:
return 1
return 0
def classify(value):
if value > 0:
return 1
return 0
For an input of zero, the two versions take different branches. A preserve-behavior review should investigate that difference rather than treating it as a cosmetic refactor. Whether the new behavior is desirable is a separate product decision.
This example explains the review question; it is not a new execution result or benchmark claim.
The public Pilot Kit uses frozen AMF v0.4 and the engine hashes referenced by RDR V2.5.7. Its default admitted scope is narrow: up to eight modified Python files, one unannotated pure positional-argument function per file, and integers from -64 to 64 plus None.
Imports, attributes, loops, classes, arbitrary calls, strings, floats and general multi-language projects fall outside that scope. Both engine observations must agree before ALLOW in the documented adapter. Syntax-derived invariant candidates are not proven invariants.
This makes the result assessable: a reviewer can see the domain and exclusions instead of reading an unqualified “safe” label.
The documented artifacts include raw engine receipts, wrapper passports, hashes, logs, timings and a minimal-in-domain counterexample when one is found. The receipt must stay associated with the analyzed base and proposed revision.
Receipts are unsigned. SHA-256 links support integrity checking; they do not authenticate the issuer. The gate does not certify arbitrary software, automatically merge a PR, or replace review of authentication, network behavior and unsupported code.
A practical institutional evaluation starts by choosing a small supported change class and defining the behavior that must remain unchanged.
Interested teams can email contact@naifgravity.com with the subject NAIF Agent Mutation Firewall — Institutional Enquiry. Include your review use case, languages, current approval workflow and a synthetic or redacted example. Do not include credentials or customer data.
Enquiries are for requirements and scope discussion. Availability will be announced separately.
Technical reference: NAIF AMF Pilot Kit documentation.
NAIF Gravity · Enterprise overview
Disclosure: Published by NAIF Gravity. AI drafted this explanation from the public pilot documentation; no new benchmark or production validation was performed for this article.