A Completeness Gate Is Only as Complete as Its Required Set. Don't Let the Model Write It. A developer argues that completeness gates conflate two distinct jobs — enumerating the required set and comparing an artifact against it — and that letting the same model write the required set undermines the check. Two cited papers support the split: AbsenceBench found 14 mid-2025 models scored 17.3–40.0 percent F1 at detecting deleted lines in GitHub pull requests, and 'Judging Is Not Enumerating' found models judge membership far better than they author the reference sets, with code models judging at F1 0.74–0.90 writing test suites that admitted only 19–42 percent of correct solutions. A completeness gate has two different problems: determining what must exist, and determining whether it exists. We keep treating them as one. One common design hands both to the same model. A generated change plan, policy bundle or control mapping goes in, the model is asked whether anything required is missing, and it answers against its own idea of what "required" means. If it reports nothing missing, the gate passes. Nothing alerts, because nothing failed. The check ran correctly against a list that was already short. The first problem is enumeration: which elements must be present? The second is comparison: is each of them present in this artifact? Any gate that checks a deployment plan against a control catalogue, or a generated policy against the clauses a standard requires, does both, whether or not the design separates them. Wherever model-generated plans, policies or change requests pass through a gate like this, it belongs to AI infrastructure architecture https://www.rack2cloud.com/ai-infrastructure-strategy-guide/ as much as to the pipeline that hosts it. Comparison is a membership question, and it can only report on elements the required set contains. Enumeration decides the ceiling. A flawless comparison against an incomplete set returns a clean result for an incomplete artifact. | Job | Question | How it fails | |---|---|---| | Enumerate the required set | What must exist? | Required elements are missing from the set, so the gate under-rejects. Requirements the specification never states are added, so it over-rejects. | | Compare artifact to set | Is each required element present? | A required element is in the set and absent from the artifact, and the comparison does not report it. | The two enumeration failures behave differently under review. An invented requirement is written down, so a reviewer can challenge it. A missing requirement is an absence, and finding it is the same problem the author failed to solve. The research below argues this asymmetry, and it is why a model-written list can survive inspection and still be short. A check that reports success while the objective went unmet already has a name on this site, False Completion 82 https://www.rack2cloud.com/ai-placement-latency-cost-tradeoff/ , and the cost of continuing to verify automated output is the subject of the Automation Validation Tax 172 https://www.rack2cloud.com/automation-validation-tax/ . This post is narrower than either: it is about how the set a gate checks against gets built. How an approval step decays into a rubber stamp is covered in GhostApproval https://www.rack2cloud.com/ai-approval-integrity/ . Two pieces of research bear on how a completeness gate should be built, and neither is about infrastructure. AbsenceBench https://arxiv.org/abs/2506.11440 gives a model an original document and an edited copy and asks which elements were removed. On its GitHub pull-request portion, 14 models scored between 17.3 and 40.0 percent F1 across their configurations, with the best at 40.0, even though the authors note that a simple program solves the task in linear time. The models are mid-2025 generation, the omissions are literal deleted lines, and the paper reports no error bars. Judging Is Not Enumerating https://arxiv.org/abs/2608.01000 measures models as authors of the sets that other systems check against: test suites, answer keys, rubrics. It tests four separate reference constructions, and in each the models judged whether a candidate belongs far better than they authored the set itself. The size of the gap depends on the construction. On the paper's algorithmic construction it was about 0.29 to 0.34 F1 and did not close across a 24-fold range of model sizes. On executable code, models that judged at F1 0.74 to 0.90 wrote test suites that admitted only 19 to 42 percent of the correct solutions. In a separate test, planted over-inclusions were detected six to seven times more often than planted omissions. A control then located the deficit: asked to emit the predicate rather than list its members, the same models reached an F1 of about 0.99. In that control the rule was already in the prompt, so the model was restating a rule, not deriving one. The authors scope the finding to one-shot authoring without test-time reasoning. Enabling reasoning on two closed frontier models closed or nearly closed the gap on the simplest, rule-based construction. It did not reliably repair test-suite authoring, where the expected behavior has to be derived from prose. The paper's one production observation is a K-12 assessment pipeline of 43,227 scored items, where errors in model-authored answer keys were omissions rather than inclusions by about ten to one, as graded by a commercial evaluator. ⚠ What this evidence does not show: It does not show that current models cannot do this: reasoning narrowed or closed the gap where a rule could be executed. It does not test semantic absence, because AbsenceBench removes literal lines. And it contains no infrastructure incident. Applying it to plans, policies and change requests is an architectural extrapolation from benchmark and research results. The completeness gate architecture that follows separates the two jobs and gives each to the component that does it reliably. The model may translate a rule that has already been specified into an executable predicate. Code then runs that predicate against an authoritative, machine-readable source to enumerate the required set. A deterministic diff compares the artifact to that set, and the gate reports the missing elements by name. | Step | Who does it | Why | |---|---|---| | Specify the rule | A person or a standard | The rule exists before the gate does | | Translate it to a predicate | The model, with the result reviewed | Restating a given rule is the task the research control shows models doing well | | Enumerate the required set | Code, from the source | The result is reproducible and adds no probabilistic step | | Compare artifact to set | A deterministic diff | Set difference is a mechanical operation | | Report | The gate | Each missing element is named | Specify, translate, enumerate, diff: the model touches only the second step. It participates in expressing the rule. It does not become the database of what must exist. When the predicate and the source are correct, deterministic enumeration gives a reproducible result without introducing another probabilistic judgment step. That is the whole argument, and it is about cost and verifiability, not about what models are capable of. A model can still propose predicates, and a reasoning model may enumerate acceptably in simple cases. The design should not make model enumeration the authoritative source of the required set when the set can be derived deterministically from an authoritative machine-readable source. Infrastructure already uses versions of this separation: policy-as-code evaluates machine-readable Terraform plans https://www.rack2cloud.com/deterministic-iac-terraform-policy-as-code/ against explicit rules rather than asking a model to invent the required checks at evaluation time. Four design properties keep the split intact in any completeness gate. The required set must be computable from something that exists outside the model: an inventory, a schema, a control catalogue, a plan diff. If no such source exists, there is nothing for code to enumerate, and the gate cannot claim completeness. The model's output is a predicate a person can read, test and version. A wrong predicate is visible text, while an omission from a model-written list is not. That visibility is this post's inference from the research finding that over-inclusions are easier to catch than omissions. Enumeration runs as code against the source and produces the same set on every run. Where a reasoning model enumerates instead, treat its output as a proposal to diff against the code-derived set, not as the set itself. A predicate that encodes a requirement the specification never stated blocks valid work. Trace every predicate to a clause in the specification, and review rejections as seriously as passes. A completeness gate built this way works where the required set can be derived from a machine-readable source. Plenty of requirements cannot be derived that way. "The runbook should cover failure of the identity provider" and "the change should consider downstream consumers" have no computable predicate, and nothing here says how to gate them. Nor does the evidence: AbsenceBench removes literal lines, and the research controls use rules a model was given. For requirements stated only in prose, a model-derived set may be the only practical option. In that case the gate is no longer proving completeness against a deterministic required set. It is performing a model-mediated review, its omissions stay hard to see, and its output should say so, so that a pass is read as weaker evidence. Two neighboring questions are out of scope here. Who is entitled to reject a gate's conclusion is a separate authority question, covered for recovery claims in Recovery Evidence Boundary https://www.rack2cloud.com/recovery-evidence-boundary/ and for AI platform evidence in AI evidence verification https://www.rack2cloud.com/ai-evidence-verification/ ; it does not depend on how the required set was built. And a probe reports on what it probes, which is the limit What Kubernetes Health Checks Actually Guarantee https://www.rack2cloud.com/kubernetes-health-checks-actually-guarantee/ covers for readiness and liveness checks. A completeness gate is itself a system that can be wrong, so test it with artifacts whose contents you already know. Run a known-complete artifact through and confirm it passes. Run one with a deliberately removed element and confirm the gate names it. If the gate cannot find a planted omission, its clean results carry no information. A fixed answer set defined before the run is the same move AI workload RPO https://www.rack2cloud.com/ai-workload-rpo/ makes for restore tests. The research points the same direction. Its one mitigation that worked for authored code test suites was discarding any suite that rejected a known-correct probe, which cut false rejection from 58 to 92 percent down to 5 percent or less, at the cost of keeping only 5 to 39 percent of suites. A planted-omission test is the mirror image, and it is cheap to run on every change to the predicate or the source. A completeness gate can only report on what its required set contains. If a model wrote that set in one pass, the gate inherits the model's omissions and returns clean results for incomplete artifacts. The mistake is treating enumeration and comparison as one job. The research shows that judging membership and enumerating requirements are different tasks, and that model performance can diverge sharply between them. Where a rule exists, let the model translate the rule and let code do the listing. Reasoning narrows the gap in simple cases and does not reliably close it elsewhere, so the design should not depend on it. Where the required set can be derived from an authoritative source, derive it. Where it cannot, say that the gate checked a model-derived list. Either way, a gate that has never found a planted omission has not shown it can find a real one. A completeness gate is only as complete as the set it checks against. Originally published at rack2cloud.com https://www.rack2cloud.com/completeness-gate-required-set/