The Copy Ceiling: An Input-Exposure Control for Ontology-Grounded Generation over Curated Corpora A study of ten language models grounded in a maintained ontology corpus raised target-name recall from 0.265 unaided to about 0.92, but a verbatim copy baseline scored 0.964, leaving every model 0.022 to 0.067 below it, according to arXiv paper 2609.24885v2. The authors report exposure accounting and a model-judged audit of 423 sampled item observations, plus a separate paired production study showing a model-judged quality gain of +0.27 [+0.11, +0.45] on a 0-5 scale. They recommend reporting exposure accounting alongside quality judgements, not in place of them, and note that rephrasing questions out of the graph's vocabulary cut exposure from 0.964 to 0.328 while the absence-keyed fallback would have fired on only 2 of 506. arXiv:2609.24885v2 Announce Type: replace Abstract: We built a node that grounds a replaceable language model in a maintained ontology corpus, then asked what its successful-looking evaluation could support. Across ten models, grounding raised target-name recall from 0.265 unaided to about 0.92. A copy baseline, the recall a verbatim copy of the shown context already achieves, scores 0.964, and every model sits 0.022 to 0.067 below it. Copying therefore scores higher on this limited recall measure, which does not assess whether answers are better. The comparison tests what a recall score establishes; it does not test whether reasoning occurred, because a reasoned answer and a copy score alike when the answer name is already in context. We report exposure accounting four counts classifying each gold item by whether the context exposed it and the answer recovered it and a model-judged audit of 423 sampled item observations. A separate paired production study found a model-judged quality gain of +0.27 +0.11, +0.45 on a 0-5 scale. Operational studies found failures that recall alone would not show: rephrasing questions out of the graph's vocabulary cut exposure from 0.964 to 0.328, yet the absence-keyed fallback would have fired on only 2 of 506; and inserting extracted facts degraded judged pages in every arm, so that step was disabled. Five-arm controls show that any well-formed on-corpus block beats no context but do not establish that the specific content matters, and no matched comparison against flat-text retrieval was run. The corpus is public and largely LLM-generated, which establishes neither training exposure nor novelty. Each study has its own outcome measure. Where gold derives from the injected corpus, we recommend reporting the accounting beside quality judgements, not in place of them.