Replication Without Persistence in Hosted LLMs: Measurement Sensitivity in Action-Time Belief Evaluation A study of hosted language models found that a previously reported Gemini 3.1 Flash-Lite deficit in the Regent Chess sequential environment recurs on fresh games under its historical configuration at +0.0530 (95% CI [+0.0329,+0.0714]), but the model-minus-uniform endpoint drops 0.0429 under a rebuilt configuration sharing the same public identifier (95% CI for the H-minus-R contrast [+0.0182,+0.0667]). Under the rebuilt configuration, the prospectively frozen interleaved same-window 4K comparison reverses sign between Gemini 3.1 and Gemini 3.7, identifiers that differ in release and product tier, while any additional serving-period contribution remains unresolved at -0.0166 ([-0.0483,+0.0157]). The authors conclude that replication, measurement sensitivity, and persistence can yield different conclusions within a single evaluation, motivating explicit indexing of hosted-model behavioural claims by tested identifier, serving period, measurement instrument, and inference configuration. arXiv:2609.22478v1 Announce Type: new Abstract: Behavioural evaluations of hosted language models can vary because the evaluated service, the measurement instrument, or both differ across runs. We separate three validation questions: whether a prior finding recurs on fresh data under its historical configuration replication , whether the endpoint changes when the evaluation-and-inference configuration is rebuilt under the same identifier measurement sensitivity , and whether the finding persists across subsequently tested identifiers under one common instrument persistence . We study these questions in Regent Chess, a sequential environment in which a hidden, mutable state is recorded exactly, allowing stated beliefs to be scored against ground truth at action time; positive endpoint values mean worse performance than a matched-uniform comparator. The previously reported Gemini 3.1 Flash-Lite deficit recurs on fresh games under its historical configuration +0.0530, 95% CI +0.0329,+0.0714 . In a back-to-back same-day H/R comparison under the same public identifier, the model-minus-uniform endpoint is 0.0429 lower under the rebuilt configuration 95% CI for the H-minus-R contrast +0.0182,+0.0667 ; all six configuration components vary jointly, so no component is isolated. Under rebuilt R, the prospectively frozen, interleaved same-window 4K comparison reverses sign between Gemini 3.1 and Gemini 3.7, identifiers that differ in release and product tier; additional descriptive and exploratory cells show the same directional pattern. Any additional serving-period contribution remains unresolved -0.0166, -0.0483,+0.0157 . Replication, measurement sensitivity, and persistence can therefore yield different conclusions within one evaluation, motivating explicit indexing of hosted-model behavioural claims by tested identifier, serving period, measurement instrument, and inference configuration.