arXiv:2610.09111v1 Announce Type: new Abstract: Deterministic inference is essential for reliable and trustworthy machine learning. Prior studies of text generation have shown that changing factors such as batch size, batch composition, hardware, or inference engine can alter the generated text, even when the prompt, model parameters, and sampling randomness are fixed. These differences have been attributed in part to floating-point non-associativity, shape-dependent kernel selection, and other implementation-level differences in numerical execution. However, it remains unclear whether, when, and to what extent the same factors affect text classification. We present a systematic study of serving-context non-invariance in text classifiers, which prior work has measured only through generated text. We train 180 models spanning discriminative, pseudo-generative, and fully generative classifier formulations and evaluate each across four categories of serving contexts, holding the checkpoint and the text fixed. Label stability does not imply score stability. Changing only the batch shape changes no labels across fp32 comparisons, yet under bf16 it moves up to 56.7 percentage points of predicted probability mass, with label changes concentrated at small margins. Fully generative classifiers change more labels than their discriminative counterparts under the same serving changes. We derive sufficient conditions for label stability under each serving change and give a separate mitigation for each mechanism. Our results identify and quantify the serving conditions that must be fixed for reproducible text classification.
Same Text, Different Prediction: Serving-Context Nondeterminism in Text Classifiers
A systematic study of 180 text-classifier models found that serving-context changes alone can shift up to 56.7 percentage points of predicted probability mass under bf16, even when the checkpoint and input text are held fixed, according to arXiv paper 2610.09111v1. The authors report that changing only batch shape altered no labels in fp32 comparisons but did change labels under bf16, with label changes concentrated at small decision margins, and that fully generative classifiers changed more labels than discriminative ones under the same serving changes. The paper derives sufficient conditions for label stability under each serving change and gives a separate mitigation for each mechanism, identifying the serving conditions that must be fixed for reproducible text classification.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.