{"slug": "greedy-decoding-is-not-precision-invariant-cross-precision-output-divergence", "title": "Greedy Decoding Is Not Precision-Invariant: Cross-Precision Output Divergence", "summary": "A September 22, 2026 arXiv paper reports that greedy decoding in large language models is not precision-invariant: the same model, prompt and decoding algorithm produced different outputs in BF16 versus FP16 on identical hardware, with 49-100% of prompts diverging across six models (1.1B-7B parameters, four families; divergence additionally characterized at 12B) and three benchmarks. The authors' error-propagation analysis found that 22 layers of accumulated body error do not distinguish flipping from non-flipping steps, with the outcome depending primarily on the top-two logit margin at the LM head relative to the directional perturbation between the top-two candidates. Their best low-overhead intervention, selective FP32 LM head recomputation triggered when the margin falls below a threshold, delivered +22-36 percentage points exact agreement on A10G (+12-21 pp on L4 and A100) at under 4% latency overhead in low-batch (batch size <=4) single-stream inference, but the benefit vanishes at batch size >=8 and under end-to-end FP8.", "body_md": "# Computer Science > Machine Learning\n\n  [Submitted on 22 Sep 2026]\n\n# Title:Greedy Decoding Is Not Precision-Invariant: Cross-Precision Output Divergence in LLM Inference\n\n[View PDF](https://arxiv.org/pdf/2609.26621)\n\nAbstract:Greedy decoding from large language models is commonly treated as deterministic. We show it is not precision-invariant: the same model, prompt, and decoding algorithm produce different outputs in BF16 versus FP16 on identical hardware. Across our evaluations of six models (1.1B-7B parameters, four families; divergence additionally characterised at 12B) and three benchmarks, 49-100\\% of prompts diverge; a single token flip often cascades into trajectory-level divergence. We develop an empirical error-propagation analysis and find that 22 layers of accumulated body error do not distinguish flipping from non-flipping steps; the outcome depends primarily on the top-two logit margin at the LM head relative to the directional perturbation between the top-two candidates. The analysis makes five testable predictions about intervention outcomes, including that applying more FP32 compute (broader scope) makes agreement worse. The experiments match all five predictions. The best-performing low-overhead intervention we evaluate, selective FP32 LM head recomputation, triggered only when the margin falls below a threshold, delivers +22-36 pp exact agreement on A10G (+12-21 pp on L4 and A100) at less than 4\\% latency overhead in low-batch (batch size <=4) single-stream inference. We map the applicability boundary across six models and four batch sizes, and hypothesise that training-time precision stability is a determining factor. The method is a partial mitigation rather than a universal determinism guarantee: its benefit vanishes when body-originated error dominates, including at batch size >=8 and under end-to-end FP8 in our tests.\n    \n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer \n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers \n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps \n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations \n\n*(*[What are Smart Citations?](https://www.scite.ai/))\n# Code, Data and Media Associated with this Article\n\nalphaXiv \n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers \n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub \n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub \n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face \n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast \n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))\n# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower \n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender \n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))\nIArxiv Recommender\n\n*(*[What is IArxiv?](https://iarxiv.org/about))\n# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/greedy-decoding-is-not-precision-invariant-cross-precision-output-divergence", "canonical_source": "https://arxiv.org/abs/2609.26621", "published_at": "2026-09-24 02:07:08+00:00", "updated_at": "2026-09-24 02:28:25.815126+00:00", "lang": "en", "topics": ["large-language-models", "machine-learning", "ai-research", "ai-infrastructure"], "entities": ["arXiv", "A10G", "L4", "A100", "BF16", "FP16", "FP32", "FP8"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/greedy-decoding-is-not-precision-invariant-cross-precision-output-divergence", "markdown": "https://wpnews.pro/news/greedy-decoding-is-not-precision-invariant-cross-precision-output-divergence.md", "text": "https://wpnews.pro/news/greedy-decoding-is-not-precision-invariant-cross-precision-output-divergence.txt", "jsonld": "https://wpnews.pro/news/greedy-decoding-is-not-precision-invariant-cross-precision-output-divergence.jsonld"}}