{"slug": "accelerating-large-language-model-decoding-with-speculative-sampling", "title": "Accelerating Large Language Model Decoding with Speculative Sampling", "summary": "Researchers introduced speculative sampling, an algorithm that accelerates transformer decoding by generating multiple tokens per call using a faster draft model and modified rejection sampling, achieving a 2-2.5x decoding speedup on Chinchilla, a 70 billion parameter language model, without compromising sample quality.", "body_md": "# Computer Science > Computation and Language\n\n  [Submitted on 2 Feb 2023]\n\n# Title:Accelerating Large Language Model Decoding with Speculative Sampling\n\n[View PDF](/pdf/2302.01318)\n\n[HTML (experimental)](https://arxiv.org/html/2302.01318v1)\n\nAbstract:We present speculative sampling, an algorithm for accelerating transformer decoding by enabling the generation of multiple tokens from each transformer call. Our algorithm relies on the observation that the latency of parallel scoring of short continuations, generated by a faster but less powerful draft model, is comparable to that of sampling a single token from the larger target model. This is combined with a novel modified rejection sampling scheme which preserves the distribution of the target model within hardware numerics. We benchmark speculative sampling with Chinchilla, a 70 billion parameter language model, achieving a 2-2.5x decoding speedup in a distributed setup, without compromising the sample quality or making modifications to the model itself.\n    \n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer \n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers \n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps \n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations \n\n*(*[What are Smart Citations?](https://www.scite.ai/))\n# Code, Data and Media Associated with this Article\n\nalphaXiv \n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers \n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub \n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub \n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face \n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast \n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))\n# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower \n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender \n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))\n# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/accelerating-large-language-model-decoding-with-speculative-sampling", "canonical_source": "https://arxiv.org/abs/2302.01318", "published_at": "2026-09-07 09:00:00+00:00", "updated_at": "2026-09-07 21:30:25.622473+00:00", "lang": "en", "topics": ["large-language-models", "artificial-intelligence"], "entities": ["Chinchilla"], "alternates": {"html": "https://wpnews.pro/news/accelerating-large-language-model-decoding-with-speculative-sampling", "markdown": "https://wpnews.pro/news/accelerating-large-language-model-decoding-with-speculative-sampling.md", "text": "https://wpnews.pro/news/accelerating-large-language-model-decoding-with-speculative-sampling.txt", "jsonld": "https://wpnews.pro/news/accelerating-large-language-model-decoding-with-speculative-sampling.jsonld"}}