Unlocking Lossless Speedups in LLMs via Discrete Diffusion (5000 Tk/S) Researchers from the Institute of Foundation Models, Cornell Tech, Cerebras Systems, and other institutions introduced a discrete diffusion method that achieves lossless speedups in large language models (LLMs), reaching 5,000 tokens per second. The approach enables faster inference without sacrificing output quality, marking a significant advance in LLM efficiency. Unlocking Lossless Speedups in LLMs via Discrete Diffusion Subham Sekhar Sahoo †,1, Lingjie Chen †,1,2, Khiem Pham †,1,3, Jonathan Geuter †,1,4, Junlin Chen 1,5, Chaitanya Dwivedi https://scholar.google.com/citations?user=ghpn6JkAAAAJ&hl=en 1, Varad Pimpalkhute https://nightlessbaron.github.io 1, Yash Akhauri https://akhauriyash.github.io 1, Alexander Moreno https://www.linkedin.com/in/alexander-moreno-ab151542/ 1, Mikhail Yurochkin https://moonfolk.github.io 1, Zhenting Wang https://zhentingwang.github.io 1, Mostafa Elhoushi https://scholar.google.com/citations?user=y cwSKAAAAAJ&hl=en 6, Nolan Dey https://scholar.google.com/citations?user=JHUfMr0AAAAJ&hl=en 6, Shane Bergsma https://sites.google.com/site/shaneabergsma/ 6, Joel Hestness https://scholar.google.com/citations?user=wkbvCf0AAAAJ&hl=en 6, Hongyi Wang https://hwang595.github.io 1,5, John Thickstun https://johnthickstun.com 3, Eric Xing https://mbzuai.ac.ae/study/faculty/professor-eric-xing/ 1, Zhengzhong Liu https://hunterhector.github.io 1 1Institute of Foundation Models 2University of Illinois Urbana-Champaign 3Cornell Tech 4Harvard University 5Rutgers University 6Cerebras Systems †Core contributors