04:00
2026-09-04
arxiv.org
artificial-intelligence
LeanStream: A Speculate-and-Refine Streaming Framework for Efficient on-Device LLM Inference
Researchers introduced LeanStream, a speculate-and-refine streaming framework for on-device large language model (LLM) inference that reduces memory usage by 4.8x to 7.5x and improves token generation…