16:29
2026-10-01
dev.to
artificial-intelligence
The Inference Auction: Why Bidding for GPU Priority Breaks KV Cache Locality
A paper from UC Berkeley, TTIC, and Google Research authored by Keegan Harris, Siddharth Prasad, Asher Trockman, Nika Haghtalab, and Michael I. Jordan shows that sorting an LLM inference queue strictlβ¦