04:00
2026-08-17
arxiv.org
artificial-intelligence
A Year in LLM Serving: Workload Evolution, Caching and Load-Balancing
A one-year production trace from Chutes, a cloud platform, reveals how LLM serving workloads evolve over time and how user-model interactions shape traffic, according to a new arXiv paper (2608.13573vā¦