22:00
2026-07-31
systems.seas.harvard.edu
large-language-models
Rethinking Bursty Workloads and KV Cache Hierarchies for Efficient LLM Serving
Akira van de Groenendaal, a senior studying Computer Science at Carnegie Mellon University, presented research at the Harvard Systems Group showing that bursty workloads can improve the time-per-outpu…