{"slug": "rethinking-bursty-workloads-and-kv-cache-hierarchies-for-efficient-llm-serving", "title": "Rethinking Bursty Workloads and KV Cache Hierarchies for Efficient LLM Serving", "summary": "Akira van de Groenendaal, a senior studying Computer Science at Carnegie Mellon University, presented research at the Harvard Systems Group showing that bursty workloads can improve the time-per-output-token of LLM clusters and that controlling the private vs shared split of a distributed KV cache can optimize time-to-first-token, challenging conventional systems choices in LLM serving.", "body_md": "# Rethinking Bursty Workloads and KV Cache Hierarchies for Efficient LLM Serving\n\n## Abstract\n\nModern LLM inference workloads are bursty and exhibit heavy prefix reuse, and serving them efficiently requires understanding how we can turn these traits to our advantage. In this talk, I’ll present two recent projects: the first investigates how latency is impacted by bursty arrivals, and the second explores how we can make best use of a distributed KV cache. First, we’ll see how burstiness can improve the time-per-output-token of your LLM cluster and discuss the conditions which make this possible, as well as what it implies for request routing. Afterwards, we’ll look at how to control the private vs shared split of your KV cache to optimize time-to-first-token. Both projects show how intuitive systems choices can sometimes leave performance gains on the table, motivating the need to question seemingly obvious decisions.\n\n## Bio\n\nAkira van de Groenendaal is a senior studying Computer Science at Carnegie Mellon University, currently doing research with the Harvard Systems Group under Prof. Juncheng Yang. He works on ML systems, with interests in workload analysis and the queueing dynamics of LLM inference.", "url": "https://wpnews.pro/news/rethinking-bursty-workloads-and-kv-cache-hierarchies-for-efficient-llm-serving", "canonical_source": "https://systems.seas.harvard.edu/seminar/2026-08-03-akira-van-de-groenendaal/", "published_at": "2026-07-31 22:00:00+00:00", "updated_at": "2026-07-31 22:39:49.127482+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "ai-research"], "entities": ["Akira van de Groenendaal", "Carnegie Mellon University", "Harvard Systems Group", "Juncheng Yang"], "alternates": {"html": "https://wpnews.pro/news/rethinking-bursty-workloads-and-kv-cache-hierarchies-for-efficient-llm-serving", "markdown": "https://wpnews.pro/news/rethinking-bursty-workloads-and-kv-cache-hierarchies-for-efficient-llm-serving.md", "text": "https://wpnews.pro/news/rethinking-bursty-workloads-and-kv-cache-hierarchies-for-efficient-llm-serving.txt", "jsonld": "https://wpnews.pro/news/rethinking-bursty-workloads-and-kv-cache-hierarchies-for-efficient-llm-serving.jsonld"}}