{"slug": "growpage-on-demand-kv-budgeting-for-efficient-llm-reasoning-serving", "title": "GrowPage: On-Demand KV Budgeting for Efficient LLM Reasoning Serving", "summary": "Researchers introduced GrowPage, an on-demand KV budgeting framework for efficient LLM reasoning serving, which dynamically adjusts key-value cache capacity per request based on attention demand. GrowPage uses dual-timescale query summaries to estimate demand and compresses or acquires physical pages at capacity boundaries, achieving a superior performance-throughput trade-off across multiple reasoning benchmarks.", "body_md": "arXiv:2609.03494v1 Announce Type: new\nAbstract: Long-output reasoning has made the key--value (KV) cache a critical memory bottleneck for efficient LLM serving. Existing KV compression methods usually rely on a predefined per-request budget and adjust only which KV states are retained, leaving the total capacity fixed throughout decoding. However, reasoning workloads exhibit substantial demand variation: different requests require different KV capacities, and the attention demand of an individual request evolves during generation. We introduce \\textbf{GrowPage}, an on-demand KV budgeting framework that treats KV capacity as a runtime resource. GrowPage maintains lightweight dual-timescale query summaries to capture recent and long-term attention behaviors, and uses their relative attention working sets to estimate demand evolution. At each capacity boundary, GrowPage either compresses KV states within the current allocation or acquires an additional physical page when broader demand emerges. By integrating with PagedAttention's page-level memory abstraction, GrowPage preserves continuous batching and prefix caching. Experiments on reasoning benchmarks across multiple models show that GrowPage achieves a superior performance--throughput trade-off over existing approaches.", "url": "https://wpnews.pro/news/growpage-on-demand-kv-budgeting-for-efficient-llm-reasoning-serving", "canonical_source": "https://arxiv.org/abs/2609.03494", "published_at": "2026-09-04 04:00:00+00:00", "updated_at": "2026-09-04 04:24:07.799741+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "ai-infrastructure"], "entities": ["GrowPage", "PagedAttention"], "alternates": {"html": "https://wpnews.pro/news/growpage-on-demand-kv-budgeting-for-efficient-llm-reasoning-serving", "markdown": "https://wpnews.pro/news/growpage-on-demand-kv-budgeting-for-efficient-llm-reasoning-serving.md", "text": "https://wpnews.pro/news/growpage-on-demand-kv-budgeting-for-efficient-llm-reasoning-serving.txt", "jsonld": "https://wpnews.pro/news/growpage-on-demand-kv-budgeting-for-efficient-llm-reasoning-serving.jsonld"}}