04:00
2026-09-04
arxiv.org
large-language-models
GrowPage: On-Demand KV Budgeting for Efficient LLM Reasoning Serving
Researchers introduced GrowPage, an on-demand KV budgeting framework for efficient LLM reasoning serving, which dynamically adjusts key-value cache capacity per request based on attention demand. Growβ¦