04:00
2026-08-25
arxiv.org
large-language-models
KVBoost: Chunk-Level Key-Value Cache Reuse with Deviation-Guided Recomputation for Efficient Large Language Model Inference
KVBoost, a chunk-level key-value cache reuse system for HuggingFace-compatible decoder models, reduces time-to-first-token by 4.49x (142.4 ms vs. 639.1 ms) on Qwen/Qwen2.5-3B over 1,000 bug-localizatiβ¦