04:00
2026-09-30
arxiv.org
large-language-models
HeadGuard: Selective Head Protection for Low-Bit VLM KV-Cache Quantization
HeadGuard, a composable head-protection method for low-bit vision-language model KV-cache quantization, recovers a substantial fraction of lost accuracy by keeping roughly 1/8 of physical KV heads in …