04:00
2026-09-01
arxiv.org
large-language-models
SemKV: Semantic Mixed-Precision KV Cache Quantization Guided by the Quality Cliff for Long-Context LLM Inference
A new arXiv preprint (2608.28911v1) introduces SemKV, a semantic mixed-precision KV cache quantization method that achieves a measured 6.0x storage reduction for long-context LLM inference with no staβ¦