cd /news/large-language-models/semkv-semantic-mixed-precision-kv-ca… · home topics large-language-models article
[ARTICLE · art-117336] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

SemKV: Semantic Mixed-Precision KV Cache Quantization Guided by the Quality Cliff for Long-Context LLM Inference

A new arXiv preprint (2608.28911v1) introduces SemKV, a semantic mixed-precision KV cache quantization method that achieves a measured 6.0x storage reduction for long-context LLM inference with no statistically detectable quality loss, and up to 7.9x when combined with a distortion-optimized quantizer. The researchers identified a 'quality cliff' between 2.0 and 2.322 code bits per value for uniform quantization, below which performance collapses, and SemKV avoids this by assigning two adjacent above-cliff precisions based on model-internal token importance scores.

read1 min views2 publishedSep 1, 2026

arXiv:2608.28911v1 Announce Type: new Abstract: The key-value (KV) cache is the dominant memory bottleneck of long-context large language model (LLM) inference, growing linearly with context length. We show that uniform KV quantization on a fractional-bit grid does not degrade gracefully: under a prespecified multi-seed statistical protocol, Llama-3.1-8B-Instruct with an affine quantizer is statistically indistinguishable from FP16 KV down to 2.322 code bits/value and collapses at 2.0 bits - a quality cliff in (2.0, 2.322] that reappears in generation-time quantization and multi-turn dialogue and transfers to Mistral-7B. The cliff reframes importance-aware mixed precision: above it, eight model-internal importance indicators are statistically interchangeable, so the benefit of mixing is grid interpolation, reaching average precisions uniform quantization cannot realize. SemKV preserves every token, ranks tokens by a model-internal score, and assigns two adjacent above-cliff precisions, achieving a measured 6.0x storage reduction with no statistically detectable quality difference from full KV (n=900, three seeds), and outperforming FP16 token pruning granted a 1.5x larger memory budget. Replacing the affine base with a distortion-optimized quantizer (TurboQuant-MSE) lowers the cliff in every protocol tested, raising the no-detectable-loss operating point to 7.9x. The recipe: measure the cliff for the target deployment setting, then interpolate above it.

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/semkv-semantic-mixed…] indexed:0 read:1min 2026-09-01 ·