{"slug": "codec-gauge-learning-compression-friendly-gauges-for-transformer-kv-caches", "title": "Codec-Gauge: Learning Compression-Friendly Gauges for Transformer KV Caches", "summary": "A new post-training method called Codec-Gauge, introduced in arXiv:2607.20538v1, learns small orthogonal channel transforms for Transformer KV caches to improve compression fidelity. Across six models at 3, 4, and 6 bits per value, Codec-Gauge reduces zfp KL divergence by 44.0% on average compared to raw coordinates, outperforming random, Hadamard, DCT, and PCA/KLT controls without changing model weights or attention semantics.", "body_md": "arXiv:2607.20538v1 Announce Type: new\nAbstract: Long-context Transformer inference increasingly relies on KV-cache compression or quantization. Prior rotation and transform-coding results suggest that the channel basis of each key/value vector affects how faithfully a fixed backend preserves model behavior. We introduce Codec-Gauge, a post-training cache-coordinate layer that learns small orthogonal channel transforms around existing compression and quantization backends. Its frequency-distribution objective combines a token-channel DCT spectral-centroid loss with a smooth rate proxy to concentrate KV energy in low-frequency codec-facing layouts. We evaluate actual compression and decompression using measured bytes and rolling compressed-history scoring. Across six models at $3$, $4$, and $6$ bits/value, learned gauges reduce zfp KL divergence by $44.0\\%$ on average relative to raw coordinates and outperform random, Hadamard, DCT, and PCA/KLT controls. The same gauges improve quality preservation for block-uniform and KIVI-style quantization. Experiments on a 27B model and long-context task prompts reproduce the quality trend, while serial storage and timing measurements validate the implemented compressed-cache paths. These results establish cache-coordinate geometry as a practical post-training variable for improving compression fidelity without changing model weights, attention semantics, or backend coding rules.", "url": "https://wpnews.pro/news/codec-gauge-learning-compression-friendly-gauges-for-transformer-kv-caches", "canonical_source": "https://arxiv.org/abs/2607.20538", "published_at": "2026-07-24 04:00:00+00:00", "updated_at": "2026-07-24 04:12:41.813279+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research"], "entities": ["Codec-Gauge", "arXiv", "KIVI", "zfp"], "alternates": {"html": "https://wpnews.pro/news/codec-gauge-learning-compression-friendly-gauges-for-transformer-kv-caches", "markdown": "https://wpnews.pro/news/codec-gauge-learning-compression-friendly-gauges-for-transformer-kv-caches.md", "text": "https://wpnews.pro/news/codec-gauge-learning-compression-friendly-gauges-for-transformer-kv-caches.txt", "jsonld": "https://wpnews.pro/news/codec-gauge-learning-compression-friendly-gauges-for-transformer-kv-caches.jsonld"}}