{"slug": "archead-activation-metric-residual-correction-for-large-language-model-output", "title": "ARCHead: Activation-Metric Residual Correction for Large Language Model Output Heads", "summary": "Researchers introduced ARCHead, a packed language-modeling head compressor that reduces persistent storage by 3.7-3.9x while maintaining model quality. On Qwen3-8B-Base, ARCHead uses 25.6% of BF16 head storage with 1.007 relative perplexity, compared to 1.14-1.16 for storage-matched naive INT4. The method adds only 0.006-0.007 cross-entropy when replacing BF16 heads left by AWQ or bitsandbytes, with less than 2% throughput change.", "body_md": "arXiv:2608.02703v1 Announce Type: new\nAbstract: Weight-only quantization substantially reduces the storage of large language model (LLM) transformer blocks, but practical backends often retain the final language-modeling head (LM-head) in BF16 or FP16. Quantizing this projection naively can strongly perturb the vocabulary-logit distribution. We present ARCHead, a packed LM-head compressor that combines a quantized low-rank core, group-wise INT4 residuals, and a low-rank correction fitted in an activation-derived metric. ARCHead stores no dense BF16 head and reduces persistent LM-head storage by 3.7-3.9x. On Qwen3-8B-Base, it uses 25.6% of BF16 head storage while attaining 1.007 relative perplexity; storage-matched naive INT4 yields 1.14-1.16. Replacing the BF16 head left by AWQ or bitsandbytes adds only 0.006-0.007 cross-entropy, with less than 2% throughput change in our measurements. ARCHead therefore complements block quantizers by compressing the large output projection they can leave untouched. Code is available at https://github.com/suayptalha/archead.", "url": "https://wpnews.pro/news/archead-activation-metric-residual-correction-for-large-language-model-output", "canonical_source": "https://arxiv.org/abs/2608.02703", "published_at": "2026-08-05 04:00:00+00:00", "updated_at": "2026-08-05 04:03:45.153049+00:00", "lang": "en", "topics": ["machine-learning", "large-language-models", "ai-research"], "entities": ["ARCHead", "Qwen3-8B-Base", "AWQ", "bitsandbytes"], "alternates": {"html": "https://wpnews.pro/news/archead-activation-metric-residual-correction-for-large-language-model-output", "markdown": "https://wpnews.pro/news/archead-activation-metric-residual-correction-for-large-language-model-output.md", "text": "https://wpnews.pro/news/archead-activation-metric-residual-correction-for-large-language-model-output.txt", "jsonld": "https://wpnews.pro/news/archead-activation-metric-residual-correction-for-large-language-model-output.jsonld"}}