{"slug": "fampwq-fisher-information-based-adaptive-mixed-precision-weight-quantization-for", "title": "FAMPWQ: Fisher Information-based Adaptive Mixed Precision Weight Quantization for Effective LLM Inference", "summary": "Researchers proposed FAMPWQ, a Fisher information-based adaptive mixed precision weight quantization method for efficient LLM inference on commodity GPUs, achieving up to 3.39 lower perplexity, 6.87% higher accuracy, and a 76% win rate in LLM-as-a-judge comparisons across 7 models and 5 benchmarks, outperforming 7 baseline approaches.", "body_md": "arXiv:2608.24945v1 Announce Type: new\nAbstract: Recent years have witnessed remarkable achievements of Large Language Models (LLMs) in multiple domains, while the excessive resource requirements of LLMs hinder the deployment on resource-constrained devices. Although model quantization stands out as an effective approach, conventional quantization approaches typically incur severe performance degradation due to uniform bit-width or simple heuristic sensitivity evaluation. In this paper, we propose a novel Fisher information-based Adaptive Mixed Precision Weight Quantization approach, i.e., FAMPWQ, which performs layer-adaptive weight quantization for effective LLM inference on commodity GPUs. First, we propose a system model with a novel Fisher information metric to measure the layer-wise sensitivity to quantization. Second, we propose a reinforcement learning-based bit-width allocator in FAMPWQ, which generates an adaptive bit-width allocation strategy based on the Fisher information sensitivity metric. Extensive experiments on 7 models and 5 benchmarks demonstrate that FAMPWQ significantly outperforms 7 baseline approaches in terms of PPL (up to 3.39 smaller), accuracy (up to 6.87% higher), and LLM-as-a-judge comparison (up to 76% win rate).", "url": "https://wpnews.pro/news/fampwq-fisher-information-based-adaptive-mixed-precision-weight-quantization-for", "canonical_source": "https://arxiv.org/abs/2608.24945", "published_at": "2026-08-27 04:00:00+00:00", "updated_at": "2026-08-27 04:18:56.269792+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "machine-learning", "ai-research"], "entities": ["FAMPWQ", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/fampwq-fisher-information-based-adaptive-mixed-precision-weight-quantization-for", "markdown": "https://wpnews.pro/news/fampwq-fisher-information-based-adaptive-mixed-precision-weight-quantization-for.md", "text": "https://wpnews.pro/news/fampwq-fisher-information-based-adaptive-mixed-precision-weight-quantization-for.txt", "jsonld": "https://wpnews.pro/news/fampwq-fisher-information-based-adaptive-mixed-precision-weight-quantization-for.jsonld"}}