04:00
2026-08-27
arxiv.org
artificial-intelligence
FAMPWQ: Fisher Information-based Adaptive Mixed Precision Weight Quantization for Effective LLM Inference
Researchers proposed FAMPWQ, a Fisher information-based adaptive mixed precision weight quantization method for efficient LLM inference on commodity GPUs, achieving up to 3.39 lower perplexity, 6.87% …