G^2PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation A new method called G^2PTQ improves post-training quantization (PTQ) for large language models by generalizing gradient compensation, addressing two complementary limitations in GPTQ-based methods that have become the de facto standard for reducing LLM memory and computational footprint without retraining. The approach targets the local, layer-wise objective used by existing GPTQ-based techniques. Post-training quantization PTQ is a practical approach to reducing the memory and computational footprint of large language models LLMs without retraining. GPTQ-based methods have become the de facto standard, yet they suffer from two complementary limitations. Methods with local, layer-wise obj