04:00
2026-09-10
arxiv.org
large-language-models
Reasoning-Aware Compression: Identifying and Protecting Vulnerable Reasoning Circuits for Energy-Efficient LLM Deployment
A reasoning-aware compression framework that selectively restores the most quantization-sensitive circuits to FP16 achieves Pareto-optimal energy-accuracy points unreachable by uniform quantization, a…