04:00
2026-08-05
arxiv.org
artificial-intelligence
Output-Aware Rotation for INT2 KV-Cache Quantization
Researchers introduced OptR, an output-aware rotation method for INT2 KV-cache quantization that minimizes post-W_O attention-output error, improving long-context LLM inference efficiency. Across threβ¦