LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding LOCKS, a new method from arXiv, enables efficient long-context decoding by giving each page of the key-value cache its own compact spectral summary, reconstructing within-page logits, and attending only the top pages. At a 2048-token budget, LOCKS matches FullKV aggregate quality at 100K+ context while attending about 2% of tokens, and halves per-token decode latency (2.0× at 1M tokens) against dense attention. The method ships as a drop-in plugin for unmodified vLLM with batched decode running in full CUDA graphs. arXiv:2607.24555v1 Announce Type: cross Abstract: Serving large language models at long context is bottlenecked by the key-value KV cache, which is read in full at every decode step. Attention keys are locally low-rank though globally high-rank: shared low-rank bases discard page-specific directions that a page's own compact basis retains. LOCKS gives every page its own spectral summary resident, about a tenth the cache's size , reconstructs within-page logits, estimates each page's attention mass by log-sum-exp, and attends only the top pages; selection itself reads no candidate keys or values. Selecting on this summary alone stays within about a point of the full cache on long-document QA LongBench-v1 , tracks the read-every-key oracle on retrieval-dense RULER down to the smallest budgets, and shows its largest margins on long-form reasoning AIME26, MATH-500 , where baseline selectors collapse. At its shipped $2048$-token budget LOCKS matches FullKV aggregate quality at $100$K$+$ context while attending about $2\%$ of the tokens, and halves per-token decode latency $2.0\times$ at $1$M tokens against dense attention. LOCKS ships as a drop-in plugin for unmodified vLLM, with batched decode running in full CUDA graphs.