{"slug": "sedem-selective-decompression-of-hidden-state-memories-for-long-context-question", "title": "SeDeM: Selective Decompression of Hidden-State Memories for Long-Context Question Answering", "summary": "Researchers propose SeDeM, a selective decompression framework that decouples compact memory storage from decoder conditioning for long-context question answering. On four long-context QA benchmarks, SeDeM achieves higher QA scores than evaluated compression baselines in both 1B and 3B same-backbone settings, and with the 3B backbone exceeds full-context fine-tuning on three datasets. SeDeM also reduces online time-to-first-token and improves autoregressive decoding throughput relative to ICAE.", "body_md": "arXiv:2608.00311v1 Announce Type: new\nAbstract: Long-context inference with large language models (LLMs) is costly: self-attention during prefill scales quadratically with sequence length, and the key-value (KV) cache grows with the number of processed tokens. Larger context windows also do not ensure reliable evidence use. Context compression reduces this cost, but many soft-compression methods use LLMs as compressors and rely on compact memory tokens both to preserve information and to condition the decoder. We propose SeDeM, a selective decompression framework that decouples compact memory storage from decoder conditioning. An LLM extracts hidden states from a chosen intermediate Transformer layer, a lightweight compressor stores them as memory blocks, a query-conditioned selector selects relevant blocks, and a decompressor expands only the selected blocks into hidden states compatible with an intermediate decoder layer. Thus, the decoder avoids both full-context processing and direct generation from highly compressed memory slots. On four long-context QA benchmarks, SeDeM achieves higher QA scores than the evaluated compression baselines in both 1B and 3B same-backbone settings, and with the 3B backbone exceeds full-context fine-tuning on three datasets. The learned selector uses block-level evidence supervision during training. SeDeM also reduces online time-to-first-token and improves autoregressive decoding throughput relative to ICAE.", "url": "https://wpnews.pro/news/sedem-selective-decompression-of-hidden-state-memories-for-long-context-question", "canonical_source": "https://www.machinebrief.com/news/sedem-selective-decompression-of-hidden-state-memories-for-l-ueuv", "published_at": "2026-08-04 04:00:00+00:00", "updated_at": "2026-08-04 04:37:05.800713+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models"], "entities": ["SeDeM", "ICAE"], "alternates": {"html": "https://wpnews.pro/news/sedem-selective-decompression-of-hidden-state-memories-for-long-context-question", "markdown": "https://wpnews.pro/news/sedem-selective-decompression-of-hidden-state-memories-for-long-context-question.md", "text": "https://wpnews.pro/news/sedem-selective-decompression-of-hidden-state-memories-for-long-context-question.txt", "jsonld": "https://wpnews.pro/news/sedem-selective-decompression-of-hidden-state-memories-for-long-context-question.jsonld"}}