{"slug": "pi-mem-pushing-long-context-reasoning-to-3-6m-tokens-with-parallel-iterative", "title": "PI-Mem: Pushing Long-Context Reasoning to 3.6M Tokens with Parallel-Iterative Memory", "summary": "Researchers propose PI-Mem (Parallel-Iterative Memory), a mechanism that processes all chunks in parallel and iteratively refines a shared memory, enabling long-context reasoning up to 3.6 million tokens. On the HotpotQA benchmark, PI-Mem with Qwen3.5-35B-A3B and Qwen2.5-7B outperforms the recurrent-memory baseline by +6.25 and +7.81 absolute points while achieving 6.1x and 2.1x inference speedups, respectively, breaking the accuracy-efficiency trade-off in long-context reasoning.", "body_md": "arXiv:2608.03048v1 Announce Type: new\nAbstract: Long-context reasoning remains a critical bottleneck for large language models, as recent recurrent-memory approaches face two inherent challenges: sequential chunk-wise updates can overwrite early critical evidence with later irrelevant content, and serial inter-chunk dependencies limit parallelism and cause latency to increase with context length. To address these issues, we propose PI-Mem (Parallel-Iterative Memory), a mechanism that processes all chunks in parallel and iteratively refines a shared memory over a bounded number of turns. In each turn, PI-Mem reads all chunks in parallel conditioned on the current memory, selects new or complementary evidence from each chunk, and merges the selected evidence into a compact shared memory for the next turn. To discourage redundant turns, we optimize the workflow through reinforcement learning with an auxiliary turn-efficiency reward, enabling the model to adaptively exit once sufficient evidence has been accumulated. We evaluate PI-Mem with Qwen3.5-35B-A3B and Qwen2.5-7B on the HotpotQA benchmark across context lengths up to 3.6 million tokens and find that it outperforms the recurrent-memory baseline by +6.25 and +7.81 absolute points while achieving 6.1$\\times$ and 2.1$\\times$ inference speedups, respectively. These results demonstrate that PI-Mem breaks the accuracy--efficiency trade-off in long-context reasoning and provides a scalable approach to complex multi-hop question answering over extremely long documents.", "url": "https://wpnews.pro/news/pi-mem-pushing-long-context-reasoning-to-3-6m-tokens-with-parallel-iterative", "canonical_source": "https://www.machinebrief.com/news/pi-mem-pushing-long-context-reasoning-to-36m-tokens-with-par-7bai", "published_at": "2026-08-05 04:00:00+00:00", "updated_at": "2026-08-05 05:03:16.578231+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research"], "entities": ["PI-Mem", "Qwen3.5-35B-A3B", "Qwen2.5-7B", "HotpotQA"], "alternates": {"html": "https://wpnews.pro/news/pi-mem-pushing-long-context-reasoning-to-3-6m-tokens-with-parallel-iterative", "markdown": "https://wpnews.pro/news/pi-mem-pushing-long-context-reasoning-to-3-6m-tokens-with-parallel-iterative.md", "text": "https://wpnews.pro/news/pi-mem-pushing-long-context-reasoning-to-3-6m-tokens-with-parallel-iterative.txt", "jsonld": "https://wpnews.pro/news/pi-mem-pushing-long-context-reasoning-to-3-6m-tokens-with-parallel-iterative.jsonld"}}