{"slug": "lsrep-a-longitudinal-state-replay-protocol-for-evaluating-conversational-memory", "title": "LSREP: A Longitudinal State-Replay Protocol for Evaluating Conversational Memory, with ICE v2 as an Audited Local-First Architecture", "summary": "A new arXiv paper (2609.16730v1) introduces LSREP, a Longitudinal State-Replay Evaluation Protocol for conversational memory, with ICE v2 as its audited local-first architecture case study. The private single-user instantiation of ICE v2 contains 1,985 turns, 219 distinct probes, and 1,211 probe-checkpoint observations across 52 checkpoints, and on three ordinary-density datasets ICE v2 shows a near-zero mean quality difference from vector-RAG while selecting 32% fewer fragments but using 6.6% more estimated prompt tokens. On the matched public LongMemEval diagnostic, ICE v2 loses to pure vector-RAG at 50.8% versus 72.8% in the evidence-only oracle and 43.0% versus 69.5% in full-S, with paired differences of -22.0 points (95% CI [-26.6, -17.4]) and -26.5 ([-31.3, -21.8]).", "body_md": "arXiv:2609.16730v1 Announce Type: new \nAbstract: Conversational memory changes during use, so endpoint question answering alone cannot establish how a persistent state accumulates, ages, or incorporates revisions. We introduce LSREP, a Longitudinal State-Replay Evaluation Protocol combining ordered replay, explicit lifecycle schedules, repeated probes, evolving reference answers, and mechanism-fidelity checks. Its architectural case study is ICE v2, a local-first memory middleware with typed stores, retrieval fusion, and dynamic context budgets. The private, single-user instantiation contains 1,985 turns, 219 distinct probes, and 1,211 probe-checkpoint observations across 52 checkpoints. On three ordinary-density datasets, ICE v2 has a near-zero mean quality difference from vector-RAG while selecting 32% fewer fragments but using 6.6% more estimated prompt tokens. A fourth, dense dataset exposes catastrophic failures of the unbudgeted baseline. The fidelity audit limits attribution: procedural retrieval is defective, several mechanisms are unexercised, and graph utility is not established. In a complementary matched public diagnostic, ICE v2 loses decisively to pure vector-RAG on LongMemEval: 50.8% versus 72.8% in the evidence-only oracle and 43.0% versus 69.5% in full-S. Paired differences are -22.0 points (95% CI [-26.6, -17.4]) and -26.5 ([-31.3, -21.8]). Conservative abstention accompanies severe multi-session and temporal failures. ICE uses less context in this diagnostic, establishing a quality-cost trade-off rather than superior efficiency. Together, replay, fidelity auditing, and public endpoint testing expose distinct failure modes that neither architectural descriptions nor aggregate scores identify alone.", "url": "https://wpnews.pro/news/lsrep-a-longitudinal-state-replay-protocol-for-evaluating-conversational-memory", "canonical_source": "https://www.machinebrief.com/news/lsrep-a-longitudinal-state-replay-protocol-for-evaluating-co-rfhm", "published_at": "2026-09-16 04:00:00+00:00", "updated_at": "2026-09-16 04:36:14.159650+00:00", "lang": "en", "topics": ["ai-research", "large-language-models", "ai-safety"], "entities": ["LSREP", "ICE v2", "LongMemEval", "vector-RAG", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/lsrep-a-longitudinal-state-replay-protocol-for-evaluating-conversational-memory", "markdown": "https://wpnews.pro/news/lsrep-a-longitudinal-state-replay-protocol-for-evaluating-conversational-memory.md", "text": "https://wpnews.pro/news/lsrep-a-longitudinal-state-replay-protocol-for-evaluating-conversational-memory.txt", "jsonld": "https://wpnews.pro/news/lsrep-a-longitudinal-state-replay-protocol-for-evaluating-conversational-memory.jsonld"}}