LSREP: A Longitudinal State-Replay Protocol for Evaluating Conversational Memory, with ICE v2 as an Audited Local-First Architecture A new arXiv paper (2609.16730v1) introduces LSREP, a Longitudinal State-Replay Evaluation Protocol for conversational memory, with ICE v2 as its audited local-first architecture case study. The private single-user instantiation of ICE v2 contains 1,985 turns, 219 distinct probes, and 1,211 probe-checkpoint observations across 52 checkpoints, and on three ordinary-density datasets ICE v2 shows a near-zero mean quality difference from vector-RAG while selecting 32% fewer fragments but using 6.6% more estimated prompt tokens. On the matched public LongMemEval diagnostic, ICE v2 loses to pure vector-RAG at 50.8% versus 72.8% in the evidence-only oracle and 43.0% versus 69.5% in full-S, with paired differences of -22.0 points (95% CI [-26.6, -17.4]) and -26.5 ([-31.3, -21.8]). arXiv:2609.16730v1 Announce Type: new Abstract: Conversational memory changes during use, so endpoint question answering alone cannot establish how a persistent state accumulates, ages, or incorporates revisions. We introduce LSREP, a Longitudinal State-Replay Evaluation Protocol combining ordered replay, explicit lifecycle schedules, repeated probes, evolving reference answers, and mechanism-fidelity checks. Its architectural case study is ICE v2, a local-first memory middleware with typed stores, retrieval fusion, and dynamic context budgets. The private, single-user instantiation contains 1,985 turns, 219 distinct probes, and 1,211 probe-checkpoint observations across 52 checkpoints. On three ordinary-density datasets, ICE v2 has a near-zero mean quality difference from vector-RAG while selecting 32% fewer fragments but using 6.6% more estimated prompt tokens. A fourth, dense dataset exposes catastrophic failures of the unbudgeted baseline. The fidelity audit limits attribution: procedural retrieval is defective, several mechanisms are unexercised, and graph utility is not established. In a complementary matched public diagnostic, ICE v2 loses decisively to pure vector-RAG on LongMemEval: 50.8% versus 72.8% in the evidence-only oracle and 43.0% versus 69.5% in full-S. Paired differences are -22.0 points 95% CI -26.6, -17.4 and -26.5 -31.3, -21.8 . Conservative abstention accompanies severe multi-session and temporal failures. ICE uses less context in this diagnostic, establishing a quality-cost trade-off rather than superior efficiency. Together, replay, fidelity auditing, and public endpoint testing expose distinct failure modes that neither architectural descriptions nor aggregate scores identify alone.