cd /news/ai-research/lsrep-a-longitudinal-state-replay-pr… · home topics ai-research article
[ARTICLE · art-131037] src=machinebrief.com ↗ pub= topic=ai-research verified=true sentiment=· neutral

LSREP: A Longitudinal State-Replay Protocol for Evaluating Conversational Memory, with ICE v2 as an Audited Local-First Architecture

A new arXiv paper (2609.16730v1) introduces LSREP, a Longitudinal State-Replay Evaluation Protocol for conversational memory, with ICE v2 as its audited local-first architecture case study. The private single-user instantiation of ICE v2 contains 1,985 turns, 219 distinct probes, and 1,211 probe-checkpoint observations across 52 checkpoints, and on three ordinary-density datasets ICE v2 shows a near-zero mean quality difference from vector-RAG while selecting 32% fewer fragments but using 6.6% more estimated prompt tokens. On the matched public LongMemEval diagnostic, ICE v2 loses to pure vector-RAG at 50.8% versus 72.8% in the evidence-only oracle and 43.0% versus 69.5% in full-S, with paired differences of -22.0 points (95% CI [-26.6, -17.4]) and -26.5 ([-31.3, -21.8]).

by read1 min views1 publishedSep 16, 2026

arXiv:2609.16730v1 Announce Type: new Abstract: Conversational memory changes during use, so endpoint question answering alone cannot establish how a persistent state accumulates, ages, or incorporates revisions. We introduce LSREP, a Longitudinal State-Replay Evaluation Protocol combining ordered replay, explicit lifecycle schedules, repeated probes, evolving reference answers, and mechanism-fidelity checks. Its architectural case study is ICE v2, a local-first memory middleware with typed stores, retrieval fusion, and dynamic context budgets. The private, single-user instantiation contains 1,985 turns, 219 distinct probes, and 1,211 probe-checkpoint observations across 52 checkpoints. On three ordinary-density datasets, ICE v2 has a near-zero mean quality difference from vector-RAG while selecting 32% fewer fragments but using 6.6% more estimated prompt tokens. A fourth, dense dataset exposes catastrophic failures of the unbudgeted baseline. The fidelity audit limits attribution: procedural retrieval is defective, several mechanisms are unexercised, and graph utility is not established. In a complementary matched public diagnostic, ICE v2 loses decisively to pure vector-RAG on LongMemEval: 50.8% versus 72.8% in the evidence-only oracle and 43.0% versus 69.5% in full-S. Paired differences are -22.0 points (95% CI [-26.6, -17.4]) and -26.5 ([-31.3, -21.8]). Conservative abstention accompanies severe multi-session and temporal failures. ICE uses less context in this diagnostic, establishing a quality-cost trade-off rather than superior efficiency. Together, replay, fidelity auditing, and public endpoint testing expose distinct failure modes that neither architectural descriptions nor aggregate scores identify alone.

── more in #ai-research 4 stories · sorted by recency
── more on @lsrep 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/lsrep-a-longitudinal…] indexed:0 read:1min 2026-09-16 ·