{"slug": "ground-truth-first-a-longitudinal-evaluation-instrument-for-agent-memory-and-the", "title": "Ground Truth First: A Longitudinal Evaluation Instrument for Agent Memory, and the Tenure Crossover in Memory-Architecture Rankings", "summary": "A new longitudinal evaluation instrument for LLM-agent memory, Veracium, inverts the typical benchmark pipeline by generating facts before text, embedding validity intervals and trust distinctions, and finds that memory-architecture rankings invert with history length: a budgeted curated-map memory leading at three weeks (96%) drops to 72% by nine weeks, while a provenance-typed graph rises to 90%. The open-source library, released with a corpus generator and harness, shows that write-stage quality strongly correlates with downstream performance and that a layered architecture performs best overall (96.8% short-horizon).", "body_md": "arXiv:2607.21962v1 Announce Type: new\nAbstract: Benchmarks for LLM-agent memory typically generate conversations first and extract answer keys afterwards -- with documented label-error and contamination problems -- and they overwhelmingly measure short interaction histories. We invert the pipeline: a seeded life-script sampler emits facts with validity intervals, volatility classes, and source channels before any text exists; an LLM renderer writes chat and email from per-event fact manifests; a fidelity verifier confirms every planted fact; and questions are instantiated mechanically from the script, so gold answers are script-valid by construction and separately validated for answerability. The synthetic, fictionalized corpus (~380 questions, 15 types) embeds features absent from the benchmarks we survey: per-fact validity intervals, sent/received trust distinctions, injection probes in a benign harness, and as-of-date question sets. Benchmarking five memory architectures against a no-memory control (fixed answerer, versioned LLM judge, three replicates, two horizons), we find backend rankings invert with history length: the budgeted curated-map memory that leads at three weeks loses recall of evicted content by nine weeks (96% to 72%) while a provenance-typed graph rises to 90%; the inversion is positive for all six users under complete cross-family re-judging (exact p=0.031). A full-rendered-history baseline ties or exceeds the best memory system at the short horizon but shows no judge-independent advantage at nine weeks, at about twice the read cost. Write-stage quality strongly correlates with downstream quality (weakly-written facts fail 24% vs 2%), and injection resistance tracked whether provenance boundaries survive representation. A layered architecture performs best among the memory systems in both regimes (96.8% short-horizon) and is released as Veracium, an open-source library, with the corpus generator and harness.", "url": "https://wpnews.pro/news/ground-truth-first-a-longitudinal-evaluation-instrument-for-agent-memory-and-the", "canonical_source": "https://arxiv.org/abs/2607.21962", "published_at": "2026-07-27 04:00:00+00:00", "updated_at": "2026-07-27 04:24:58.388908+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-agents", "ai-research", "ai-tools"], "entities": ["Veracium", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/ground-truth-first-a-longitudinal-evaluation-instrument-for-agent-memory-and-the", "markdown": "https://wpnews.pro/news/ground-truth-first-a-longitudinal-evaluation-instrument-for-agent-memory-and-the.md", "text": "https://wpnews.pro/news/ground-truth-first-a-longitudinal-evaluation-instrument-for-agent-memory-and-the.txt", "jsonld": "https://wpnews.pro/news/ground-truth-first-a-longitudinal-evaluation-instrument-for-agent-memory-and-the.jsonld"}}