cd /news/artificial-intelligence/ground-truth-first-a-longitudinal-ev… · home topics artificial-intelligence article
[ARTICLE · art-74904] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Ground Truth First: A Longitudinal Evaluation Instrument for Agent Memory, and the Tenure Crossover in Memory-Architecture Rankings

A new longitudinal evaluation instrument for LLM-agent memory, Veracium, inverts the typical benchmark pipeline by generating facts before text, embedding validity intervals and trust distinctions, and finds that memory-architecture rankings invert with history length: a budgeted curated-map memory leading at three weeks (96%) drops to 72% by nine weeks, while a provenance-typed graph rises to 90%. The open-source library, released with a corpus generator and harness, shows that write-stage quality strongly correlates with downstream performance and that a layered architecture performs best overall (96.8% short-horizon).

read1 min views1 publishedJul 27, 2026

arXiv:2607.21962v1 Announce Type: new Abstract: Benchmarks for LLM-agent memory typically generate conversations first and extract answer keys afterwards -- with documented label-error and contamination problems -- and they overwhelmingly measure short interaction histories. We invert the pipeline: a seeded life-script sampler emits facts with validity intervals, volatility classes, and source channels before any text exists; an LLM renderer writes chat and email from per-event fact manifests; a fidelity verifier confirms every planted fact; and questions are instantiated mechanically from the script, so gold answers are script-valid by construction and separately validated for answerability. The synthetic, fictionalized corpus (~380 questions, 15 types) embeds features absent from the benchmarks we survey: per-fact validity intervals, sent/received trust distinctions, injection probes in a benign harness, and as-of-date question sets. Benchmarking five memory architectures against a no-memory control (fixed answerer, versioned LLM judge, three replicates, two horizons), we find backend rankings invert with history length: the budgeted curated-map memory that leads at three weeks loses recall of evicted content by nine weeks (96% to 72%) while a provenance-typed graph rises to 90%; the inversion is positive for all six users under complete cross-family re-judging (exact p=0.031). A full-rendered-history baseline ties or exceeds the best memory system at the short horizon but shows no judge-independent advantage at nine weeks, at about twice the read cost. Write-stage quality strongly correlates with downstream quality (weakly-written facts fail 24% vs 2%), and injection resistance tracked whether provenance boundaries survive representation. A layered architecture performs best among the memory systems in both regimes (96.8% short-horizon) and is released as Veracium, an open-source library, with the corpus generator and harness.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @veracium 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ground-truth-first-a…] indexed:0 read:1min 2026-07-27 ·