Long-term conversational memory is the ability to recall and reason over months of past conversations. It is one of the hardest unsolved problems in AI assistants, and the field has converged on a shared assumption about how to solve it: the model cannot be trusted to find what it needs, so the system has to pre-digest the conversation for it. Extract atomic facts. Build entity graphs. Summarize each session. Page memory in and out of a custom memory-management layer.
We think that assumption is worth testing, and we have the numbers to argue against it.
We put raw conversation turns into a single Postgres table, exposed three search tools over it via MCP, and let the agent decide how to look things up. No fact extraction. No summarization. No graphs. No memory manager. That system scores F1 = 0.665 on the LoCoMo benchmark with Claude Sonnet (raw, full 10 samples), above Omni-SimpleMem's best reported configuration at 0.613. With the much smaller Claude Haiku it reaches F1 = 0.638.
The interesting part is not the score. It is what the score implies. Every one of those pre-digestion layers exists to compensate for retrieval the model was assumed to be incapable of doing. Once the model can search for itself, they stop paying for themselves. We tested the most popular one directly, and removing LLM-extracted atomic facts did not lower F1 while cutting roughly 50% from ingestion time.
The complexity did not disappear. It moved. It moved out of the ingestion pipeline and into the search tool interface. That is the single idea this post is about, and the sharpest evidence for it is a result we did not expect: swapping a simple one-parameter search tool for a nine-parameter one, with the same database underneath, cost 0.09 F1. The interface is not packaging. It is the system.
There is a second idea running alongside it, which took us longer to see. Adversarial rejection is the binding constraint. Nearly every change that made the agent better at answering made it worse at correctly refusing to answer, and the papers that exclude adversarial questions from scoring have hidden that tradeoff rather than solved it.
By agentic search we mean something specific: the agent chooses its retrieval strategy at query time. Which modes to search in, how to combine them, what to search for next given what the last search returned, when to stop. Nothing about the strategy is fixed in advance. A fixed pipeline decides all of that once, at design time, for every question.
What everyone else builds
The published LoCoMo systems share a family resemblance. Omni-SimpleMem uses pyramid expansion, LLM summarization, BM25 hybrid retrieval, and adaptive top-k, with the whole arrangement discovered through automated architecture search. MemGPT implements a memory-management OS with paging. A-MEM builds associative memory graphs. Others extract entities, maintain summaries, or run dedicated consolidation passes.
Each layer is a reasonable response to the same constraint: a fixed retrieval step gets one shot at finding the right context, so the corpus has to be pre-shaped to make that one shot count. Extraction, summarization, and graph-building are all ways of doing retrieval work in advance, because you cannot do it at query time.
That constraint is now optional.
What we built instead
Three components, one table.
The memory table
CREATE TABLE memory ( id uuid PRIMARY KEY, content text NOT NULL, meta jsonb NOT NULL, tree ltree NOT NULL, temporal tstzrange,
embedding halfvec(1536),
created_at timestamptz DEFAULT now()
);
Indexes: HNSW for vector search (halfvec_cosine_ops), BM25 for full text via pg_textsearch, GIST for tree paths and temporal ranges, GIN for metadata. One table, five access patterns.
Ingestion
Each conversation turn becomes a row. Content is the speaker's text with any shared image description appended:
Melanie: Take a look at this. [shared image: a photo of a painting of a sunset over a lake]
Turns link via prev_id / next_id in metadata for navigation, and sit in an ltree hierarchy: conv.{speaker}.s{session_number}. Embeddings come from OpenAI's text-embedding-3-small. No LLM calls at ingest, one embedding per turn.
That is the entire ingestion path. No extraction, no summarization, no entity resolution.
Three tools over MCP
me_memory_search runs hybrid search across five modes, any of which can be combined, fused with Reciprocal Rank Fusion:
Semantic: vector similarity via HNSW
Fulltext: BM25 keyword matching via pg_textsearch
Grep: Postgres regex (~*) as an additional filter Tree: ltree path filtering, typically to scope to one speaker
Temporal: tstzrange containment and overlap
me_memory_get retrieves a memory by ID along with a window of surrounding turns, so the agent can read around a search hit. me_memory_tree exposes the path hierarchy.
The agent picks the modes, the combination, and the follow-up queries. There is no retrieval algorithm in our code. What there is instead is a system prompt describing how the tools behave, and that distinction matters for the rest of this post: several of our largest gains were prompt changes rather than code changes. The policy did not vanish. It became text.
The full stack, for anyone reproducing this: Postgres on Tiger Cloud with pgvector, pg_textsearch for BM25, and ltree. Embeddings from OpenAI text-embedding-3-small (1536 dimensions, stored as halfvec). Agent is Claude Sonnet or Haiku driven through the Claude Code CLI over MCP. Scoring runs through src/scorer.py, a port of the LoCoMo paper's evaluation.py.
The benchmark
LoCoMo tests long-term conversational memory across five categories:
Multi-hop (282): aggregation across sessions ("What has Melanie painted?")
Temporal (321): date-sensitive questions ("When did Caroline go to the pride parade?")
Open-domain (96): inference from personality ("Would Caroline be considered religious?")
Adversarial (446): questions that should be rejected, because they describe the wrong person's activity
Ten conversations, 1,986 questions. Each conversation spans roughly 20 sessions over several months between two speakers. Scoring uses token-level F1 with Porter stemming, matching the original LoCoMo evaluation.
What we learned (45+ experiments) We ran 45+ experiments over two weeks, one change at a time.
A note on how to read the numbers: every experiment was selected on a single conversation (conv-26, ~200 questions), where identical reruns vary by about ±0.04 F1. Adopted changes were then validated on all 10 conversations (1,986 questions) at checkpoints. Unless labeled otherwise, the numbers below come from those full runs, and deltas are paired on the same questions with 95% bootstrap intervals. One identical-config repeat of the full benchmark, run three hours apart, differed by 0.0002 F1, while its individual conversations swung by up to 0.06. That is the whole argument for validating on all ten conversations. The exception is findings 1 and 2, where the fixed pipeline was only ever measured on the tuning conversation; those numbers are labeled as such.
- Agentic search beats fixed retrieval
Our first version was a fixed pipeline: embed the question, run hybrid search, stuff the top 10 results into the prompt. On the tuning conversation it peaked at F1 0.481 after adding session dates to the results. The same day, on the same conversation and the same model, giving the agent a single search tool and letting it decide what to query scored 0.584. That gap of 0.10 is well outside the rerun noise and is the cleanest fixed-versus-agentic comparison we have. The final system scores 0.627 on that conversation.
The agent adapts its strategy per question. For "What has Melanie painted?" it might run:
For "When did Caroline go to the pride parade?" it combines:
search(semantic="Caroline pride parade", grep="pride|parade|march")
No fixed pipeline produces both, because no fixed pipeline knows which one this question needs.
- Tool schema complexity is a first-order variable
This is the finding that most changed how we think about the problem, and we learned it the hard way. When we swapped the simple single-parameter tool for the production-matching nine-parameter search schema, F1 on the tuning conversation dropped by 0.09, from 0.584 to 0.493. Same database. Same indexes. Same data. The agent spent its limited tool budget reading schemas instead of searching.
A good part of the program that followed was recovering ground the richer interface had cost. If the complexity lives in the tool interface, then the interface has a budget, and every parameter you add spends some of it before the agent retrieves anything.
- Adversarial rejection is the binding constraint
The pattern recurred throughout: any change that made the model more willing to answer, or that added more retrievable content, hurt adversarial rejection. A blanket instruction to infer from available evidence collapsed adversarial on the tuning conversation, from 0.681 to 0.362. Three-turn sliding-window memories took it to 0.468. Per-speaker profile memories, 0.553. "Use exact words from the memories" took it to zero. Most of these were reverted for that reason alone. The one inference change that survived was narrow, limited to questions phrased as "might," "would," or "could," and it held adversarial on the full benchmark.
The full-benchmark trajectory shows the cost we did accept. From the first 10-sample checkpoint to the final Sonnet run, overall F1 rose +0.050 (CI +0.035 to +0.065): multi-hop +0.105, single-hop +0.079, temporal +0.062. Adversarial fell from 0.922 to 0.870 (CI −0.076 to −0.027).
Anyone scoring without adversarial questions would have seen only the gains. That is the strongest argument we know for keeping them in the metric, and it is why the appendix spends as long as it does on the papers that drop them.
- Let the agent search as much as it wants
The largest single validated gain came from deleting a constraint. Removing the tool-call cap (previously six) moved full-benchmark F1 from 0.615 to 0.646 with Sonnet, a paired gain of +0.031 (95% CI +0.018 to +0.044). All ten conversations moved in the same direction, from +0.010 to +0.067. Multi-hop rose +0.048 and single-hop +0.042. Adversarial did not move.
The agent self-regulates: mean tool calls per question went from 2.7 to 3.2, and the final system averages 3.9 with a median of 3.
- Grep as a filter: a multi-hop gain with an adversarial cost
We added Postgres regex (~*) as a search parameter that filters the ranked semantic and full-text results. The agent uses it for synonym expansion on list questions:
grep: "painted|drew|art|canvas|sketch" On the full benchmark, the window in which grep shipped shows a trade rather than a free win: multi-hop +0.046 (CI +0.015 to +0.079) and single-hop +0.022, against adversarial −0.036 (CI −0.058 to −0.014). Overall F1 moved +0.012, not distinguishable from zero. Grep appears in about half of all searches.
Two prompt changes shipped in the same window, one telling the agent it could filter on the facts subtree and one tightening adversarial rejection, so the adversarial cost belongs to the window as a whole. Grep is the largest change in it and the one that most increases what the agent retrieves, which is exactly the profile finding 3 predicts.
We also found a bug worth knowing about: 13% of searches used grep alone, which fell into a filter-only path ordered by insertion time, effectively random. Grep-only searches now return an error and must combine with a ranked mode. That share dropped below 1%.
- Tree paths, image captions, and context windows
Four changes shipped between two full-benchmark checkpoints, along with the grep-only fix above: speaker-first tree paths, a case fix to them, image captions, and a wider context window.
Speaker-first tree paths. Memories are organized as conv.{speaker}.s{N}, and the agent may filter with tree: "conv.melanie.". We had prohibited speaker filtering to protect adversarial accuracy; with the attribution prompt in place the prohibition was no longer needed. A silent bug mattered more than the design: the agent wrote conv.Melanie. while paths are lowercase, so 30 of 42 filters matched nothing until the prompt said "speaker is lowercase". On the full benchmark the filter is used in 10% of searches with Haiku, 27% with Sonnet.
Image captions. LoCoMo turns include shared images with blip_caption descriptions, stored in metadata and invisible to search. Appending them to the content made 1,226 turns searchable for the first time.
Context windows on me_memory_get. The tool returns surrounding turns. We tested windows of 1, 2, and 3 on the tuning conversation; the differences were inside the noise floor, and we settled on 2 as a design choice rather than a measured optimum.
Together these moved full-benchmark F1 by +0.024 with Haiku (CI +0.010 to +0.039), and adversarial held at 0.887 to 0.893 on 441 questions. Which of the five carried the gain is not knowable from our data. On the tuning conversation the same bundle looked like +0.061, roughly the 3x inflation you should expect from selecting changes on one conversation.
- Facts are useless (for this task) We had Haiku extract atomic facts per session and stored them alongside raw turns. A clean ablation on the tuning conversation showed no difference. On the full benchmark, the run without facts scored 0.642 against 0.641 for the last run with them, with a paired interval spanning zero.
That comparison also absorbed two other changes, including removing a category hint from the prompt that we expected to cost F1, so the safe statement is that dropping facts did not lower F1. Facts added about 50% to ingestion time and diluted search results, so we removed extraction entirely.
- Multi-hop: iterative search improves recall, not yet F1
Multi-hop is our weakest category. Inspecting the zero-recall failures in the full run showed that 15 of 38 stopped after one or two searches. We added a prompt instruction to decompose multi-fact questions into sub-queries and let each search inform the next.
On the full benchmark with Sonnet, multi-hop recall rose from 0.592 to 0.651 (error-corrected; the paired interval on all questions is +0.028 to +0.092), temporal recall also rose, and the number of multi-hop questions with zero recall fell from 38 to 26 of 282. F1 did not move.
The mechanism is less clear than we first thought: average tool calls on multi-hop questions barely changed (4.0 to 4.2), so the gain appears to come from different queries rather than more of them, and an ingestion change that made image queries searchable shipped in the same run.
- What didn't help
Reverted with no signal or negative on the tuning conversation, none validated further: RRF weight tuning, score-based truncation, showing relevance scores to the agent, set-union merging from the Omni-SimpleMem paper, category-aware prompting, semantic-only search for inferential questions, and relative-to-absolute date prompting.
Candidate-pool depth had no measurable F1 effect at any setting; we kept 60 candidates and 15 results for recall. A standalone grep tool scored the same as the grep parameter and was dropped to keep a three-tool interface, which finding 2 suggests was the right call for reasons beyond tidiness. No fact-extraction variant survived: more thorough extraction, cross-session aggregation, and speaker profiles each hurt, and the base version added nothing on the full benchmark.
- Open-domain has a low ceiling on this benchmark
Analysis of open-domain failures showed that many gold answers are creative inferences never stated in the conversation. "What hobby could Andrew pick up?" expects "install a bird feeder," which appears nowhere in the text. Others require recognizing a location from a shared photo. We reported two benchmark errors in this category where the gold answer is supported by no text or image metadata.
Encouraging the model to infer rather than refuse helped a little. The practical limit is the benchmark, not the retrieval system.
What the numbers actually mean
Comparison against results reported in Omni-SimpleMem (Liu et al., 2026), the prior best on LoCoMo. All scores are raw F1, no error correction. Their paper re-ran every baseline under a single protocol, which is why the GPT-4o rows are directly comparable to each other.
Method
Model
Multi-hop
Single-hop
Temporal
Open-domain
Adversarial
Overall
MemVerse
GPT-4o
0.260
0.157
0.196
0.192
0.944
0.365
Claude-Mem
GPT-4o
0.294
0.153
0.167
0.243
0.915
0.383
Mem0
GPT-4o
0.309
0.156
0.217
0.295
0.857
0.397
A-MEM
GPT-4o
0.295
0.174
0.200
0.266
0.898
0.394
MemGPT
GPT-4o
0.305
0.188
0.246
0.305
0.843
0.404
SimpleMem
GPT-4o
0.318
0.195
0.235
0.308
0.802
0.432
Omni-SimpleMem
GPT-4o
0.556
0.365
0.255
0.641
0.835
0.598
Omni-SimpleMem
GPT-5.1
0.598
0.367
0.307
0.676
0.747
0.613
Ours
Claude Haiku
0.420
0.645
0.567
0.311
0.883
0.638
Ours
Claude Sonnet
0.453
0.673
0.625
0.400
0.870
0.665
Against Omni-SimpleMem's strongest configuration, our Sonnet run leads on single-hop (+0.306), temporal (+0.318), and adversarial (+0.123), and trails on multi-hop (−0.145) and open-domain (−0.276). Overall lead is +0.052.
Two things in that table are worth more than the ranking.
The temporal gap is the clearest evidence for the thesis. Dates live in a tstzrange column with a GIST index, and the agent can query them directly. A fixed retrieval pipeline has to hope the date survives the embedding, which mostly it does not. That is what "the complexity moved into the tool interface" buys you: the agent gets to use a real index instead of a vector approximation of one.
Their backbone sweep is a partial control on our model confound. Moving Omni-SimpleMem from GPT-4o to GPT-5.1 adds 0.015 overall to their pipeline. If a newer backbone were the whole story behind our result, that is roughly the size of effect we would expect it to explain. It raised their multi-hop and open-domain and dropped their adversarial from 0.835 to 0.747, which is the same pattern finding 3 describes: the newer model answers more and rejects less. It is suggestive, not decisive. See Limitations.
Our per-category results
Full 10 samples, error-corrected. We exclude 164 of 1,986 questions (8.3%) with benchmark errors: wrong gold answers, unsupported evidence citations, or gold answers requiring image understanding from photos with no text equivalent (for example, "Voyageurs National Park" as the expected answer when no text or metadata anywhere contains the park name). 156 of the exclusions come from the LoCoMo Audit; 7 are ours.
Category
Haiku F1
Sonnet F1
Sonnet Acc
Sonnet Recall
Multi-hop
0.445
0.485
0.825
0.651
Temporal
0.581
0.648
0.860
0.916
Open-domain
0.328
0.441
0.679
0.555
Single-hop
0.670
0.696
0.932
0.883
Adversarial
0.893
0.880
0.896
n/a
Overall
0.666
0.694
0.887
0.831
Four F1 numbers circulate in this post, so here they are in one place. Raw: 0.665 Sonnet, 0.638 Haiku.Error-corrected: 0.694 Sonnet, 0.666 Haiku. Raw is the number to compare against other papers, and it is the one we lead with because it is conservative. Note that Haiku's corrected score (0.666) sits a thousandth above Sonnet's raw score (0.665) by coincidence; they are not comparable to each other.
One disclosure about the corrections: 5 of our 7 additions to the error list are adversarial questions, which is also our strongest category. On the headline run those 5 move adversarial F1 by +0.010 and overall by +0.002. Small, but self-audited corrections in your best category are exactly the kind of thing a reader should be able to check, which is why adversarial-errors.json ships with the harness.
What it costs to run
We did not instrument this properly. The eval harness records tool calls per question but discards the CLI's cost, token, and latency fields, and the runs went through a subscription, so there is no bill to reconstruct. What the saved results do show:
Sonnet
Haiku
Mean tool calls per question
3.9
2.6
Median
3
2
So roughly four to five model turns per question, against one for a fixed pipeline. On the ingest side we make no LLM calls at all, one embedding per turn, while Omni-SimpleMem's own latency table shows its pipelines paying LLM calls at ingest.
Cheap ingest and expensive queries, versus the reverse. That is the actual trade, and which side of it you want depends on your read-to-write ratio. For a high-QPS product, paying four model turns per question may well be worse than paying once at ingest. We are not claiming our side of the trade is universally correct, only that it is the one nobody in this literature has been measuring.
Where the complexity went
Our system is a single Postgres table with standard indexes, exposed as three MCP tools. Theirs are multi-stage pipelines with summarizers, graph builders, and memory managers. We score higher on the deterministic metric.
That gap is not because we found a better retrieval algorithm. We did not write a retrieval algorithm. The gap is that we spent our effort on the tool interface rather than the ingestion pipeline, and that is where the measurable movement was: removing the tool-call cap (+0.031 overall), the nine-parameter schema that cost 0.09 on the tuning conversation before we clawed it back, grep as a filter (+0.046 multi-hop against −0.036 adversarial), the tree and caption bundle (+0.024 with Haiku), a prompt instruction to decompose questions (+0.059 multi-hop recall). None of those are storage decisions.
These are not additive and we are not claiming they sum to the difference between the fixed baseline and the final system. Several are bundles where attribution is not recoverable, one is measured on a single conversation, and some of the movement belongs to the model rather than the interface. The claim is directional: when we went looking for improvements, the interface is where we kept finding them, and the ingestion pipeline is where we kept finding nothing.
The corollary is uncomfortable for anyone building a memory product. If the pre-digestion layers exist to compensate for weak retrieval, and the model can now do retrieval itself, then those layers are carrying cost without carrying weight. You are maintaining an extraction pipeline, a summarization pass, and a graph, plus the operational surface of whatever specialized store holds them, to make up for a capability the model already has.
Why Postgres was the right substrate
The single-table design only works if one query engine can serve every access pattern the agent might want. Vector similarity, BM25 ranking, regex, hierarchical paths, and temporal ranges, all composable in one query, all consistent, no fan-out across services.
Postgres can do that now. The piece that used to be missing is real BM25 scoring as a native index, and pg_textsearch supplies it. We did not ablate this, so we cannot tell you what the system scores with ts_rank instead of BM25. What we can tell you is that the design assumes true BM25 ranking as a composable index, and the fallback if you do not have one is Postgres plus a search service, at which point you have a split architecture and a sync problem and the simplicity argument collapses. We run on Tiger Cloud partly because it is one of the few hosted providers that offers pg_textsearch. The hybrid search and RRF pattern is documented if you want a starting point.
The second reason is about process rather than architecture: 45 experiments in two weeks, across 119 eval runs. Each one needed an isolated database with the full corpus loaded, and database forking gave us that without a reload. That changes which experiments are worth running. The finding that facts are useless required building the extraction pipeline, re-ingesting the whole corpus through it, and measuring against a clean baseline. That is the kind of experiment you run once and interpret charitably when setup is expensive, and the kind you run properly and kill when forking is cheap.
There is a general version of this. Agent memory looks like the strongest possible case for specialized infrastructure: novel access patterns, unusual scale characteristics, a new workload nobody had tuned for. It turned out to need one table, five indexes, and three well-designed tools, at least at LoCoMo's scale. The reason to reach for a purpose-built system is that your database cannot express the workload. That is worth checking before you build the pipeline, because Postgres has been quietly absorbing the primitives.
Limitations
If you are inclined to argue with this post, start here. These are the objections we think are strongest, including the ones we cannot answer. The comparison is model-confounded, and we cannot fully control for it. Every baseline in the F1 table runs a GPT model. Ours runs Claude Sonnet or Haiku. That means "simple architecture with good tools beats complex pipeline" and "newer model reads dialogue better" predict the same result. The clean controls would be running our harness on GPT-4o and running a pipeline system on Sonnet. The first is not really possible: our harness shells out to the Claude Code CLI, so swapping the model means swapping the entire agent runtime, and we would be comparing two different harnesses rather than two models. The second is feasible, since SimpleMem's code is public, and it is the top of our list.
Two partial controls exist in the meantime. Our own fixed pipeline ran on Sonnet before we went agentic, on the same conversation, and peaked at 0.481 against the final system's 0.627. That comparison holds the model fixed and is the number the thesis actually rests on, with the caveat that the fixed pipeline was naive, tuned over seven runs on one conversation, and used a category hint we later removed as benchmark leakage. Separately, Omni-SimpleMem's own backbone sweep moves their pipeline only 0.015 from GPT-4o to GPT-5.1, which bounds how much a newer backbone plausibly explains. Neither is decisive. Read the architectural claim as well-supported rather than proven.
We did not record model versions. Only the CLI alias (--model sonnet) is logged, and print-mode transcripts were not saved, so the exact Sonnet and Haiku versions behind each run are unrecoverable. Comparisons across dates assume the alias resolved to the same model. The one identical-config repeat we have, three hours apart, showed no drift; across days it is untested. This is a real gap and we cannot close it retroactively.
Selection was on one conversation. Every change was chosen on conv-26 and then measured on all ten. The full-benchmark figures are not held-out in the strict sense, though the other nine conversations were seen only at the ten checkpoints, not during selection. Where the tuning conversation and the other nine disagree, we report the full-benchmark number: finding 6 is the clearest case, where a bundle that looked like +0.061 on conv-26 was +0.024 on the full set.
Attribution inside bundles is not recoverable. Findings 5, 6, and 7 each report the effect of everything that shipped between two checkpoints. Nothing isolates a single change on the full benchmark. Where we name a specific change as the likely cause, that is judgment, not measurement.
We do impose structure, just cheaply. We describe the ingestion path as raw turns, and relative to fact extraction it is. But we impose an ltree speaker and session hierarchy, and prev_id / next_id links are a linked list, which is a graph. The tree structure was part of one of our larger bundles. The difference from the systems we compare against is that our structure is deterministic and costs no LLM calls to build, not that it is absent.
We changed ingestion partway through. Appending image captions to the content column made 1,226 turns searchable. That is content enrichment at ingest time, which is the same category of move as fact extraction, and it was part of a bundle that helped. The distinction we would draw is that it made existing information searchable rather than generating new derived information, but a skeptic is entitled to call that a fine line.
LoCoMo is small. Roughly 20 sessions and two speakers per conversation, 5,882 turns total. Summarization and consolidation exist in production systems largely because raw turns eventually exceed any workable retrieval budget and recall degrades as the corpus grows. We have no scaling data: no turn count at which recall falls off, no index size, no p99 latency. "One table was enough" is demonstrated at a scale where a lot of things would be enough.
What's next
Run SimpleMem on Sonnet. Their code is public and this is the control that would separate architecture from backbone. It is ahead of everything else on this list.
Multi-hop is still the weakest category (F1 0.485 with Sonnet). Iterative search raised recall without moving F1, which means the agent is now finding evidence it is not successfully using. That is a different problem from the one we thought we had, and a more tractable one.
Instrument cost properly. Tokens, latency, and dollars per question, against a fixed pipeline on the same questions. We should have been recording this from the start.
We also have not touched the embedding model, tried re-ranking, or added query expansion at the retrieval level. And the interesting frontier is that a fixed pipeline can only exploit search strategies its authors thought of, while an agent can compose the tools in ways we did not plan for. Whether it actually discovers strategies we would not have written is something we can now test rather than assume.
LoCoMo has become the default benchmark for long-term conversational memory, and the published numbers have gotten high. Several systems now report above 90% accuracy. Those numbers are not comparable to each other, and some of them are measuring the benchmark's mistakes rather than the system's capability. We ran into this while trying to work out where we stood, and ended up auditing both the benchmark and the evaluation methodology.
Judge variance swamps the differences between systems
Since Mem0 (Chhikara et al., 2025), the field has moved toward LLM-as-judge accuracy as the primary LoCoMo metric. A judge model, usually GPT-4o-mini, compares the generated answer against the gold answer with generous grading: "as long as it touches on the same topic, count it as CORRECT."
There is a real reason for this. F1 penalizes valid paraphrases, since "May 7th" and "7 May 2023" are the same answer and score badly against each other. A judge catches semantic equivalence that token overlap misses.
The problem is that the judge is a hyperparameter, and almost nobody reports its value. We took one fixed set of predictions (Sonnet, full 10 samples) and ran it through four judge configurations. Same predictions, same gold answers, only the judge model and grading prompt changed.
Judge / Prompt
Multi-hop
Single-hop
Temporal
Open-domain
Overall (w/o Adv) Haiku / Prompt A
45.0
77.2
65.7
49.0
67.1
Haiku / Prompt B
55.0
87.3
77.6
64.6
77.9
GPT-4o-mini / Prompt A 46.5
79.8
75.1
51.0
70.9
GPT-4o-mini / Prompt B 79.8
90.6
82.2
62.5
85.1
18 points of spread in overall accuracy. 35 points on multi-hop. From the same predictions.
Prompt A is a neutral evaluation prompt: "decide whether the ground-truth content is present in the model's response." Prompt B is the Mem0 and APEX-MEM generous grading prompt. The prompt matters more than the judge model does. Holding the judge fixed and switching prompts moves overall accuracy 10.8 points on Haiku and 14.2 on GPT-4o-mini, with GPT-4o-mini multi-hop swinging 33.3 points on prompt alone. Holding the prompt fixed and switching judges moves overall accuracy 3.8 points (Prompt A) and 7.2 points (Prompt B).
Adversarial is the one stable category, at roughly a point of spread, because it is a binary match-or-reject decision requiring no semantic judgment. It is also the category most papers exclude.
Here is the accuracy comparison, with the judge configuration shown where it is known:
Method
Model
Judge
Judge Prompt
Multi-hop
Single-hop
Temporal
Open-domain
Adversarial
Overall (w/o Adv)
Overall (w/ Adv)
Mem0
GPT-4o
GPT-4o-mini Mem0
n/a
n/a
n/a
n/a
n/a
68.4
n/a
GAAMA
GPT-4o-mini
GPT-4o-mini
fact coverage*
72.2
87.2
71.9
49.3
n/a
78.9
n/a
Ours
Claude Haiku
GPT-4o-mini Mem0
70.9
81.3
72.9
42.7
89.7
75.3
78.5
APEX-MEM
Claude 4.5 Haiku
undisclosed
undisclosed
n/a
n/a
n/a
n/a
n/a
84.9
n/a
Ours
Claude Sonnet
GPT-4o-mini Mem0
79.8
90.6
82.2
62.5
88.6
85.1
85.9
APEX-MEM
Claude 4.5 Sonnet
undisclosed
undisclosed
n/a
n/a
n/a
n/a
n/a
88.4
n/a
APEX-MEM
GPT-5
undisclosed
undisclosed
86.3
89.9
90.6
91.7
86.8
89.5
88.9
MemMachine
GPT-4.1-mini
GPT-4o-mini
Mem0
88.3
95.1
91.6
71.9
n/a
91.7
n/a
HyperMem
GPT-4.1-mini
GPT-4o-mini
Mem0
93.6
96.1
89.7
70.8
n/a
92.7
n/a
- GAAMA uses a continuous key fact coverage score rather than binary CORRECT/WRONG, so its numbers are not directly comparable to the rest of the column. Three of the nine rows have an undisclosed judge. Four omit adversarial. One uses a different scoring function entirely. Note where we land: several systems report higher accuracy than we do. We are not claiming they are worse. We are claiming nobody can currently tell, which is a different and more fixable problem.
The benchmark has an 8.3% error rate
We excluded 164 of 1,986 questions (8.3%) as unreliable. Three failure types, of varying severity:
Wrong gold answers. The expected answer contradicts the conversation.
Unsupported citations. The evidence turns cited for a question do not contain the answer. This is the mildest of the three: the answer may still be recoverable elsewhere in the conversation, so an unsupported citation makes a question unreliable for measuring retrieval rather than strictly unanswerable.
Answers requiring image understanding with no text equivalent. The gold answer is "Voyageurs National Park," and no text, caption, or metadata anywhere names the park. The information exists only in the pixels of a shared trail-map photo.
156 of the exclusions come from the LoCoMo Audit; 7 are ours, 5 of them adversarial. We are confident there are more than 164.
One disclosure, since this appendix is about disclosure: adversarial is our own strongest category and our corrections are concentrated there. That is a conflict of interest, and the only reasonable response is to publish the corrections so anyone can disagree with individual calls.
The consequence is a ceiling. Any system reporting above roughly 91% accuracy is partly being evaluated on agreement with benchmark noise. Above that line there is no way to tell from outside whether a point of improvement is capability or coincidence. That is not an argument against using LoCoMo, which remains the best comparison point available. It is an argument for reporting both raw and error-corrected numbers, and for treating the top of the leaderboard as saturated rather than as a live race.
Dropping adversarial questions removes the hallucination guardrail
Most papers exclude LoCoMo's 446 adversarial questions from accuracy scoring. Those questions describe an activity belonging to the wrong speaker, and the correct response is to reject them.
Excluding them is not a neutral simplification. Adversarial questions are the only part of LoCoMo that penalizes guessing, and we have a direct measurement of what that hides. Over our full development trajectory, overall F1 rose +0.050 while adversarial fell from 0.922 to 0.870. Multi-hop gained 0.105, single-hop 0.079, temporal 0.062. Every one of those gains is visible without adversarial in the metric, and the entire cost of them is invisible. We did not set out to trade adversarial accuracy for the other categories; we discovered we had been doing it, and only because the number was in front of us.
In production that cost is hallucination: a memory system that confidently attributes one person's activities to another. If we had been scoring the way most papers score, we would have shipped it and called it progress.
What we would ask for
None of this requires a new benchmark. Four disclosures would make LoCoMo numbers comparable again:
Report F1 alongside judge accuracy. F1 is imperfect and penalizes paraphrase, but it is deterministic and reproducible with no configuration to disclose. It is the only number two teams can currently compare without coordinating.
Publish the full judge configuration. Model, exact prompt text, temperature. An appendix is fine.
Report adversarial separately, whether or not it is in your headline denominator.
Report raw and error-corrected numbers, and say which error list you used.
We have done all four. The harness, the corrections files, and the evidence file deriving every number in this post are in the repository. We would rather be checked than cited.
Why a 40-Year-Old Database Is Still Winning in the AI Era
Why Postgres wins in the AI era: 40 years of reliability, extensibility, and an open ecosystem. Agents need trusted foundations, not fragile architectures.