Agent Memory After the Hype: What Jev-Mem, KnowledgeX and BlogWriter Taught Me About My Own Grep… A follow-up analysis of agent memory benchmarks found that the widely cited claim that a plain filesystem grep agent scored 74.0 percent on LoCoMo versus Mem0's graph configuration at 68.5 percent rests on two self-reported numbers from different pipelines, with Letta's own blog post acknowledging the comparison is apples-to-oranges. The author also cited the Penfield Labs audit finding roughly 6.4 percent of LoCoMo's answer key is wrong and that a GPT-4o mini judge accepted about 63 percent of deliberately wrong but topically adjacent answers, undercutting the 5.5-point gap. The piece examines Jev-Mem (arXiv 2609.23986), which uses a non-generative System-One controller for frequent memory decisions and reserves the System-Two LLM for final reasoning, plus the KnowledgeX MCP server and Jesse Liberty's BlogWriter memory approach. A week and a half ago I published “Your Agent’s Memory Is Probably Wrong.” It opened with a number I was a little too pleased with: on the LoCoMo benchmark, a plain filesystem agent using grep scored 74.0 percent, and Mem0’s graph configuration scored 68.5 percent. Simple beats clever. Folder beats graph. People shared it for exactly that reason. Then three things landed in my reading list in the same week, and each one poked at a different part of that claim. A paper called Jev-Mem that beats everything on LoCoMo by using a graph and fewer LLM calls. An open-source MCP server called KnowledgeX that builds a markdown knowledge graph across Claude conversations. And a short, almost boring post from Jesse Liberty about how his BlogWriter app handles memory, which turned out to be the most useful of the three. So I did what I should have done before publishing the first time. I went back to the source of my own headline, rebuilt a small version of my four-layer design in .NET with EF Core and SQLite, added supersession and TTL properly, and ran a small retrieval comparison: grep-style search against embedding search with a local Ollama model. This is the write-up. Some of my original claim held. Some of it didn’t, and one part was framed in a way I’m now a bit embarrassed about. The 74.0 percent figure comes from Letta’s blog post “Benchmarking AI Agent Memory: Is a Filesystem All You Need?” I cited it as if it were a neutral result. It isn’t. It’s a vendor benchmarking its own agent framework, and the 68.5 percent it compares against is Mem0’s own self-reported number for its best graph variant. Two self-reported numbers, two different pipelines, one table. To be fair to Letta, they say this themselves. The post explicitly calls comparing agent frameworks and memory tools an apples-to-oranges exercise. I read that sentence the first time and still led with the comparison. The second thing I glossed over is what “grep” meant in that setup. The Letta agent ran on GPT-4o mini with tool rules constraining its calls, and it had grep, a search files tool that does semantic search, plus open and close for files. So it wasn't grep against a graph. It was an agent with a few familiar file tools, including a semantic one, against an extraction pipeline. That's a different and less tweetable claim. The third thing is the one that stings. Not long after the memory piece, I wrote about the Penfield Labs audit of LoCoMo: roughly 6.4 percent of the answer key is wrong, and the standard GPT-4o mini judge accepted about 63 percent of deliberately wrong but topically adjacent answers. I wrote that a LoCoMo score near the top of the leaderboard tells you how closely a system agrees with a flawed key under a lenient judge. Then I never went back and applied that to my own opening line. A 5.5 point gap on that benchmark, between two systems using different pipelines and reporting their own numbers, is not the knockout I presented it as. So the honest version of my original claim is: “a vendor reported that its file-tool agent beat a competitor’s self-reported number on a benchmark with known label and judge problems.” Still interesting. Not proof of anything. Jev-Mem is a paper arXiv 2609.23986 with public code at github.com/libingzheren/Jev-Mem. The idea is clean. Most memory systems put an autoregressive LLM in the middle of every memory decision: what kind of observation this is, which memories it relates to, which retrieval path to take, when to stop searching. Jev-Mem hands those small, frequent decisions to a lightweight, non-generative "System-One" controller, and only calls the "System-Two" LLM for final reasoning and answer synthesis. The memory itself is a shared graph where each observation is a node, connected by four kinds of edges: semantic, temporal, causal and entity. The headline numbers, all reported by the paper’s own authors: Jev-Mem on LoCoMo, as reported by the authors GPT-4o mini answers and judges Metric Jev-Mem Comparison in the paper----------------------------- ---------- ----------------------------------Overall LLM-as-a-Judge 0.777 MAGMA 0.700 strongest baseline Memory build time 158 s Nemori 1,044 s 6.6x slower Average query latency 0.93 s MAGMA 1.47 sBy category Jev-Mem MAGMA----------------------------- ---------- ----------Single-hop 0.802 0.776Multi-hop 0.623 0.569Temporal 0.637 0.650 <- MAGMA still winsOpen-domain 0.618 0.517Adversarial 0.962 0.742 A few things I want to be explicit about before anyone quotes this. These are the authors’ numbers. I couldn’t find an independent reproduction, and I haven’t run their code myself. The judge is GPT-4o mini, the same setup the LoCoMo audit showed accepts a lot of vague wrong answers. And the biggest single jump in that table is the adversarial category, 0.742 to 0.962, which is also a category many LoCoMo evaluations leave out entirely. I’d want to see the overall score with and without it before I’d treat 0.777 as the number. The most popular write-up of the paper I saw this week was from Agent Native, and it reads like a launch post with a paid program link at the bottom. I’m treating it as promotion. The paper is the source. The detail I find most interesting is the one row where Jev-Mem loses: temporal. MAGMA still edges it out there. Keep that in mind, because temporal is exactly where my own rebuild fell over. What Jev-Mem does to my original claim is subtle. On the surface, a multi-relational graph winning on LoCoMo argues against “folder beats graph.” But its actual thesis rhymes with the reason I gave for the grep result: keep expensive, flexible generation out of the hot path of memory control, and make the routine decisions cheap and predictable. Grep is cheap and predictable. So is a classifier. The difference is that Jev-Mem keeps structure time, cause, entity that a folder of text files doesn’t have. KnowledgeX, by Rahul Nayak, is an open-source MCP server that lets Claude build a personal knowledge library out of your conversations, stored as markdown with links between notes, so it’s a graph you can open in any editor and take with you. I’ll be upfront: I haven’t run it yet. The repository page wouldn’t load for me this week and the article didn’t render either, so everything I can say about it comes from the article summary and the project’s public pull request list one of them is titled “Offer to save knowledge when it comes up”, which suggests saving is meant to be explicit rather than silent . I’m not going to describe internals I haven’t seen. What makes it relevant here is the storage choice. Markdown files with explicit links sit right in the middle of my original argument. They’re as grep-able as Letta’s folder, and as model-familiar, but they carry edges like a graph. If my “familiarity” explanation for the grep result is right, this is the kind of format that should get the benefit of both. The questions I’d want answered before I trust it as agent memory rather than a personal notebook are the same ones my rebuild below kept tripping over. When a note goes stale, does anything mark it as superseded, or do the old and new versions sit side by side as equally valid markdown? Is there provenance on each note, so I know whether the user said it or the model inferred it? And what happens when two sessions write to the same note at once? A markdown graph is a great format. It isn’t a lifecycle. Jesse Liberty’s post “Memory in BlogWriter” is short, and it reads like a code tour rather than an argument, which is probably why it changed my thinking the most. He notes at the bottom that BlogWriter itself wrote the first draft and he edited it, which I enjoyed. The design, as he describes it: Here’s what that did to my four-layer design. My “buffer” layer was a deque of turns. Jesse’s equivalent is a typed document of workflow state, and the conversation history isn’t kept at all. Continuity comes from the app deciding what matters and writing it down in a known shape. That’s not a lesser form of memory. For a workflow agent it’s probably the right one, because nothing has to be retrieved: the state is the context. It also reminded me that the concurrency token I had in my original .NET code wasn’t decoration. BlogWriter’s ETag check is the same idea as the RowVersion column I wrote about. The difference is that his is the core of the design, and mine was a footnote. The piece that pushed back hardest on my design was Nitin Bisht’s “AI Agents Don’t Need Bigger Context Windows. They Need Memory.” It’s member-only, so I’ve only read part of it, but the framing alone is worth stealing. He opens with a coding agent that hit a missing DB URL error, fixed it with a local setting, forgot the fix the next day, made it again, and on the third day removed it as a hack. His metaphor is that the context window is a desk that gets cleared at the end of every session, and memory is what decides what goes back on the desk. His loop has three steps: write, manage, read. Managing is where decay and supersession live: facts lose weight unless they're reaffirmed, and newer facts explicitly replace older ones. That middle step is the thing I under-built. My original design had TTL and supersession on write, but no decay on read. Everything active was equally loud forever. His loop treats forgetting as routine maintenance rather than an exception, and after the rebuild below I think he’s right. I didn’t benchmark decay this time, though, because sixty synthetic days in a test file is not enough history for a decay curve to mean anything. That’s on my list. I wanted something small enough to read in one sitting and run without a cloud account. One SQLite file, one table, EF Core for the lifecycle rules, and Ollama for embeddings. Setup: dotnet new console -n MemoryLabcd MemoryLabdotnet add package Microsoft.EntityFrameworkCore.Sqlite Local embeddings, no API keyollama pull nomic-embed-textollama serve skip if the Ollama app is already running If you’d rather not install Ollama on the host, the Docker route works the same way: docker run -d --name ollama -p 11434:11434 -v ollama:/root/.ollama ollama/ollamadocker exec ollama ollama pull nomic-embed-text The .csproj only needs the one package and to copy the test file to the output folder: