cd /news/ai-agents/agent-memory-after-the-hype-what-jev… · home › topics › ai-agents › article
[ARTICLE · art-148856] src=pub.towardsai.net ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Agent Memory After the Hype: What Jev-Mem, KnowledgeX and BlogWriter Taught Me About My Own Grep…

A follow-up analysis of agent memory benchmarks found that the widely cited claim that a plain filesystem grep agent scored 74.0 percent on LoCoMo versus Mem0's graph configuration at 68.5 percent rests on two self-reported numbers from different pipelines, with Letta's own blog post acknowledging the comparison is apples-to-oranges. The author also cited the Penfield Labs audit finding roughly 6.4 percent of LoCoMo's answer key is wrong and that a GPT-4o mini judge accepted about 63 percent of deliberately wrong but topically adjacent answers, undercutting the 5.5-point gap. The piece examines Jev-Mem (arXiv 2609.23986), which uses a non-generative System-One controller for frequent memory decisions and reserves the System-Two LLM for final reasoning, plus the KnowledgeX MCP server and Jesse Liberty's BlogWriter memory approach.

by read34 min views2 publishedOct 10, 2026

A week and a half ago I published “Your Agent’s Memory Is Probably Wrong.” It opened with a number I was a little too pleased with: on the LoCoMo benchmark, a plain filesystem agent using grep scored 74.0 percent, and Mem0’s graph configuration scored 68.5 percent. Simple beats clever. Folder beats graph. People shared it for exactly that reason.

Then three things landed in my reading list in the same week, and each one poked at a different part of that claim. A paper called Jev-Mem that beats everything on LoCoMo by using a graph and fewer LLM calls. An open-source MCP server called KnowledgeX that builds a markdown knowledge graph across Claude conversations. And a short, almost boring post from Jesse Liberty about how his BlogWriter app handles memory, which turned out to be the most useful of the three.

So I did what I should have done before publishing the first time. I went back to the source of my own headline, rebuilt a small version of my four-layer design in .NET with EF Core and SQLite, added supersession and TTL properly, and ran a small retrieval comparison: grep-style search against embedding search with a local Ollama model.

This is the write-up. Some of my original claim held. Some of it didn’t, and one part was framed in a way I’m now a bit embarrassed about.

The 74.0 percent figure comes from Letta’s blog post “Benchmarking AI Agent Memory: Is a Filesystem All You Need?” I cited it as if it were a neutral result. It isn’t. It’s a vendor benchmarking its own agent framework, and the 68.5 percent it compares against is Mem0’s own self-reported number for its best graph variant. Two self-reported numbers, two different pipelines, one table.

To be fair to Letta, they say this themselves. The post explicitly calls comparing agent frameworks and memory tools an apples-to-oranges exercise. I read that sentence the first time and still led with the comparison.

The second thing I glossed over is what “grep” meant in that setup. The Letta agent ran on GPT-4o mini with tool rules constraining its calls, and it had grep, a search_files tool that does semantic search, plus open and close for files. So it wasn't grep against a graph. It was an agent with a few familiar file tools, including a semantic one, against an extraction pipeline. That's a different and less tweetable claim.

The third thing is the one that stings. Not long after the memory piece, I wrote about the Penfield Labs audit of LoCoMo: roughly 6.4 percent of the answer key is wrong, and the standard GPT-4o mini judge accepted about 63 percent of deliberately wrong but topically adjacent answers. I wrote that a LoCoMo score near the top of the leaderboard tells you how closely a system agrees with a flawed key under a lenient judge. Then I never went back and applied that to my own opening line. A 5.5 point gap on that benchmark, between two systems using different pipelines and reporting their own numbers, is not the knockout I presented it as.

So the honest version of my original claim is: “a vendor reported that its file-tool agent beat a competitor’s self-reported number on a benchmark with known label and judge problems.” Still interesting. Not proof of anything.

Jev-Mem is a paper (arXiv 2609.23986) with public code at github.com/libingzheren/Jev-Mem. The idea is clean. Most memory systems put an autoregressive LLM in the middle of every memory decision: what kind of observation this is, which memories it relates to, which retrieval path to take, when to stop searching. Jev-Mem hands those small, frequent decisions to a lightweight, non-generative "System-One" controller, and only calls the "System-Two" LLM for final reasoning and answer synthesis.

The memory itself is a shared graph where each observation is a node, connected by four kinds of edges: semantic, temporal, causal and entity.

The headline numbers, all reported by the paper’s own authors:

Jev-Mem on LoCoMo, as reported by the authors (GPT-4o mini answers and judges)Metric                         Jev-Mem     Comparison in the paper-----------------------------  ----------  ----------------------------------Overall LLM-as-a-Judge         0.777       MAGMA 0.700 (strongest baseline)Memory build time              158 s       Nemori 1,044 s (6.6x slower)Average query latency          0.93 s      MAGMA 1.47 sBy category                    Jev-Mem     MAGMA-----------------------------  ----------  ----------Single-hop                     0.802       0.776Multi-hop                      0.623       0.569Temporal                       0.637       0.650   <- MAGMA still winsOpen-domain                    0.618       0.517Adversarial                    0.962       0.742

A few things I want to be explicit about before anyone quotes this.

These are the authors’ numbers. I couldn’t find an independent reproduction, and I haven’t run their code myself. The judge is GPT-4o mini, the same setup the LoCoMo audit showed accepts a lot of vague wrong answers. And the biggest single jump in that table is the adversarial category, 0.742 to 0.962, which is also a category many LoCoMo evaluations leave out entirely. I’d want to see the overall score with and without it before I’d treat 0.777 as the number.

The most popular write-up of the paper I saw this week was from Agent Native, and it reads like a launch post with a paid program link at the bottom. I’m treating it as promotion. The paper is the source.

The detail I find most interesting is the one row where Jev-Mem loses: temporal. MAGMA still edges it out there. Keep that in mind, because temporal is exactly where my own rebuild fell over.

What Jev-Mem does to my original claim is subtle. On the surface, a multi-relational graph winning on LoCoMo argues against “folder beats graph.” But its actual thesis rhymes with the reason I gave for the grep result: keep expensive, flexible generation out of the hot path of memory control, and make the routine decisions cheap and predictable. Grep is cheap and predictable. So is a classifier. The difference is that Jev-Mem keeps structure (time, cause, entity) that a folder of text files doesn’t have.

KnowledgeX, by Rahul Nayak, is an open-source MCP server that lets Claude build a personal knowledge library out of your conversations, stored as markdown with links between notes, so it’s a graph you can open in any editor and take with you.

I’ll be upfront: I haven’t run it yet. The repository page wouldn’t load for me this week and the article didn’t render either, so everything I can say about it comes from the article summary and the project’s public pull request list (one of them is titled “Offer to save knowledge when it comes up”, which suggests saving is meant to be explicit rather than silent). I’m not going to describe internals I haven’t seen.

What makes it relevant here is the storage choice. Markdown files with explicit links sit right in the middle of my original argument. They’re as grep-able as Letta’s folder, and as model-familiar, but they carry edges like a graph. If my “familiarity” explanation for the grep result is right, this is the kind of format that should get the benefit of both.

The questions I’d want answered before I trust it as agent memory rather than a personal notebook are the same ones my rebuild below kept tripping over. When a note goes stale, does anything mark it as superseded, or do the old and new versions sit side by side as equally valid markdown? Is there provenance on each note, so I know whether the user said it or the model inferred it? And what happens when two sessions write to the same note at once? A markdown graph is a great format. It isn’t a lifecycle.

Jesse Liberty’s post “Memory in BlogWriter” is short, and it reads like a code tour rather than an argument, which is probably why it changed my thinking the most. (He notes at the bottom that BlogWriter itself wrote the first draft and he edited it, which I enjoyed.)

The design, as he describes it:

Here’s what that did to my four-layer design. My “buffer” layer was a deque of turns. Jesse’s equivalent is a typed document of workflow state, and the conversation history isn’t kept at all. Continuity comes from the app deciding what matters and writing it down in a known shape. That’s not a lesser form of memory. For a workflow agent it’s probably the right one, because nothing has to be retrieved: the state is the context.

It also reminded me that the concurrency token I had in my original .NET code wasn’t decoration. BlogWriter’s ETag check is the same idea as the RowVersion column I wrote about. The difference is that his is the core of the design, and mine was a footnote.

The piece that pushed back hardest on my design was Nitin Bisht’s “AI Agents Don’t Need Bigger Context Windows. They Need Memory.” It’s member-only, so I’ve only read part of it, but the framing alone is worth stealing.

He opens with a coding agent that hit a missing DB_URL error, fixed it with a local setting, forgot the fix the next day, made it again, and on the third day removed it as a hack. His metaphor is that the context window is a desk that gets cleared at the end of every session, and memory is what decides what goes back on the desk. His loop has three steps: write, manage, read. Managing is where decay and supersession live: facts lose weight unless they're reaffirmed, and newer facts explicitly replace older ones.

That middle step is the thing I under-built. My original design had TTL and supersession on write, but no decay on read. Everything active was equally loud forever. His loop treats forgetting as routine maintenance rather than an exception, and after the rebuild below I think he’s right. I didn’t benchmark decay this time, though, because sixty synthetic days in a test file is not enough history for a decay curve to mean anything. That’s on my list.

I wanted something small enough to read in one sitting and run without a cloud account. One SQLite file, one table, EF Core for the lifecycle rules, and Ollama for embeddings.

Setup:

dotnet new console -n MemoryLabcd MemoryLabdotnet add package Microsoft.EntityFrameworkCore.Sqlite# Local embeddings, no API keyollama pull nomic-embed-textollama serve   # skip if the Ollama app is already running

If you’d rather not install Ollama on the host, the Docker route works the same way:

docker run -d --name ollama -p 11434:11434 -v ollama:/root/.ollama ollama/ollamadocker exec ollama ollama pull nomic-embed-text

The .csproj only needs the one package and to copy the test file to the output folder:

<Project Sdk="Microsoft.NET.Sdk"><PropertyGroup>    <OutputType>Exe</OutputType>    <TargetFramework>net10.0</TargetFramework>    <ImplicitUsings>enable</ImplicitUsings>    <Nullable>enable</Nullable>  </PropertyGroup>  <ItemGroup>    <!-- Added with: dotnet add package Microsoft.EntityFrameworkCore.Sqlite -->    <PackageReference Include="Microsoft.EntityFrameworkCore.Sqlite" Version="10.*" />  </ItemGroup>  <ItemGroup>    <None Update="cases.json" CopyToOutputDirectory="PreserveNewest" />  </ItemGroup></Project>

The entry carries the four layers, a status instead of a delete flag, provenance, a TTL, a pointer to whatever superseded it, and the embedding as a blob. SQLite has no rowversion, so I use a Guid that changes on every update. It does the same job as BlogWriter's ETag: a stale writer's UPDATE matches zero rows and EF Core throws.

using System.Runtime.InteropServices;using Microsoft.EntityFrameworkCore;using Microsoft.EntityFrameworkCore.ChangeTracking;using Microsoft.EntityFrameworkCore.Storage.ValueConversion;namespace MemoryLab;public enum MemoryLayer { Buffer, Episodic, Semantic, Procedural }public enum MemoryStatus { Active, Superseded, Expired, Forgotten }public sealed class MemoryEntry{    public long Id { get; set; }    public required string Key { get; set; }    public required string Value { get; set; }    public MemoryLayer Layer { get; set; }    public MemoryStatus Status { get; set; } = MemoryStatus.Active;    public required string Source { get; set; }          // provenance: "user", "tool:repo_read", "agent"    public DateTime CreatedAtUtc { get; set; }           // DateTime, not DateTimeOffset: SQLite can't compare DTO server side    public DateTime? ExpiresAtUtc { get; set; }    public long? SupersededById { get; set; }    public float[]? Embedding { get; set; }    // SQLite has no rowversion column. A Guid that changes on every update does the same job    // as Cosmos DB's ETag or SQL Server's rowversion: a stale writer's UPDATE matches zero rows.    public Guid ConcurrencyStamp { get; set; } = Guid.NewGuid();}public sealed class MemoryDb(string connectionString) : DbContext{    public DbSet<MemoryEntry> Memories => Set<MemoryEntry>();    protected override void OnConfiguring(DbContextOptionsBuilder options) =>        options.UseSqlite(connectionString);    protected override void OnModelCreating(ModelBuilder model)    {        var e = model.Entity<MemoryEntry>();        e.HasIndex(m => new { m.Key, m.Status });        e.Property(m => m.Layer).HasConversion<string>();        e.Property(m => m.Status).HasConversion<string>();        e.Property(m => m.ConcurrencyStamp).IsConcurrencyToken();        e.Property(m => m.Embedding).HasConversion(            new ValueConverter<float[], byte[]>(v => Floats.ToBytes(v), v => Floats.FromBytes(v)),            new ValueComparer<float[]>((a, b) => Floats.Same(a, b), v => Floats.Hash(v), v => Floats.Copy(v)));    }    public override Task<int> SaveChangesAsync(CancellationToken ct = default)    {        foreach (var entry in ChangeTracker.Entries<MemoryEntry>())        {            if (entry.State == EntityState.Modified)                entry.Entity.ConcurrencyStamp = Guid.NewGuid();        }        return base.SaveChangesAsync(ct);    }}// Helpers live outside the lambdas because expression trees can't touch Span<T>.public static class Floats{    public static byte[] ToBytes(float[] v) => MemoryMarshal.AsBytes(v.AsSpan()).ToArray();    public static float[] FromBytes(byte[] v) => MemoryMarshal.Cast<byte, float>(v.AsSpan()).ToArray();    public static bool Same(float[]? a, float[]? b) => a is null ? b is null : b is not null && a.AsSpan().SequenceEqual(b);    public static int Hash(float[] v) => v.Length == 0 ? 0 : HashCode.Combine(v.Length, v[0], v[^1]);    public static float[] Copy(float[] v) => (float[])v.Clone();}

One gotcha worth knowing up front: EF Core’s SQLite provider can’t compare or sort DateTimeOffset values in SQL, so the timestamps are UTC DateTime.

Every non-episodic write looks for the active entry with the same key and marks it superseded in the same transaction. Episodes never replace each other, because they’re a log of things that happened. Reads check expiry again even if the sweep hasn’t run, so a late background job can’t leak an expired fact.

using Microsoft.EntityFrameworkCore;namespace MemoryLab;public sealed class MemoryService(MemoryDb db, OllamaEmbedder? embedder, TimeProvider clock){    /// <summary>    /// Writes a memory. Anything that isn't episodic supersedes the active entry with the same key.    /// Nothing is overwritten or deleted: the old row is marked Superseded and points at its replacement.    /// </summary>    public async Task<MemoryEntry> WriteAsync(        string key, string value, MemoryLayer layer, string source,        TimeSpan? ttl = null, CancellationToken ct = default)    {        var now = clock.GetUtcNow().UtcDateTime;        var entry = new MemoryEntry        {            Key = key,            Value = value,            Layer = layer,            Source = source,            CreatedAtUtc = now,            ExpiresAtUtc = ttl is null ? null : now + ttl.Value,        };        if (embedder is not null)            entry.Embedding = await embedder.EmbedDocumentAsync($"{key}: {value}", ct);        await using var tx = await db.Database.BeginTransactionAsync(ct);        // Episodes are a log of things that happened. They never replace each other.        List<MemoryEntry> current = layer == MemoryLayer.Episodic            ? []            : await db.Memories.Where(m => m.Key == key && m.Status == MemoryStatus.Active).ToListAsync(ct);        db.Memories.Add(entry);        await db.SaveChangesAsync(ct); // need entry.Id for the back pointer        foreach (var old in current)        {            old.Status = MemoryStatus.Superseded;            old.SupersededById = entry.Id;        }        await db.SaveChangesAsync(ct); // throws DbUpdateConcurrencyException if someone else got there first        await tx.CommitAsync(ct);        return entry;    }    /// <summary>TTL sweep. Run it from a BackgroundService or a scheduled job.</summary>    public Task<int> ExpireAsync(CancellationToken ct = default)    {        var now = clock.GetUtcNow().UtcDateTime;        return db.Memories            .Where(m => m.Status == MemoryStatus.Active && m.ExpiresAtUtc != null && m.ExpiresAtUtc < now)            .ExecuteUpdateAsync(s => s.SetProperty(m => m.Status, MemoryStatus.Expired), ct);    }    /// <summary>    /// What retrieval is allowed to see. Governed = active and not past its TTL (checked again at    /// read time, so a late sweep can't leak expired facts). Raw = everything except forgotten rows,    /// which is what my original one-table design effectively searched.    /// </summary>    public Task<List<MemoryEntry>> CandidatesAsync(bool governed, CancellationToken ct = default)    {        var now = clock.GetUtcNow().UtcDateTime;        var q = db.Memories.AsNoTracking().Where(m => m.Status != MemoryStatus.Forgotten);        if (governed)            q = q.Where(m => m.Status == MemoryStatus.Active && (m.ExpiresAtUtc == null || m.ExpiresAtUtc > now));        return q.ToListAsync(ct);    }    /// <summary>Walks the supersession chain backwards, for "what was it before?" questions.</summary>    public async Task<List<MemoryEntry>> HistoryAsync(string key, CancellationToken ct = default) =>        await db.Memories.AsNoTracking()            .Where(m => m.Key == key)            .OrderByDescending(m => m.CreatedAtUtc)            .ToListAsync(ct);}/// <summary>A clock the benchmark can move forward, so TTLs and supersession play out over "days".</summary>public sealed class SimClock(DateTimeOffset start) : TimeProvider{    public DateTimeOffset Now { get; set; } = start;    public override DateTimeOffset GetUtcNow() => Now;}

The grep retriever is deliberately dumb: lowercase the question, drop stop words, count how many remaining terms appear anywhere in the key and value, break ties by recency. The embedding retriever calls Ollama’s /api/embed and ranks by cosine similarity. Note the search_query: and search_document: prefixes. nomic-embed-text was trained with them, and leaving them off is an easy way to make embeddings look worse than they are.

using System.Net.Http.Json;using System.Text.RegularExpressions;namespace MemoryLab;public interface IRetriever{    string Name { get; }    Task<IReadOnlyList<MemoryEntry>> SearchAsync(        string query, IReadOnlyList<MemoryEntry> candidates, int k, CancellationToken ct = default);}/// <summary>/// Roughly "grep -i -c" for every meaningful word in the question, over key + value./// Ties go to the newest entry, which is what I'd do with ls -t | xargs grep./// </summary>public sealed partial class GrepRetriever : IRetriever{    private static readonly HashSet<string> Stop =    [        "the", "and", "for", "are", "was", "what", "which", "who", "how", "does", "did",        "do", "is", "it", "of", "to", "in", "on", "a", "an", "we", "i", "am", "can", "be",        "this", "that", "with", "from", "now", "use", "using", "should", "these", "days",        "right", "still", "before", "after", "need", "kind", "like", "where", "current",    ];    [GeneratedRegex(@"[^a-z0-9_./-]+")]    private static partial Regex Splitter();    public string Name => "grep";    public static string[] Terms(string text) =>        Splitter().Split(text.ToLowerInvariant())            .Select(t => t.Trim('.', '/', '-'))            .Where(t => t.Length >= 3 && !Stop.Contains(t))            .Distinct()            .ToArray();    public Task<IReadOnlyList<MemoryEntry>> SearchAsync(        string query, IReadOnlyList<MemoryEntry> candidates, int k, CancellationToken ct = default)    {        var terms = Terms(query);        IReadOnlyList<MemoryEntry> hits = candidates            .Select(m => (m, score: terms.Count(t => $"{m.Key} {m.Value}".ToLowerInvariant().Contains(t))))            .Where(x => x.score > 0)            .OrderByDescending(x => x.score)            .ThenByDescending(x => x.m.CreatedAtUtc)            .ThenByDescending(x => x.m.Id)            .Take(k)            .Select(x => x.m)            .ToList();        return Task.FromResult(hits);    }}/// <summary>Cosine similarity over embeddings from a local Ollama model. No API key, no per-call cost.</summary>public sealed class EmbeddingRetriever(OllamaEmbedder embedder) : IRetriever{    public string Name => $"embed:{embedder.Model}";    public async Task<IReadOnlyList<MemoryEntry>> SearchAsync(        string query, IReadOnlyList<MemoryEntry> candidates, int k, CancellationToken ct = default)    {        var q = await embedder.EmbedQueryAsync(query, ct);        return candidates            .Where(m => m.Embedding is not null)            .Select(m => (m, score: Cosine(q, m.Embedding!)))            .OrderByDescending(x => x.score)            .Take(k)            .Select(x => x.m)            .ToList();    }    private static double Cosine(float[] a, float[] b)    {        double dot = 0, na = 0, nb = 0;        for (var i = 0; i < a.Length; i++)        {            dot += a[i] * b[i];            na += a[i] * a[i];            nb += b[i] * b[i];        }        return dot / (Math.Sqrt(na) * Math.Sqrt(nb) + 1e-12);    }}public sealed class OllamaEmbedder(HttpClient http, string model = "nomic-embed-text"){    public string Model => model;    // nomic-embed-text was trained with task prefixes. Leaving them off quietly costs recall.    public Task<float[]> EmbedDocumentAsync(string text, CancellationToken ct = default) =>        EmbedAsync("search_document: " + text, ct);    public Task<float[]> EmbedQueryAsync(string text, CancellationToken ct = default) =>        EmbedAsync("search_query: " + text, ct);    private async Task<float[]> EmbedAsync(string input, CancellationToken ct)    {        using var response = await http.PostAsJsonAsync("/api/embed", new { model, input }, ct);        response.EnsureSuccessStatusCode();        var body = await response.Content.ReadFromJsonAsync<EmbedResponse>(ct)                   ?? throw new InvalidOperationException("Empty response from Ollama.");        return body.Embeddings[0];    }    private sealed record EmbedResponse(float[][] Embeddings);}

I wrote 33 memory writes spread over sixty simulated days for a fictional payments team, then 26 questions asked on day sixty. Six writes supersede earlier ones (staging moved from MySQL to Postgres, the gateway timeout changed, the feature flag vendor changed, a procedural rule got stricter), two buffer and promo entries expire, and there’s a deliberate distractor: a reporting replica that is still MySQL, to tempt a stale staging answer.

I wrote these myself, which is the obvious weakness. Twenty-six questions written by the person who also wrote the retrievers is a smoke test, not a benchmark. Every case lives in this file so you can see exactly what I measured and swap in your own:

{  "nowDay": 60,  "writes": [    { "t": 0,  "layer": "Semantic",   "key": "staging.db",        "value": "Staging database is MySQL 8 running on db-stg-01.", "source": "user" },    { "t": 0,  "layer": "Semantic",   "key": "deploy.target",     "value": "Default deploy target for the billing service is the staging cluster.", "source": "user" },    { "t": 2,  "layer": "Semantic",   "key": "gateway.timeout",   "value": "Payment gateway timeout is controlled by PAYGW_TIMEOUT_MS, set to 30000.", "source": "tool:repo_read" },    { "t": 3,  "layer": "Semantic",   "key": "owner.invoice",     "value": "The billing service owns invoice numbering.", "source": "user" },    { "t": 4,  "layer": "Procedural", "key": "rule.tests",        "value": "Before proposing a commit in billing-service, run dotnet test and attach the summary.", "source": "user" },    { "t": 5,  "layer": "Semantic",   "key": "test.framework",    "value": "Integration tests use xUnit with Testcontainers for Postgres.", "source": "tool:repo_read" },    { "t": 6,  "layer": "Semantic",   "key": "ticket.format",     "value": "Tickets for the payments repo use the Why / What / Acceptance Criteria layout.", "source": "user" },    { "t": 8,  "layer": "Semantic",   "key": "oncall.rotation",   "value": "Ali is on call for payments the first week of every month.", "source": "user" },    { "t": 9,  "layer": "Episodic",   "key": "ep.incident.dupes", "value": "Incident: duplicate charges on 3 merchants traced to a missing idempotency key on the capture endpoint.", "source": "tool:incident_log" },    { "t": 10, "layer": "Semantic",   "key": "webhook.retries",   "value": "Webhook retries go through an Azure Service Bus queue named pay-retry.", "source": "user" },    { "t": 12, "layer": "Semantic",   "key": "cache.fx",          "value": "Exchange rates are cached in Redis for 15 minutes.", "source": "tool:repo_read" },    { "t": 13, "layer": "Semantic",   "key": "cache.session",     "value": "User sessions are stored in Redis with a 30 minute sliding expiry.", "source": "tool:repo_read" },    { "t": 14, "layer": "Semantic",   "key": "logging",           "value": "Structured logs go to Seq locally and to Grafana Loki in production.", "source": "user" },    { "t": 15, "layer": "Semantic",   "key": "api.versioning",    "value": "Public API versions are passed in the URL path, for example /v2/payments.", "source": "tool:repo_read" },    { "t": 16, "layer": "Semantic",   "key": "money.storage",     "value": "Amounts are stored as integer minor units; banker's rounding is used only for FX conversion.", "source": "user" },    { "t": 18, "layer": "Semantic",   "key": "pr.style",          "value": "The team wants small diffs, never full file rewrites, in agent pull requests.", "source": "user" },    { "t": 20, "layer": "Semantic",   "key": "secrets",           "value": "Secrets are read from Azure Key Vault; nothing sensitive goes in appsettings.json.", "source": "user" },    { "t": 20, "layer": "Semantic",   "key": "promo.summer",      "value": "Temporary promo: merchants onboarding this summer pay no fees for 60 days.", "source": "user", "ttlDays": 30 },    { "t": 22, "layer": "Semantic",   "key": "feature.flags",     "value": "Feature flags are managed in LaunchDarkly under the project payments-core.", "source": "user" },    { "t": 25, "layer": "Semantic",   "key": "refund.limit",      "value": "Refunds above 5000 EUR require a second approver from finance.", "source": "user" },    { "t": 26, "layer": "Semantic",   "key": "db.reporting",      "value": "The reporting replica is a read-only MySQL 8 instance used by the BI team.", "source": "user" },    { "t": 28, "layer": "Semantic",   "key": "data.residency",    "value": "Customer data for EU merchants must stay in the West Europe region.", "source": "user" },    { "t": 30, "layer": "Semantic",   "key": "staging.db",        "value": "Staging was migrated to Postgres 16 on pg-stg-02 and the old MySQL box is retired.", "source": "user" },    { "t": 33, "layer": "Episodic",   "key": "ep.release.4-12",   "value": "Release 4.12 shipped the new payout scheduler; the rollback drill took 6 minutes.", "source": "tool:release_notes" },    { "t": 35, "layer": "Procedural", "key": "rule.migrations",   "value": "Never edit an applied EF Core migration; add a new one instead.", "source": "user" },    { "t": 40, "layer": "Semantic",   "key": "deploy.target",     "value": "Billing now ships straight to the production cluster behind a feature flag.", "source": "user" },    { "t": 45, "layer": "Semantic",   "key": "gateway.timeout",   "value": "PAYGW_TIMEOUT_MS was lowered to 12000 after the latency review.", "source": "user" },    { "t": 47, "layer": "Episodic",   "key": "ep.review.invoices","value": "Design review agreed the orders service should not generate invoice numbers.", "source": "tool:meeting_notes" },    { "t": 50, "layer": "Semantic",   "key": "webhook.retries",   "value": "Failed webhook deliveries now use a Postgres outbox table called webhook_outbox; Service Bus is gone.", "source": "user" },    { "t": 52, "layer": "Procedural", "key": "rule.tests",        "value": "Before proposing a commit in billing-service, run dotnet test plus the contract tests in tests/Contracts.", "source": "user" },    { "t": 55, "layer": "Semantic",   "key": "feature.flags",     "value": "We replaced the old flag vendor with a self-hosted Flagsmith instance at flags.internal.", "source": "user" },    { "t": 56, "layer": "Buffer",     "key": "scratch.branch",    "value": "Working branch for the current task is feature/fx-rounding-fix.", "source": "agent", "ttlDays": 1 },    { "t": 58, "layer": "Buffer",     "key": "scratch.ticket",    "value": "Current ticket being worked on is PAY-881, refund approval UI.", "source": "agent", "ttlDays": 7 }  ],  "questions": [    { "id": "q01", "q": "What database does staging run on?",                              "expect": "postgres",            "stale": "mysql 8 running",   "tags": ["supersession"] },    { "id": "q02", "q": "Where do billing deployments go by default these days?",           "expect": "production",          "stale": "staging cluster",   "tags": ["supersession", "paraphrase"] },    { "id": "q03", "q": "Which setting controls the payment gateway timeout?",              "expect": "paygw_timeout_ms",                                  "tags": ["identifier"] },    { "id": "q04", "q": "What is the current value of PAYGW_TIMEOUT_MS?",                   "expect": "12000",               "stale": "30000",             "tags": ["supersession", "identifier"] },    { "id": "q05", "q": "Who is responsible for generating invoice numbers?",               "expect": "billing service owns",                              "tags": ["paraphrase"] },    { "id": "q06", "q": "How should I structure a new ticket for the payments repo?",       "expect": "acceptance criteria",                               "tags": ["recall"] },    { "id": "q07", "q": "What test tooling do the integration tests rely on?",              "expect": "xunit",                                             "tags": ["recall"] },    { "id": "q08", "q": "How are failed webhook deliveries retried now?",                   "expect": "outbox",              "stale": "service bus queue", "tags": ["supersession"] },    { "id": "q09", "q": "How long do we keep FX rates before refreshing them?",             "expect": "15 minutes",                                        "tags": ["paraphrase"] },    { "id": "q10", "q": "Where can I look at production logs?",                             "expect": "loki",                                              "tags": ["recall"] },    { "id": "q11", "q": "How do clients pick an API version?",                              "expect": "url path",                                          "tags": ["paraphrase"] },    { "id": "q12", "q": "How should money amounts be persisted?",                           "expect": "minor units",                                       "tags": ["paraphrase"] },    { "id": "q13", "q": "What kind of diffs does the team expect from the agent?",          "expect": "small diffs",                                       "tags": ["recall"] },    { "id": "q14", "q": "Where do we keep secrets like connection strings?",                "expect": "key vault",                                         "tags": ["recall"] },    { "id": "q15", "q": "Which feature flag tool are we using?",                            "expect": "flagsmith",           "stale": "launchdarkly",      "tags": ["supersession", "paraphrase"] },    { "id": "q16", "q": "Who needs to sign off on a large refund?",                         "expect": "second approver",                                   "tags": ["paraphrase"] },    { "id": "q17", "q": "Can EU merchant data be stored in the US?",                        "expect": "west europe",                                       "tags": ["recall"] },    { "id": "q18", "q": "What do I need to run before proposing a commit in billing-service?", "expect": "contract tests",    "stale": "attach the summary","tags": ["supersession", "procedural"] },    { "id": "q19", "q": "Is it OK to change a migration that was already applied?",         "expect": "never edit",                                        "tags": ["procedural", "paraphrase"] },    { "id": "q20", "q": "What caused the duplicate charges incident?",                      "expect": "idempotency key",                                   "tags": ["episodic"] },    { "id": "q21", "q": "How long did the 4.12 rollback take?",                             "expect": "6 minutes",                                         "tags": ["episodic", "identifier"] },    { "id": "q22", "q": "Which branch am I working on?",                                    "expect": null,                  "stale": "fx-rounding-fix",   "tags": ["expired"] },    { "id": "q23", "q": "Is the summer no-fee promo still running?",                        "expect": null,                  "stale": "no fees",           "tags": ["expired"] },    { "id": "q24", "q": "Which database does the BI team query?",                           "expect": "reporting replica",                                 "tags": ["distractor"] },    { "id": "q25", "q": "Which ticket am I working on right now?",                          "expect": "pay-881",                                           "tags": ["recall"] },    { "id": "q26", "q": "What did staging run on before Postgres?",                         "expect": "mysql 8 running",                                   "tags": ["history"] }  ]}

It replays the writes on a simulated clock, sweeps TTLs, then asks every question in two modes. “Raw” searches everything that was ever written, which is what my original one-table design did. “Governed” only searches active, unexpired entries. At the end, two contexts edit the same row to check that the concurrency stamp actually rejects the stale write.

using System.Diagnostics;using System.Text.Json;using Microsoft.EntityFrameworkCore;using MemoryLab;// dotnet run                 -> grep and Ollama embeddings// dotnet run -- --no-embed   -> grep only (no Ollama needed)var useEmbeddings = !args.Contains("--no-embed");const string ConnectionString = "Data Source=memorylab.db";var cases = JsonSerializer.Deserialize<Cases>(    File.ReadAllText("cases.json"), new JsonSerializerOptions(JsonSerializerDefaults.Web))!;var day0 = new DateTimeOffset(2026, 8, 1, 9, 0, 0, TimeSpan.Zero);var clock = new SimClock(day0);OllamaEmbedder? embedder = useEmbeddings    ? new OllamaEmbedder(new HttpClient { BaseAddress = new Uri("http://localhost:11434") })    : null;await using (var setup = new MemoryDb(ConnectionString)){    await setup.Database.EnsureDeletedAsync();    await setup.Database.EnsureCreatedAsync();}// 1. Replay sixty "days" of writes. Supersession happens inside WriteAsync.var buildTimer = Stopwatch.StartNew();await using (var db = new MemoryDb(ConnectionString)){    var memory = new MemoryService(db, embedder, clock);    foreach (var w in cases.Writes.OrderBy(w => w.T))    {        clock.Now = day0.AddDays(w.T);        await memory.WriteAsync(w.Key, w.Value, Enum.Parse<MemoryLayer>(w.Layer), w.Source,            w.TtlDays is { } d ? TimeSpan.FromDays(d) : null);    }    clock.Now = day0.AddDays(cases.NowDay);    var expired = await memory.ExpireAsync();    Console.WriteLine($"Built {cases.Writes.Count} writes in {buildTimer.ElapsedMilliseconds} ms, TTL sweep expired {expired}.");}// 2. Ask every question against every retriever, with and without lifecycle filtering.var retrievers = new List<IRetriever> { new GrepRetriever() };if (embedder is not null) retrievers.Add(new EmbeddingRetriever(embedder));Console.WriteLine();Console.WriteLine("+----------------------------+----------+--------+--------+---------+-----------+");Console.WriteLine("| Retriever                  | Mode     | hit@1  | hit@3  | stale@1 | median ms |");Console.WriteLine("+----------------------------+----------+--------+--------+---------+-----------+");var perTag = new List<string>();foreach (var governed in new[] { false, true }){    List<MemoryEntry> candidates;    await using (var db = new MemoryDb(ConnectionString))        candidates = await new MemoryService(db, embedder, clock).CandidatesAsync(governed);    foreach (var retriever in retrievers)    {        int hit1 = 0, hit3 = 0, stale1 = 0, staleN = 0;        var latencies = new List<double>();        var tags = new Dictionary<string, (int ok, int n)>();        foreach (var q in cases.Questions)        {            var sw = Stopwatch.StartNew();            var results = await retriever.SearchAsync(q.Q, candidates, 3);            latencies.Add(sw.Elapsed.TotalMilliseconds);            bool ok1 = Judge(q, results, 1), ok3 = Judge(q, results, 3);            hit1 += ok1 ? 1 : 0;            hit3 += ok3 ? 1 : 0;            if (q.Stale is not null && q.Expect is not null)            {                staleN++;                if (results.Count > 0 && Has(results[0], q.Stale)) stale1++;            }            foreach (var tag in q.Tags)            {                var (ok, n) = tags.GetValueOrDefault(tag);                tags[tag] = (ok + (ok1 ? 1 : 0), n + 1);            }            if (!ok1)                perTag.Add($"  MISS [{retriever.Name}/{Mode(governed)}] {q.Id} -> {(results.Count > 0 ? results[0].Value : "(nothing)")}");        }        latencies.Sort();        var n = cases.Questions.Count;        Console.WriteLine(            $"| {retriever.Name,-26} | {Mode(governed),-8} | {hit1,2}/{n,-3} | {hit3,2}/{n,-3} | {stale1,3}/{staleN,-3} | {latencies[latencies.Count / 2],9:F2} |");        perTag.Add($"[{retriever.Name}/{Mode(governed)}] " +                   string.Join("  ", tags.OrderBy(t => t.Key).Select(t => $"{t.Key} {t.Value.ok}/{t.Value.n}")));    }}Console.WriteLine("+----------------------------+----------+--------+--------+---------+-----------+");Console.WriteLine();perTag.ForEach(Console.WriteLine);// 3. Two agents edit the same memory at once. The second, stale write must lose.await ConcurrencyDemoAsync(ConnectionString);static string Mode(bool governed) => governed ? "governed" : "raw";static bool Has(MemoryEntry m, string needle) =>    m.Value.Contains(needle, StringComparison.OrdinalIgnoreCase);static bool Judge(Question q, IReadOnlyList<MemoryEntry> results, int k){    var top = results.Take(k).ToList();    return q.Expect is null        ? !top.Any(m => Has(m, q.Stale!))          // expired fact must not come back        : top.Any(m => Has(m, q.Expect));}static async Task ConcurrencyDemoAsync(string cs){    await using var agentA = new MemoryDb(cs);    await using var agentB = new MemoryDb(cs);    var a = await agentA.Memories.FirstAsync(m => m.Key == "ticket.format");    var b = await agentB.Memories.FirstAsync(m => m.Key == "ticket.format");    a.Value += " Add a Rollback section.";    await agentA.SaveChangesAsync();    b.Value += " Drop the Why section.";    try    {        await agentB.SaveChangesAsync();        Console.WriteLine("\nConcurrency: agent B silently overwrote agent A. That is the bug.");    }    catch (DbUpdateConcurrencyException)    {        Console.WriteLine("\nConcurrency: agent B's stale write was rejected. Reload and retry.");    }}record Cases(int NowDay, List<Write> Writes, List<Question> Questions);record Write(int T, string Layer, string Key, string Value, string Source, int? TtlDays);record Question(string Id, string Q, string? Expect, string? Stale, List<string> Tags);

Run it with dotnet run, or dotnet run -- --no-embed if Ollama isn't running and you only want the grep half.

Scoring is simple. A question is a hit if the expected phrase appears in the returned memory. For the two expired questions, it’s a hit if the expired fact does not come back. “Stale@1” counts how often, on the six supersession questions, the top result was the old version of the fact.

26 questions, 33 writes (25 active, 6 superseded, 2 expired at day 60)+----------------------------+----------+--------+--------+---------+-----------+| Retriever                  | Mode     | hit@1  | hit@3  | stale@1 | median ms |+----------------------------+----------+--------+--------+---------+-----------+| grep                       | raw      | 19/26  | 24/26  |   2/6   |   ~0.03   || grep                       | governed | 21/26  | 25/26  |   0/6   |   ~0.03   || embed:nomic-embed-text     | raw      | [[MY RUN]] | [[MY RUN]] | [[MY RUN]] | [[MY RUN]] || embed:nomic-embed-text     | governed | [[MY RUN]] | [[MY RUN]] | [[MY RUN]] | [[MY RUN]] |+----------------------------+----------+--------+--------+---------+-----------+Embedding latency includes the HTTP round trip to the local Ollama server, which ismost of it. Grep latency is in-process string matching over 25 to 33 rows.

[[MY RUN: per-tag breakdown for the embedding rows, especially paraphrase and supersession, copied from the MISS lines the runner prints.]]

The grep misses are more interesting than the totals, so here they are in governed mode:

q01  "What database does staging run on?"     -> returned the billing-service commit rule. The new staging fact matches only        "staging" (it never says "database"). The commit rule matches only "run".        One term each, so the recency tie-break picked the newer, wrong one.q02  "Where do billing deployments go by default these days?"     -> same failure. The new deploy fact says "Billing now ships straight to        production", so it matches only "billing", ties with the commit rule        ("billing-service"), and loses on recency again.q05  "Who is responsible for generating invoice numbers?"     -> returned the design review saying the orders service should NOT generate them.        Right topic, opposite answer.q16  "Who needs to sign off on a large refund?"     -> returned the current ticket about a refund approval UI. Both it and the real        answer match "refund"; the ticket is newer. Bonus noise: "sign" matched        "design" in an unrelated episode.q26  "What did staging run on before Postgres?"     -> governed mode can never answer this. The answer is a superseded fact.

Lifecycle mattered more than the retriever. This is the part I’m most confident about, because it’s the one thing the grep rows show on their own. Switching from raw to governed didn’t change the retrieval algorithm at all. It took stale answers on the supersession questions from 2 out of 6 to zero, and expired facts coming back from 2 out of 2 to zero. No ranking trick does that, because neither grep nor cosine similarity has any idea what “still true” means. That was the central argument of the original piece, and it survived.

The cheap baseline is worth running first. Grep with a lifecycle filter got 21 of 26 at the top position in a few hundredths of a millisecond, with no model, no index and no dependency beyond SQLite. Whatever you put on top of it has to beat that by enough to pay for itself.

[[MY RUN: if the governed embedding row also shows stale@1 at or near zero while raw embedding shows stale answers, say so here. That would mean the lifecycle result holds for both retrievers, which is the stronger version of this claim.]]

“Grep beat a graph” was the wrong headline. The number came from a vendor, the agent had a semantic search tool alongside grep, and the benchmark has documented label and judge problems. I’d stand by “start with something simple and familiar.” I wouldn’t stand by the comparison I opened with.

Grep breaks exactly where people say it does. Paraphrase cost it q05 and q16. Worse, supersession only worked when the new fact reused the old fact’s words. When a user says “billing now ships straight to production” instead of “the deploy target is production”, grep has no way to connect the two, and the governance layer can only hide the old fact, not find the new one. My recency tie-break, which I added because it felt like a cheap way to prefer newer facts, actively made q01 worse.

[[MY RUN: whether embeddings fixed q01, q02, q05 and q16, and whether they introduced misses grep didn’t have, for example on the PAYGW_TIMEOUT_MS identifier questions q03 and q04.]]

Active-only memory can’t answer history. q26 is unanswerable in governed mode by construction. I added a HistoryAsync method that walks the supersession chain, but the retriever doesn't know when to call it. That's the temporal weakness again, and it's the category where MAGMA still beats Jev-Mem. Time is not a filter you bolt on at the end. It's structure, and I under-built it twice now.

I conflated memory with state. BlogWriter made this obvious. A lot of what I was calling buffer and procedural memory is really workflow state with a known shape. It doesn’t need retrieval at all, it needs a typed document and a concurrency check.

This is how I’d lay them out now. Anything I didn’t measure myself is labelled with where the claim comes from.

+-------------------------+---------------------+-----------------------+-------------------------------+| Approach                | Cost                | Latency               | How it fails                  |+-------------------------+---------------------+-----------------------+-------------------------------+| Full context            | Tokens grow with    | Grows with history    | Fact buried mid-context,      || (everything in prompt)  | history, every turn |                       | model misses it               |+-------------------------+---------------------+-----------------------+-------------------------------+| Grep over files         | ~zero, no model     | ~0.03 ms in my test   | Paraphrase; new fact in new   || (my lab, raw)           |                     | (measured)            | words; returns stale facts    |+-------------------------+---------------------+-----------------------+-------------------------------+| Grep + lifecycle        | ~zero, plus a sweep | ~0.03 ms in my test   | Paraphrase; can't answer      || (my lab, governed)      | job                 | (measured)            | "what was it before?"         |+-------------------------+---------------------+-----------------------+-------------------------------+| Local embeddings        | Free per call, one  | [[MY RUN]] ms, mostly | Exact identifiers; no idea    || (Ollama + SQLite)       | local model to run  | the Ollama call       | what is still true            |+-------------------------+---------------------+-----------------------+-------------------------------+| LLM fact extraction     | An LLM call on      | Write path slow,      | Wrong fact extracted once,    || (Mem0 style)            | every write         | read path fast        | retrieved forever             |+-------------------------+---------------------+-----------------------+-------------------------------+| Temporal graph          | Entity resolution   | Graph traversal on    | Edge never invalidated;       || (Zep/Graphiti, MAGMA)   | on write            | read (MAGMA 1.47 s,   | entity merged wrongly         ||                         |                     | per Jev-Mem paper)    |                               |+-------------------------+---------------------+-----------------------+-------------------------------+| Jev-Mem                 | Fewer LLM calls,    | 0.93 s query, 158 s   | Not independently reproduced; ||                         | Jev classifier for  | build (authors'       | still behind MAGMA on         ||                         | control             | numbers, unverified)  | temporal                      |+-------------------------+---------------------+-----------------------+-------------------------------+| Markdown knowledge      | Files you own, MCP  | Not measured          | Unknown to me: supersession,  || graph (KnowledgeX)      | server              |                       | provenance, concurrent edits  |+-------------------------+---------------------+-----------------------+-------------------------------+| Session state document  | One read and one    | One point read        | Not memory across sessions;   || (BlogWriter)            | write per run       |                       | stale save rejected by ETag   ||                         | (Cosmos DB)         |                       | (by design)                   |+-------------------------+---------------------+-----------------------+-------------------------------+

The latency numbers in that table aren’t comparable to each other, and I want to say so plainly before someone puts them in a slide. Jev-Mem’s 0.93 seconds includes the answering LLM. My 0.03 milliseconds is a string match over a few dozen rows. Different jobs.

I don’t know how much of the lifecycle win survives at real scale. Thirty-three writes is nothing. With thousands of entries, the question stops being “is the old fact hidden” and becomes “did the supersession ever get recorded in the first place,” which depends on the write path recognising that “billing now ships to production” replaces “deploy target is staging.” My lab cheats by giving both writes the same key. A real agent has to infer that, and inference is exactly where Mem0-style extraction goes wrong.

I don’t know whether Jev-Mem’s numbers hold up outside the authors’ setup. The code is public, so this is testable, and I’d rather someone with a LoCoMo harness and the corrected answer key from the audit ran it than read another summary of the abstract.

I don’t know whether decay helps or just adds a knob. Nitin’s argument is convincing on paper. I’d want a few months of real agent memory before I trust any half-life I pick.

And I don’t know yet whether KnowledgeX handles any of the things above, because I haven’t run it.

Three things, all small.

I’m moving workflow state out of the memory store. Anything with a known shape, like the current task, the ticket, the draft, goes into one typed document per session with a concurrency token, the way BlogWriter does it.

I’m keeping grep plus lifecycle as the default for semantic and procedural memory, but adding the embedding retriever as a fallback when grep returns nothing or ties, rather than as a replacement. [[MY RUN: adjust this sentence if the embedding rows say otherwise.]]

And I’m adding a history path that the agent can call explicitly, because “what was it before?” is a real question people ask, and hiding superseded facts is only half of handling time.

The original article was right that most memory bugs are lifecycle bugs. It was wrong to dress that up as “grep beats graphs.” The lesson I’m keeping from this week is less exciting and more useful: decide what’s true now, keep what used to be true, and only then argue about how to search it.

Tags: AI Agents, Agent Memory, LLM Engineering, DotNet, Entity Framework Core, Ollama, Retrieval Augmented Generation, SQLite

Agent Memory After the Hype: What Jev-Mem, KnowledgeX and BlogWriter Taught Me About My Own Grep… was originally published in Towards AI on Medium, where people are continuing the conversation by highlighting and responding to this story.

── more in #ai-agents 4 stories · sorted by recency
── more on @letta 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/agent-memory-after-t…] indexed:0 read:34min 2026-10-10 · —