{"slug": "building-an-agentic-system-in-net-part-3-durable-memory-with-postgres-and", "title": "Building an Agentic System in .NET, Part 3: Durable Memory with Postgres and pgvector", "summary": "A developer has published the third part of a series on building an agentic system in .NET, showing how to add durable long-term memory using Postgres with the pgvector extension and EF Core. The approach stores distilled facts rather than raw conversation turns, scores each memory item for importance, and creates an HNSW index via raw migration DDL since EF Core lacks a fluent API for it. The writeup includes working entity definitions, package versions, and a recall query intended for production use.", "body_md": "The previous two parts wired up the agent loop and gave it tools. The missing piece is memory. Session context resets with every new conversation, so anything the agent learned about a user, their preferences or an earlier decision is gone. Long term memory fixes that, but only if you store the right things, index them properly, and put them in front of the model at the right moment.\n\nThis part covers all of it with working EF Core code, a real HNSW index, and a recall query you can put into production.\n\nThe most common mistake is storing raw conversation turns. Transcripts grow without limit and are mostly noise: small talk, clarifications, the same question asked three ways. They belong in a separate session store with a time limit, and Redis with a TTL does that job well.\n\nLong term memory should hold *distilled facts*, things that last and are worth pulling back in a later session. Good candidates:\n\nEvery item gets an importance score between 0.0 and 1.0 at write time, plus a timestamp. Both feed the ranking later.\n\nPackages first:\n\n```\ndotnet add package Pgvector.EntityFrameworkCore --version 0.3.0\ndotnet add package Npgsql.EntityFrameworkCore.PostgreSQL --version 10.0.3\n```\n\nWatch the version. `Pgvector.EntityFrameworkCore` v0.3.x targets EF Core 9 and 10. On EF Core 8, pin to v0.2.2.\n\n```\nusing Microsoft.EntityFrameworkCore;\nusing Pgvector;\n\npublic class MemoryItem\n{\n    public Guid Id { get; set; } = Guid.NewGuid();\n    public string UserId { get; set; } = string.Empty;\n    public string Content { get; set; } = string.Empty;\n    public Vector Embedding { get; set; } = null!;\n    public float Importance { get; set; }          // 0.0 to 1.0\n    public DateTimeOffset CreatedAt { get; set; } = DateTimeOffset.UtcNow;\n    public DateTimeOffset LastAccessedAt { get; set; } = DateTimeOffset.UtcNow;\n}\n\npublic class AgentDbContext(DbContextOptions<AgentDbContext> options)\n    : DbContext(options)\n{\n    public DbSet<MemoryItem> Memories => Set<MemoryItem>();\n\n    protected override void OnModelCreating(ModelBuilder modelBuilder)\n    {\n        modelBuilder.HasPostgresExtension(\"vector\");\n\n        modelBuilder.Entity<MemoryItem>(e =>\n        {\n            e.HasKey(m => m.Id);\n            e.Property(m => m.Embedding).HasColumnType(\"vector(1536)\");\n            e.HasIndex(m => m.UserId);\n        });\n    }\n}\n```\n\nRegister the context in `Program.cs`:\n\n```\nbuilder.Services.AddNpgsql<AgentDbContext>(\n    connectionString,\n    npgsqlOptions => npgsqlOptions.UseVector());\n```\n\nCreate the migration with `dotnet ef migrations add AddMemoryItem`. It will produce the table with a `vector(1536)` column, but EF Core has no fluent API for an HNSW index, so you add that DDL to the migration yourself:\n\n```\npublic partial class AddMemoryItem : Migration\n{\n    protected override void Up(MigrationBuilder migrationBuilder)\n    {\n        migrationBuilder.Sql(\"CREATE EXTENSION IF NOT EXISTS vector;\");\n\n        migrationBuilder.CreateTable(\n            name: \"Memories\",\n            columns: table => new\n            {\n                Id            = table.Column<Guid>(nullable: false),\n                UserId        = table.Column<string>(nullable: false),\n                Content       = table.Column<string>(nullable: false),\n                Embedding     = table.Column<Vector>(type: \"vector(1536)\", nullable: false),\n                Importance    = table.Column<float>(nullable: false),\n                CreatedAt     = table.Column<DateTimeOffset>(nullable: false),\n                LastAccessedAt = table.Column<DateTimeOffset>(nullable: false)\n            },\n            constraints: t => t.PrimaryKey(\"PK_Memories\", x => x.Id));\n\n        // EF Core has no fluent API for HNSW, so add it directly.\n        migrationBuilder.Sql(\n            \"\"\"CREATE INDEX ON \"Memories\" USING hnsw (\"Embedding\" vector_cosine_ops)\n               WITH (m = 16, ef_construction = 64);\"\"\");\n\n        migrationBuilder.CreateIndex(\n            name: \"IX_Memories_UserId\",\n            table: \"Memories\",\n            column: \"UserId\");\n    }\n\n    protected override void Down(MigrationBuilder migrationBuilder)\n        => migrationBuilder.DropTable(\"Memories\");\n}\n```\n\n`m = 16` is the default maximum connections per layer and `ef_construction = 64` is the search width at build time. Higher values give better recall but cost build time and memory. For an agent memory store where items arrive one at a time rather than in a bulk load, HNSW is the right index. IVFFlat needs an `ANALYZE` training step after a bulk insert and gets worse if you skip it.\n\n**One gotcha that will stop you.** HNSW in pgvector has a hard ceiling of 2,000 dimensions. `text-embedding-3-small` produces 1,536, which fits. `text-embedding-3-large` produces 3,072 and fails at index creation with `column cannot have more than 2000 dimensions for hnsw index`. Stay on `text-embedding-3-small`, or use the `halfvec` column type if you need the bigger model.\n\nAgent Framework 1.0 (GA on 3 April 2026) and the current Semantic Kernel both sit on `Microsoft.Extensions.AI`, so write against that:\n\n```\nusing Microsoft.Extensions.AI;\n\npublic interface IMemoryEmbeddingService\n{\n    Task<ReadOnlyMemory<float>> EmbedAsync(\n        string text, CancellationToken ct = default);\n}\n\npublic sealed class MeaiEmbeddingService(\n    IEmbeddingGenerator<string, Embedding<float>> generator)\n    : IMemoryEmbeddingService\n{\n    public async Task<ReadOnlyMemory<float>> EmbedAsync(\n        string text, CancellationToken ct = default)\n    {\n        var result = await generator.GenerateAsync(\n            [text], cancellationToken: ct);\n        return result[0].Vector;\n    }\n}\n```\n\nRegister whichever backend you want. The interface does not change:\n\n```\n// Azure OpenAI\nbuilder.Services.AddAzureOpenAIEmbeddingGenerator(\n    deploymentName: \"text-embedding-3-small\",\n    endpoint: new Uri(config[\"AzureOpenAI:Endpoint\"]!),\n    credential: new DefaultAzureCredential());\n\nbuilder.Services.AddSingleton<IMemoryEmbeddingService, MeaiEmbeddingService>();\n```\n\nSwitching to Ollama or another local model is one line of DI. Nothing else in the pipeline notices.\n\nOne long paragraph makes a poor memory item, because the embedding averages the whole thing and the similarity scores get mushy. Chunk first. **SemanticChunker.NET** by Gregor Biswanger works with `Microsoft.Extensions.AI`, and the recommended threshold is percentile 95. Start there and move in steps of five, 90 to 95 to 98, depending on whether you are getting too many tiny fragments or too few useful ones. Leave room in the token budget too. An 8,192 token context has to hold the system prompt, the recalled memories and the new user turn.\n\nAt demo scale, a few hundred rows, a sequential scan is invisible. At 10,000 rows it starts to hurt. At 100,000 rows it is not acceptable anywhere a user is waiting.\n\nHNSW walks a layered graph from coarse to fine, so query cost grows logarithmically instead of linearly. The DBI-services benchmark from March 2026, using 25,000 Wikipedia articles embedded with `text-embedding-3-large`, shows the pattern clearly. On a small table you even have to disable `enable_seqscan` to make the planner use the index at all. On real production data the difference is not subtle.\n\nAt 50 million vectors, pgvectorscale, which is DiskANN based, reaches 471 QPS at 99 percent recall. That is well past the point where plain pgvector is the right tool. Under roughly 10 million vectors, with a mix of relational and vector queries, pgvector with HNSW is the right call.\n\nOne tuning detail catches people out. `hnsw.ef_search` defaults to 40, which caps the candidate list the index returns at query time. Tune it with `SET LOCAL` inside a transaction:\n\n```\nBEGIN;\nSET LOCAL hnsw.ef_search = 100;\nSELECT ... FROM \"Memories\" ORDER BY ... LIMIT 10;\nCOMMIT;\n```\n\nNever use a session level `SET` when you have connection pooling. The setting survives pooler reuse and quietly changes unrelated queries on that connection.\n\nCosine similarity on its own is not enough. A memory from three years ago that scores 0.97 may be less useful than a slightly less similar one from last week that the user marked as critical. The recall score blends three signals:\n\n```\npublic async Task<IReadOnlyList<MemoryItem>> RecallAsync(\n    string userId,\n    string queryText,\n    int topK = 5,\n    CancellationToken ct = default)\n{\n    var queryVector = await _embedding.EmbedAsync(queryText, ct);\n    var pgVector    = new Vector(queryVector.ToArray());\n\n    // Weights: tune these per use-case\n    const float wSimilarity = 0.6f;\n    const float wRecency    = 0.2f;\n    const float wImportance = 0.2f;\n\n    // Recency: exponential decay, half-life = 30 days\n    // EF Core translates CosineDistance to the <=> operator\n    var now = DateTimeOffset.UtcNow;\n\n    var results = await _db.Memories\n        .Where(m => m.UserId == userId)\n        .Select(m => new\n        {\n            Item       = m,\n            Similarity = 1f - m.Embedding.CosineDistance(pgVector),\n            Recency    = (float)Math.Exp(\n                -0.693f * EF.Functions\n                    .DateDiffDay(m.LastAccessedAt, now) / 30.0),\n        })\n        .Select(x => new\n        {\n            x.Item,\n            Score = wSimilarity * x.Similarity\n                  + wRecency    * x.Recency\n                  + wImportance * x.Item.Importance\n        })\n        .OrderByDescending(x => x.Score)\n        .Take(topK)\n        .Select(x => x.Item)\n        .ToListAsync(ct);\n\n    // Update last-accessed timestamp\n    var ids = results.Select(m => m.Id).ToList();\n    await _db.Memories\n        .Where(m => ids.Contains(m.Id))\n        .ExecuteUpdateAsync(s =>\n            s.SetProperty(m => m.LastAccessedAt, now), ct);\n\n    return results;\n}\n```\n\nEF Core translates `CosineDistance` into the `<=>` pgvector operator and the HNSW index picks it up on its own. The recency decay uses a 30 day half life. Halve it or double it depending on how fast facts go stale in your domain.\n\nA recall query nobody calls is useless. Wire it into a `SessionStart` hook, a middleware or agent lifecycle event that runs before the first user message reaches the model:\n\n```\npublic sealed class MemoryInjectionMiddleware(\n    IMemoryRecallService recall,\n    ILogger<MemoryInjectionMiddleware> logger)\n{\n    public async Task<AgentContext> OnSessionStartAsync(\n        AgentContext context, CancellationToken ct = default)\n    {\n        var memories = await recall.RecallAsync(\n            context.UserId,\n            context.InitialMessage,\n            topK: 5, ct);\n\n        if (memories.Count > 0)\n        {\n            var block = string.Join(\"\\n\",\n                memories.Select((m, i) => $\"[Memory {i + 1}] {m.Content}\"));\n\n            context.SystemPrompt = $\"\"\"\n                {context.SystemPrompt}\n\n                ## Recalled context from prior sessions\n                {block}\n                \"\"\";\n\n            logger.LogInformation(\n                \"Injected {Count} memories for user {UserId}\",\n                memories.Count, context.UserId);\n        }\n\n        return context;\n    }\n}\n```\n\nThe agent never sees a database row. It sees its own system prompt with the five most relevant memories attached, ranked by the blended score. That closes the loop: distil facts when you write, rank by similarity plus recency plus importance when you read, and inject before the model sees a single token.\n\nPart 4 takes on the other half of the problem, memory that is stored but no longer true. Timestamps and decay scoring, deduplication, contradiction detection, verification on read, and pruning as a hosted service, so the recall query you just built keeps returning facts the agent can rely on.", "url": "https://wpnews.pro/news/building-an-agentic-system-in-net-part-3-durable-memory-with-postgres-and", "canonical_source": "https://dev.to/mehdimohseni82/building-an-agentic-system-in-net-part-3-durable-memory-with-postgres-and-pgvector-3aaa", "published_at": "2026-10-04 13:08:21+00:00", "updated_at": "2026-10-04 13:12:51.099813+00:00", "lang": "en", "topics": ["ai-agents", "large-language-models", "ai-infrastructure", "developer-tools", "mlops"], "entities": [".NET", "Postgres", "pgvector", "EF Core", "Npgsql", "Redis"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/building-an-agentic-system-in-net-part-3-durable-memory-with-postgres-and", "markdown": "https://wpnews.pro/news/building-an-agentic-system-in-net-part-3-durable-memory-with-postgres-and.md", "text": "https://wpnews.pro/news/building-an-agentic-system-in-net-part-3-durable-memory-with-postgres-and.txt", "jsonld": "https://wpnews.pro/news/building-an-agentic-system-in-net-part-3-durable-memory-with-postgres-and.jsonld"}}