cd /news/ai-agents/building-an-agentic-system-in-net-pa… · home › topics › ai-agents › article
[ARTICLE · art-144838] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Building an Agentic System in .NET, Part 3: Durable Memory with Postgres and pgvector

A developer has published the third part of a series on building an agentic system in .NET, showing how to add durable long-term memory using Postgres with the pgvector extension and EF Core. The approach stores distilled facts rather than raw conversation turns, scores each memory item for importance, and creates an HNSW index via raw migration DDL since EF Core lacks a fluent API for it. The writeup includes working entity definitions, package versions, and a recall query intended for production use.

by read7 min views2 publishedOct 4, 2026

The previous two parts wired up the agent loop and gave it tools. The missing piece is memory. Session context resets with every new conversation, so anything the agent learned about a user, their preferences or an earlier decision is gone. Long term memory fixes that, but only if you store the right things, index them properly, and put them in front of the model at the right moment.

This part covers all of it with working EF Core code, a real HNSW index, and a recall query you can put into production.

The most common mistake is storing raw conversation turns. Transcripts grow without limit and are mostly noise: small talk, clarifications, the same question asked three ways. They belong in a separate session store with a time limit, and Redis with a TTL does that job well.

Long term memory should hold distilled facts, things that last and are worth pulling back in a later session. Good candidates:

Every item gets an importance score between 0.0 and 1.0 at write time, plus a timestamp. Both feed the ranking later.

Packages first:

dotnet add package Pgvector.EntityFrameworkCore --version 0.3.0
dotnet add package Npgsql.EntityFrameworkCore.PostgreSQL --version 10.0.3

Watch the version. Pgvector.EntityFrameworkCore v0.3.x targets EF Core 9 and 10. On EF Core 8, pin to v0.2.2.

using Microsoft.EntityFrameworkCore;
using Pgvector;

public class MemoryItem
{
    public Guid Id { get; set; } = Guid.NewGuid();
    public string UserId { get; set; } = string.Empty;
    public string Content { get; set; } = string.Empty;
    public Vector Embedding { get; set; } = null!;
    public float Importance { get; set; }          // 0.0 to 1.0
    public DateTimeOffset CreatedAt { get; set; } = DateTimeOffset.UtcNow;
    public DateTimeOffset LastAccessedAt { get; set; } = DateTimeOffset.UtcNow;
}

public class AgentDbContext(DbContextOptions<AgentDbContext> options)
    : DbContext(options)
{
    public DbSet<MemoryItem> Memories => Set<MemoryItem>();

    protected override void OnModelCreating(ModelBuilder modelBuilder)
    {
        modelBuilder.HasPostgresExtension("vector");

        modelBuilder.Entity<MemoryItem>(e =>
        {
            e.HasKey(m => m.Id);
            e.Property(m => m.Embedding).HasColumnType("vector(1536)");
            e.HasIndex(m => m.UserId);
        });
    }
}

Register the context in Program.cs:

builder.Services.AddNpgsql<AgentDbContext>(
    connectionString,
    npgsqlOptions => npgsqlOptions.UseVector());

Create the migration with dotnet ef migrations add AddMemoryItem. It will produce the table with a vector(1536) column, but EF Core has no fluent API for an HNSW index, so you add that DDL to the migration yourself:

public partial class AddMemoryItem : Migration
{
    protected override void Up(MigrationBuilder migrationBuilder)
    {
        migrationBuilder.Sql("CREATE EXTENSION IF NOT EXISTS vector;");

        migrationBuilder.CreateTable(
            name: "Memories",
            columns: table => new
            {
                Id            = table.Column<Guid>(nullable: false),
                UserId        = table.Column<string>(nullable: false),
                Content       = table.Column<string>(nullable: false),
                Embedding     = table.Column<Vector>(type: "vector(1536)", nullable: false),
                Importance    = table.Column<float>(nullable: false),
                CreatedAt     = table.Column<DateTimeOffset>(nullable: false),
                LastAccessedAt = table.Column<DateTimeOffset>(nullable: false)
            },
            constraints: t => t.PrimaryKey("PK_Memories", x => x.Id));

        // EF Core has no fluent API for HNSW, so add it directly.
        migrationBuilder.Sql(
            """CREATE INDEX ON "Memories" USING hnsw ("Embedding" vector_cosine_ops)
               WITH (m = 16, ef_construction = 64);""");

        migrationBuilder.CreateIndex(
            name: "IX_Memories_UserId",
            table: "Memories",
            column: "UserId");
    }

    protected override void Down(MigrationBuilder migrationBuilder)
        => migrationBuilder.DropTable("Memories");
}

m = 16 is the default maximum connections per layer and ef_construction = 64 is the search width at build time. Higher values give better recall but cost build time and memory. For an agent memory store where items arrive one at a time rather than in a bulk load, HNSW is the right index. IVFFlat needs an ANALYZE training step after a bulk insert and gets worse if you skip it.

One gotcha that will stop you. HNSW in pgvector has a hard ceiling of 2,000 dimensions. text-embedding-3-small produces 1,536, which fits. text-embedding-3-large produces 3,072 and fails at index creation with column cannot have more than 2000 dimensions for hnsw index. Stay on text-embedding-3-small, or use the halfvec column type if you need the bigger model.

Agent Framework 1.0 (GA on 3 April 2026) and the current Semantic Kernel both sit on Microsoft.Extensions.AI, so write against that:

using Microsoft.Extensions.AI;

public interface IMemoryEmbeddingService
{
    Task<ReadOnlyMemory<float>> EmbedAsync(
        string text, CancellationToken ct = default);
}

public sealed class MeaiEmbeddingService(
    IEmbeddingGenerator<string, Embedding<float>> generator)
    : IMemoryEmbeddingService
{
    public async Task<ReadOnlyMemory<float>> EmbedAsync(
        string text, CancellationToken ct = default)
    {
        var result = await generator.GenerateAsync(
            [text], cancellationToken: ct);
        return result[0].Vector;
    }
}

Register whichever backend you want. The interface does not change:

// Azure OpenAI
builder.Services.AddAzureOpenAIEmbeddingGenerator(
    deploymentName: "text-embedding-3-small",
    endpoint: new Uri(config["AzureOpenAI:Endpoint"]!),
    credential: new DefaultAzureCredential());

builder.Services.AddSingleton<IMemoryEmbeddingService, MeaiEmbeddingService>();

Switching to Ollama or another local model is one line of DI. Nothing else in the pipeline notices.

One long paragraph makes a poor memory item, because the embedding averages the whole thing and the similarity scores get mushy. Chunk first. SemanticChunker.NET by Gregor Biswanger works with Microsoft.Extensions.AI, and the recommended threshold is percentile 95. Start there and move in steps of five, 90 to 95 to 98, depending on whether you are getting too many tiny fragments or too few useful ones. Leave room in the token budget too. An 8,192 token context has to hold the system prompt, the recalled memories and the new user turn.

At demo scale, a few hundred rows, a sequential scan is invisible. At 10,000 rows it starts to hurt. At 100,000 rows it is not acceptable anywhere a user is waiting.

HNSW walks a layered graph from coarse to fine, so query cost grows logarithmically instead of linearly. The DBI-services benchmark from March 2026, using 25,000 Wikipedia articles embedded with text-embedding-3-large, shows the pattern clearly. On a small table you even have to disable enable_seqscan to make the planner use the index at all. On real production data the difference is not subtle.

At 50 million vectors, pgvectorscale, which is DiskANN based, reaches 471 QPS at 99 percent recall. That is well past the point where plain pgvector is the right tool. Under roughly 10 million vectors, with a mix of relational and vector queries, pgvector with HNSW is the right call.

One tuning detail catches people out. hnsw.ef_search defaults to 40, which caps the candidate list the index returns at query time. Tune it with SET LOCAL inside a transaction:

BEGIN;
SET LOCAL hnsw.ef_search = 100;
SELECT ... FROM "Memories" ORDER BY ... LIMIT 10;
COMMIT;

Never use a session level SET when you have connection pooling. The setting survives pooler reuse and quietly changes unrelated queries on that connection.

Cosine similarity on its own is not enough. A memory from three years ago that scores 0.97 may be less useful than a slightly less similar one from last week that the user marked as critical. The recall score blends three signals:

public async Task<IReadOnlyList<MemoryItem>> RecallAsync(
    string userId,
    string queryText,
    int topK = 5,
    CancellationToken ct = default)
{
    var queryVector = await _embedding.EmbedAsync(queryText, ct);
    var pgVector    = new Vector(queryVector.ToArray());

    // Weights: tune these per use-case
    const float wSimilarity = 0.6f;
    const float wRecency    = 0.2f;
    const float wImportance = 0.2f;

    // Recency: exponential decay, half-life = 30 days
    // EF Core translates CosineDistance to the <=> operator
    var now = DateTimeOffset.UtcNow;

    var results = await _db.Memories
        .Where(m => m.UserId == userId)
        .Select(m => new
        {
            Item       = m,
            Similarity = 1f - m.Embedding.CosineDistance(pgVector),
            Recency    = (float)Math.Exp(
                -0.693f * EF.Functions
                    .DateDiffDay(m.LastAccessedAt, now) / 30.0),
        })
        .Select(x => new
        {
            x.Item,
            Score = wSimilarity * x.Similarity
                  + wRecency    * x.Recency
                  + wImportance * x.Item.Importance
        })
        .OrderByDescending(x => x.Score)
        .Take(topK)
        .Select(x => x.Item)
        .ToListAsync(ct);

    // Update last-accessed timestamp
    var ids = results.Select(m => m.Id).ToList();
    await _db.Memories
        .Where(m => ids.Contains(m.Id))
        .ExecuteUpdateAsync(s =>
            s.SetProperty(m => m.LastAccessedAt, now), ct);

    return results;
}

EF Core translates CosineDistance into the <=> pgvector operator and the HNSW index picks it up on its own. The recency decay uses a 30 day half life. Halve it or double it depending on how fast facts go stale in your domain.

A recall query nobody calls is useless. Wire it into a SessionStart hook, a middleware or agent lifecycle event that runs before the first user message reaches the model:

public sealed class MemoryInjectionMiddleware(
    IMemoryRecallService recall,
    ILogger<MemoryInjectionMiddleware> logger)
{
    public async Task<AgentContext> OnSessionStartAsync(
        AgentContext context, CancellationToken ct = default)
    {
        var memories = await recall.RecallAsync(
            context.UserId,
            context.InitialMessage,
            topK: 5, ct);

        if (memories.Count > 0)
        {
            var block = string.Join("\n",
                memories.Select((m, i) => $"[Memory {i + 1}] {m.Content}"));

            context.SystemPrompt = $"""
                {context.SystemPrompt}

                ## Recalled context from prior sessions
                {block}
                """;

            logger.LogInformation(
                "Injected {Count} memories for user {UserId}",
                memories.Count, context.UserId);
        }

        return context;
    }
}

The agent never sees a database row. It sees its own system prompt with the five most relevant memories attached, ranked by the blended score. That closes the loop: distil facts when you write, rank by similarity plus recency plus importance when you read, and inject before the model sees a single token.

Part 4 takes on the other half of the problem, memory that is stored but no longer true. Timestamps and decay scoring, deduplication, contradiction detection, verification on read, and pruning as a hosted service, so the recall query you just built keeps returning facts the agent can rely on.

── more in #ai-agents 4 stories · sorted by recency
── more on @.net 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/building-an-agentic-…] indexed:0 read:7min 2026-10-04 · —