cd /news/ai-agents/ai-agent-memory-vs-rag-what-s-the-di… · home topics ai-agents article
[ARTICLE · art-127014] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

AI agent memory vs RAG — what's the difference?

A developer behind Statewave argues that RAG and agent memory are distinct patterns that are frequently conflated, noting that RAG retrieves content an agent doesn't already know while memory retrieves context the agent has already participated in. The writeup details three failure modes that emerge when teams stretch RAG to cover memory, including embedding-nearest results that aren't decision-relevant, no compaction of history into durable facts, and no invalidation model for superseded facts. Statewave is positioned as a memory layer built on pgvector that handles ranking, provenance, and token-bounded context bundles, intended to run alongside existing RAG stacks rather than replace them.

by read4 min views1 publishedSep 11, 2026

Most teams building on LLMs end up with two patterns in the same codebase: RAG for looking things up in a corpus, and some hand-rolled memory for remembering what the agent has done or what the user has said. The two are often confused, and the confusion costs real engineering time when one is used in place of the other.

RAG retrieves content the agent doesn't already know. Memory retrieves context the agent has already participated in.

They share a vector store but they answer different questions, store different shapes of data, and have different correctness requirements.

Both patterns sit on top of an embedding store — usually pgvector or a dedicated vector database — and both use approximate-nearest-neighbour search at retrieval time. That's where the overlap ends.

If you only need to ground answers in static documentation, you need RAG. If you need the agent to remember the user across sessions, you need memory. RAG was designed to answer "what does our documentation say about X?" — not "what did this user tell us last month?" When teams stretch RAG to cover memory, three failure modes show up:

You can paper over each of these in your application code. Teams that do end up with a memory runtime in everything but name.

A memory runtime adds three things on top of the vector store:

None of those are properties of "RAG" in the literature sense. They're what makes memory infrastructure rather than retrieval over chat logs.

RAG Memory
Question shape "What does our content say about X?" "What does this user / agent / project need to know right now?"
Data shape Document chunks Episodes → compiled memories with provenance
Ranking signal Cosine similarity Similarity + kind priority + recency + validity + token budget
Mutability Append-only chunks; reindex on doc update Episodes append-only; memories supersede; compaction is idempotent
Output Top-K chunks Token-bounded bundle ready to drop into a prompt

Most production agents need both. The grounding corpus (docs, knowledge base) lives in RAG. The user / account / project context lives in memory.

Trying to make either pattern do the other's job is the common architecture mistake — and it's the one we built Statewave to stop people from making.

Statewave is the memory layer — episodes in, ranked context bundles out, with deterministic ranking and provenance. It uses pgvector under the hood (no separate vector DB to operate) but it's not a RAG framework: there's no document , no chunker, no retriever interface for grounding-over-corpora. If you want RAG, plug Statewave alongside your existing RAG stack — Statewave handles the who you're talking to layer, your RAG framework handles the what does the knowledge base say layer.

The architecture page on the docs site goes deeper on the ranking signals and the compile-vs-retrieve split. The getting-started guide is a five-minute Docker Compose path if you want to try it side-by-side with your current RAG setup.

RAG retrieves content the agent doesn't already know — document chunks ranked by cosine similarity. Memory retrieves context the agent has already participated in — episodes and compiled facts ranked by recency, kind, validity, and similarity.

You can stretch it, but three failure modes show up: embedding-nearest isn't decision-relevant (an allergy note won't be the closest embedding to a lunch question), there's no compaction of history into durable facts, and there's no invalidation model for facts that get superseded.

Most production agents do. The grounding corpus — docs, knowledge base — lives in RAG. The user, account, or project context lives in memory. Trying to make either pattern do the other's job is the common architecture mistake.

Three things: compilation (turning raw episodes into typed facts with confidence and validity), deterministic ranking (the same query always returns the same bundle), and provenance (every compiled memory carries the IDs of the episodes it came from).

No. It uses pgvector under the hood but ships no document , chunker, or retriever for grounding over a corpus. It's the who-you're-talking-to layer, meant to run alongside your existing RAG stack rather than replace it.

Look at the shape of the question. "What does our content say about X?" is RAG. "What does this user, agent, or project need to know right now?" is memory.

Originally published on the Statewave blog. Statewave is an open-source, self-hosted memory runtime for AI agents — GitHub.

── more in #ai-agents 4 stories · sorted by recency
── more on @statewave 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ai-agent-memory-vs-r…] indexed:0 read:4min 2026-09-11 ·