cd /news/large-language-models/i-tried-building-llm-memory-the-stup… · home topics large-language-models article
[ARTICLE · art-90496] src=dev.to ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

I Tried Building LLM Memory the Stupid Way — It Worked Surprisingly Well

A developer experimenting with local RAG systems found that a simple line-length filter—keeping only lines over 10 words—can effectively preserve conversational memory for LLMs. In tests with 13-14 specific questions about past chats, the system recovered almost all details, including exact amounts, dates, and cross-conversation connections, despite retrieving only eight chunks. The approach failed on short code snippets, but the developer considers it a useful context builder rather than a production-grade memory system.

read2 min views1 publishedAug 10, 2026

This started as a fun experiment.

I've been working on a local RAG system , and at some point I started wondering about something that sounds simple but is actually quite difficult:

Can an LLM remember old conversations without sending its entire history back into the context window?

I didn't want to build a complicated memory architecture just to answer that question. So I tried something almost ridiculous. I exported my old chats and extracted conversations from Json files.

The entire memory extractor was basically one rule

If a line had more than 10 words, keep it.

If it had 10 or fewer, throw it away.

That's pretty much it.

No LLM summarization during ingestion. No expensive memory model. No elaborate memory classification.

The resulting text was then chunked and fed into my existing RAG. For every question, I allowed retrieval of only 8 chunks.

So the experiment looked roughly like this:

Old conversations

Keep lines with >10 words

Create chunks

BM25 + semantic retrieval

Top 8 chunks

Feed back to same LLM

Then I decided to actually try to break it.

I used my own old conversations

Instead of asking generic questions, I went back to conversations from the extracted chats and asked the LLM very specific questions.

Things like old procurement discussions, financial figures, dates, coding decisions, project discussions and relationships between completely separate conversations.

I asked roughly 13–14 questions.

LLM failed.

Then I fed the chunks from my RAG

And honestly, I expected it to fail much more often.

It didn't.

It recovered almost everything I asked about.

It could retrieve exact amounts and dates. It could recover old technical discussions. In some cases, it could connect information from separate conversations and reconstruct what had happened.

And remember:

The model only received eight retrieved chunks.

It wasn't given my entire historical conversation corpus.

Then I found the obvious weakness

The one clear failure was actually quite revealing.

It was terrible at remembering code and commands may be since they are often less than 10 words. So my brilliant memory algorithm simply throws them away. But I tried to make it as a context builder rather than a storage. So I am satisfied with results.

The LLM was restricted to use internet or think only to answer from its memory. I think if that was allowed that single miss would haven’t occurred. Its not a production grade memory management system. But it works gr8.

── more in #large-language-models 4 stories · sorted by recency
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/i-tried-building-llm…] indexed:0 read:2min 2026-08-10 ·