cd /news/artificial-intelligence/lai-146-can-you-trust-your-ais-memor… · home › topics › artificial-intelligence › article
[ARTICLE · art-147648] src=pub.towardsai.net ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

LAI #146: Can You Trust Your AI’s Memory?

Claude Opus 5.5 ranked first overall among 181 AI model configurations tested across 9,050 generated drafts on 10 writing tasks, but GLM-5.3 Flash scored less than 3% lower at roughly 190 times lower cost, according to results published by Towards AI co-founder Louis-François Bouchard. Bouchard also reported that GPT-6 Astra was particularly good at avoiding common AI writing patterns while being less successful at matching his voice, and that the choice of judge significantly affected the rankings. The same issue described a 45-minute Towards AI Mentorship call in which Omar walked through building an AI tutor, advising teams to start with a small evaluation set of real queries and add retrieval only where the simpler setup fails.

read7 min views2 publishedOct 8, 2026

Good morning, AI enthusiasts!

Giving agents more memory is relatively straightforward. Keeping that information accurate, up to date, and under the user’s control is much harder. This week’s issue looks at that challenge, along with several interesting experiments on how we build, evaluate, and understand AI systems. You’ll learn:

We also have my latest experiment comparing 181 AI model configurations across 9,050 writing drafts, including some surprising results on model quality, cost, and evaluation bias.

Let’s get into it.

This week in What’s AI, I’m sharing the results of testing 181 AI model configurations on my own YouTube scripts. I generated 9,050 drafts across 10 writing tasks to see which models could best reproduce my writing style, how much quality improves with higher spending, and whether AI judges can reliably tell the difference.

Claude Opus 5.5 ranked first overall, but GLM-5.3 Flash scored less than 3% lower at roughly 190 times lower cost. GPT-6 Astra was particularly good at avoiding common AI writing patterns, although it was less successful at matching my voice. I also found that the choice of judge can significantly affect the rankings.

I share six findings from these experiments, including the trade-offs between writing quality and cost, when generating multiple drafts makes sense, and how to build a similar evaluation for your own writing. Read the full article here.

This week, Omar spent 45 minutes on one of our Mentorship live calls, walking through how we built our AI tutor. It is a useful example if you are deciding whether your own system needs RAG and where to start.

As he mentioned on the call, start with the questions the system needs external knowledge to answer. Collect a small set of real queries and, for each one, record:

Start by checking whether the system can actually surface the information needed for those questions. If it cannot, retrieval is the first problem to solve. If it can, there is no point changing the retriever yet.

With our tutor, that meant testing real student questions against the retrieval setup before adding more infrastructure. We only added complexity where the simpler setup was actually failing.

So if you are deciding whether to build RAG, start with a small evaluation set and find the exact point where your current system loses access to the information it needs. The architecture should follow that failure, not come before it.

The 45-minute tutor walkthrough was part of our weekly Towards AI Mentorship live call, where we work through these kinds of implementation decisions using both our systems and what members are building.

A quick update on the book before I wrap up. We are deep into the final rounds of AI Engineering for Production now: editing chapters, tightening examples, checking the technical details, and making sure the book reflects how we actually build with these systems today.

Over the next few weeks, we will start sharing more from behind the scenes: parts of the book, ideas that changed substantially while we were writing, lessons from building the examples, and some material that will not make it into the final manuscript. If you want to follow the book as it comes together and get those updates first, you can sign up here*.*

And if there is something you particularly want us to cover, reply to one of those emails and tell me. We are far enough along that the structure is set, but there is still room to make sure we answer the questions practitioners actually have.

— Louis-François Bouchard, Towards AI Co-founder & Head of Community

Agent_hellboy built Cully, an intelligent workspace that observes your coding workflow, remembers what matters, and helps you steer Claude Code, Codex, Cursor, or other coding agents across sessions. It also has a live Agent Health bar that shows the agent, project, branch, session time, context pressure, loop warnings, and unchecked edits; every number is measured, and anything unknown is shown as missing rather than guessed. Check it out here and support a fellow community member. If you have any questions or feedback, share them in the thread!

The Learn AI Together Discord community is flooding with collaboration opportunities. If you are excited to dive into applied AI, want a study partner, or even want to find a partner for your passion project, join the collaboration channel! Keep an eye on this section, too — we share cool opportunities every week!

  1. Inc0gnit08125 is looking for an accountability partner, so everyone can share what they have achieved. The main goal is to speed up learning and become job-ready in the next 6 months. If you are working towards the same thing, connect with them in the thread!

2. .d0n0x is looking for groups/people to learn fine-tuning, post-training, theory, and other fundamentals for experimentation, and also work on practical workflows/agents. If this sounds interesting to you, reach out to them in the thread!

  1. Divyanshijoshi is participating in the AI Builder Cup Competition and looking for 1 teammate with experience in full-stack/GCP dev (React, Cloud Run, Firebase). If you want to join, contact her in the thread!

Meme shared by drdub_

From Token to Trajectory by Lian Kim The same token starts with the same embedding, but its hidden states change as the model processes context. The author traces the word “light” through six weight-related and color-related contexts using Pythia-160M. The representations begin separating at state 3, but the distinction weakens at the final state under cosine similarity, even as raw Euclidean distance increases because of larger vector norms. A second experiment compares JEV, Claude, and GPT on banking-intent classification, measuring both decision accuracy and whether they return valid labels. The article shows why representation metrics need careful interpretation and why classification systems should be evaluated on more than accuracy alone.

  1. Fractal Recursion, Bounded Infinity, and Adapted Systems by Lian Kim

A model can reach lower training loss without improving its ability to generalize or adapt. The author uses the Koch snowflake, whose perimeter grows indefinitely while its area remains finite, to introduce the difference between convergence and stability. The author uses gradient descent and Hessian eigenvalues to examine how models converge during training, including why a flatter loss minimum does not necessarily mean better generalization. It also explains the difference between a system that stays close to its original state after a disturbance and one that eventually returns to it, and connects these ideas to memory and adaptive systems.

  1. The Memory Problem OpenAI’s Dots Makes Impossible To Ignore by allglenn

OpenAI’s dots privacy FAQ states that users cannot individually inspect, edit, or delete saved memories. The author uses this limitation to examine how persistent agents should handle information that becomes outdated, incorrect, or needs to be forgotten. He compares memory controls across dots, Meta’s Muse, xAI’s Team Bots, and Claude Managed Agents, then explains why similarity-based retrieval alone cannot reliably distinguish old facts from current ones. The article proposes separate memory write and retrieval paths, validity periods, and provenance tracking so developers can correct or delete specific information without resetting an agent’s entire memory.

  1. AI Agents That Don’t Wait for a Prompt: How Microsoft Foundry Routines Work by Dave R

Microsoft Foundry Routines lets agents start automatically on schedules, at specified times, or in response to GitHub issues and Teams messages. The author explains how Foundry manages triggers, execution identities, queues, retries, and run histories, including how agents can schedule their own follow-up tasks using a reminder tool. He also covers the difference between running an agent under its own identity and the routine creator’s permissions.

  1. Your Knowledge Graph Is Measuring Your Pipeline, Not Your Subject: A Replication Attempt by Ali Süleyman TOPUZ

Knowledge graphs can tell us how concepts relate to one another, but their structure also depends on how the LLM extracts those relationships. The author tested the same extraction pipeline on three unrelated subjects: ASP.NET Core middleware, SQL optimization, and cellulose acetate manufacturing. Surprisingly, the resulting graphs had almost the same number of connections relative to their size. Even small changes to how relationships were classified affected some metrics more than changing the subject entirely. The article explains how to check whether your knowledge graph metrics reflect the source material or simply the way your extraction pipeline was configured.

If you are interested in publishing with Towards AI, check our guidelines and sign up. We will publish your work to our network if it meets our editorial policies and standards. LAI #146: Can You Trust Your AI’s Memory? was originally published in Towards AI on Medium, where people are continuing the conversation by highlighting and responding to this story.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @claude opus 5.5 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/lai-146-can-you-trus…] indexed:0 read:7min 2026-10-08 · —