LAI #146: Can You Trust Your AI’s Memory? Claude Opus 5.5 ranked first overall among 181 AI model configurations tested across 9,050 generated drafts on 10 writing tasks, but GLM-5.3 Flash scored less than 3% lower at roughly 190 times lower cost, according to results published by Towards AI co-founder Louis-François Bouchard. Bouchard also reported that GPT-6 Astra was particularly good at avoiding common AI writing patterns while being less successful at matching his voice, and that the choice of judge significantly affected the rankings. The same issue described a 45-minute Towards AI Mentorship call in which Omar walked through building an AI tutor, advising teams to start with a small evaluation set of real queries and add retrieval only where the simpler setup fails. Good morning, AI enthusiasts Giving agents more memory is relatively straightforward. Keeping that information accurate, up to date, and under the user’s control is much harder. This week’s issue looks at that challenge, along with several interesting experiments on how we build, evaluate, and understand AI systems. You’ll learn: We also have my latest experiment comparing 181 AI model configurations across 9,050 writing drafts, including some surprising results on model quality, cost, and evaluation bias. Let’s get into it. This week in What’s AI, I’m sharing the results of testing 181 AI model configurations on my own YouTube scripts. I generated 9,050 drafts across 10 writing tasks to see which models could best reproduce my writing style, how much quality improves with higher spending, and whether AI judges can reliably tell the difference. Claude Opus 5.5 ranked first overall, but GLM-5.3 Flash scored less than 3% lower at roughly 190 times lower cost. GPT-6 Astra was particularly good at avoiding common AI writing patterns, although it was less successful at matching my voice. I also found that the choice of judge can significantly affect the rankings. I share six findings from these experiments, including the trade-offs between writing quality and cost, when generating multiple drafts makes sense, and how to build a similar evaluation for your own writing. Read the full article here https://louisbouchard.substack.com/p/i-tested-181-ai-models-on-my-own . This week, Omar spent 45 minutes on one of our Mentorship https://towardsai.com/academy/mentorship/?utm source=newsletter&utm medium=email&utm id=AItips live calls, walking through how we built our AI tutor. It is a useful example if you are deciding whether your own system needs RAG and where to start. As he mentioned on the call, start with the questions the system needs external knowledge to answer. Collect a small set of real queries and, for each one, record: Start by checking whether the system can actually surface the information needed for those questions. If it cannot, retrieval is the first problem to solve. If it can, there is no point changing the retriever yet. With our tutor, that meant testing real student questions against the retrieval setup before adding more infrastructure. We only added complexity where the simpler setup was actually failing. So if you are deciding whether to build RAG, start with a small evaluation set and find the exact point where your current system loses access to the information it needs. The architecture should follow that failure, not come before it. The 45-minute tutor walkthrough was part of our weekly Towards AI Mentorship https://towardsai.com/academy/mentorship/?utm source=newsletter&utm medium=email&utm id=AItips live call, where we work through these kinds of implementation decisions using both our systems and what members are building. A quick update on the book before I wrap up. We are deep into the final rounds of AI Engineering for Production now: editing chapters, tightening examples, checking the technical details, and making sure the book reflects how we actually build with these systems today. Over the next few weeks, we will start sharing more from behind the scenes: parts of the book, ideas that changed substantially while we were writing, lessons from building the examples, and some material that will not make it into the final manuscript. If you want to follow the book as it comes together and get those updates first, you can sign up here . And if there is something you particularly want us to cover, reply to one of those emails and tell me. We are far enough along that the structure is set, but there is still room to make sure we answer the questions practitioners actually have. — Louis-François Bouchard, Towards AI Co-founder & Head of Community Agent hellboy https://discord.com/channels/702624558536065165/983037843532308500/1557526037108105306 built Cully, an intelligent workspace that observes your coding workflow, remembers what matters, and helps you steer Claude Code, Codex, Cursor, or other coding agents across sessions. It also has a live Agent Health bar that shows the agent, project, branch, session time, context pressure, loop warnings, and unchecked edits; every number is measured, and anything unknown is shown as missing rather than guessed. Check it out here https://github.com/mcp-runtime/cully and support a fellow community member. If you have any questions or feedback, share them in the thread https://discord.com/channels/702624558536065165/983037843532308500/1557526037108105306 The Learn AI Together Discord community is flooding with collaboration opportunities. If you are excited to dive into applied AI, want a study partner, or even want to find a partner for your passion project, join the collaboration channel https://discord.gg/rj6m9AF7eC Keep an eye on this section, too — we share cool opportunities every week 1. Inc0gnit08125 https://discord.com/channels/702624558536065165/784477688551178240/1555205787909881978 is looking for an accountability partner, so everyone can share what they have achieved. The main goal is to speed up learning and become job-ready in the next 6 months. If you are working towards the same thing, connect with them in the thread https://discord.com/channels/702624558536065165/784477688551178240/1555205787909881978 2. .d0n0x https://discord.com/channels/702624558536065165/784477688551178240/1557707157565349962 is looking for groups/people to learn fine-tuning, post-training, theory, and other fundamentals for experimentation, and also work on practical workflows/agents. If this sounds interesting to you, reach out to them in the thread https://discord.com/channels/702624558536065165/784477688551178240/1557707157565349962 3. Divyanshijoshi https://discord.com/channels/702624558536065165/998978160605540454/1555491383626698873 is participating in the AI Builder Cup Competition and looking for 1 teammate with experience in full-stack/GCP dev React, Cloud Run, Firebase . If you want to join, contact her in the thread https://discord.com/channels/702624558536065165/998978160605540454/1555491383626698873 Meme shared by drdub https://discord.com/channels/702624558536065165/830572933197201459/1556140816328687646 From Token to Trajectory https://pub.towardsai.net/from-token-to-trajectory-8300f311f6a2?sharedUserId=tai-tech by Lian Kim https://liankimco.medium.com/ The same token starts with the same embedding, but its hidden states change as the model processes context. The author traces the word “light” through six weight-related and color-related contexts using Pythia-160M. The representations begin separating at state 3, but the distinction weakens at the final state under cosine similarity, even as raw Euclidean distance increases because of larger vector norms. A second experiment compares JEV, Claude, and GPT on banking-intent classification, measuring both decision accuracy and whether they return valid labels. The article shows why representation metrics need careful interpretation and why classification systems should be evaluated on more than accuracy alone. 1. Fractal Recursion, Bounded Infinity, and Adapted Systems https://pub.towardsai.net/fractal-recursion-bounded-infinity-and-adapted-systems-b11046c9e1e1?sharedUserId=tai-tech by Lian Kim https://liankimco.medium.com/ A model can reach lower training loss without improving its ability to generalize or adapt. The author uses the Koch snowflake, whose perimeter grows indefinitely while its area remains finite, to introduce the difference between convergence and stability. The author uses gradient descent and Hessian eigenvalues to examine how models converge during training, including why a flatter loss minimum does not necessarily mean better generalization. It also explains the difference between a system that stays close to its original state after a disturbance and one that eventually returns to it, and connects these ideas to memory and adaptive systems. 2. The Memory Problem OpenAI’s Dots Makes Impossible To Ignore https://pub.towardsai.net/the-memory-problem-openais-dots-makes-impossible-to-ignore-525e93363d64?sk=2e301ff5621d915ca195a3ff1658bf17 by allglenn https://medium.com/@glennlenormand OpenAI’s dots privacy FAQ states that users cannot individually inspect, edit, or delete saved memories. The author uses this limitation to examine how persistent agents should handle information that becomes outdated, incorrect, or needs to be forgotten. He compares memory controls across dots, Meta’s Muse, xAI’s Team Bots, and Claude Managed Agents, then explains why similarity-based retrieval alone cannot reliably distinguish old facts from current ones. The article proposes separate memory write and retrieval paths, validity periods, and provenance tracking so developers can correct or delete specific information without resetting an agent’s entire memory. 3. AI Agents That Don’t Wait for a Prompt: How Microsoft Foundry Routines Work https://pub.towardsai.net/microsoft-foundry-routines-ai-agents-schedule-events-15e9a4b0c25e?sk=6964401bad5b927db7f340cb95d27f4f by Dave R https://blog.azinsider.net/ Microsoft Foundry Routines lets agents start automatically on schedules, at specified times, or in response to GitHub issues and Teams messages. The author explains how Foundry manages triggers, execution identities, queues, retries, and run histories, including how agents can schedule their own follow-up tasks using a reminder tool. He also covers the difference between running an agent under its own identity and the routine creator’s permissions. 4. Your Knowledge Graph Is Measuring Your Pipeline, Not Your Subject: A Replication Attempt https://pub.towardsai.net/your-knowledge-graph-is-measuring-your-pipeline-not-your-subject-a-replication-attempt-6b62462820c6?sharedUserId=tai-tech by Ali Süleyman TOPUZ https://topuzas.medium.com/ Knowledge graphs can tell us how concepts relate to one another, but their structure also depends on how the LLM extracts those relationships. The author tested the same extraction pipeline on three unrelated subjects: ASP.NET Core middleware, SQL optimization, and cellulose acetate manufacturing. Surprisingly, the resulting graphs had almost the same number of connections relative to their size. Even small changes to how relationships were classified affected some metrics more than changing the subject entirely. The article explains how to check whether your knowledge graph metrics reflect the source material or simply the way your extraction pipeline was configured. If you are interested in publishing with Towards AI, check our guidelines and sign up https://contribute.towardsai.net/ . We will publish your work to our network if it meets our editorial policies and standards. LAI 146: Can You Trust Your AI’s Memory? https://pub.towardsai.net/lai-146-can-you-trust-your-ais-memory-f30f47fd9bd1 was originally published in Towards AI https://pub.towardsai.net on Medium, where people are continuing the conversation by highlighting and responding to this story.