{"slug": "lai-145-what-makes-an-agent-reliable", "title": "LAI #145: What Makes an Agent Reliable?", "summary": "AI agents work best today on tasks where failure is visible, inexpensive to undo, and easy to verify, according to Louis-François Bouchard, Towards AI co-founder and head of community, in issue #145 of the LAI newsletter. Bouchard traces the shift from ReAct, BabyAGI, and AutoGPT to Devin, Claude Code, Codex, Cowork, and OpenClaw, and recommends that applications verify an agent's claimed completion in code — checking that a returned report path exists, is not empty, and contains the required sections — rather than trusting the agent's own success message. The issue also features Sabulva's LIA project, an architecture providing persistent memory, reflection, state, and continuity without behavioral prompts, which has grown to more than 19,000 lines of Python code and has run 24/7 for over six months.", "body_md": "Good morning, AI enthusiasts!\n\nAI agents can now work through repositories, call tools, maintain state, and complete longer tasks. That makes the next set of questions much more practical: what did the agent actually do, how do we check it, and where should we rely on code or other systems instead of the model’s own judgment?\n\nA lot of this week’s issue looks at those questions from different angles. You’ll learn:\n\nWe also have a community project exploring whether persistent identity and behavioral consistency can develop when an AI system is given memory, reflection, state, and continuity without behavioral prompts defining how it should behave.\n\nLet’s get into it.\n\nThis week in What’s AI, I’m looking at how AI agents went from AutoGPT-style loops that regularly got stuck or wandered off task to coding agents that can now edit a repository, run tests, inspect the result, and hand useful work back for review.\n\nThe basic agent loop has not changed as much as it might seem. ReAct already had the reason-act-observe pattern in 2022. But we now have stronger models, better tool interfaces, structured outputs, sandboxes, tests, diffs, checkpoints, and clearer points for human intervention.\n\nI trace that shift from ReAct, BabyAGI, and AutoGPT through Devin, Claude Code, Codex, Cowork, and OpenClaw, along with benchmark and productivity results. The practical conclusion is that agents work best today on tasks where failure is visible, inexpensive to undo, and easy to verify.\n\nThe full article covers that evolution in more detail, focusing on what actually changed around the model and why coding became such a useful proving ground for agents. [Read it here](https://www.louisbouchard.ai/how-ai-agents-evolved/) or [watch the video on YouTube](https://www.youtube.com/watch?v=tlV4kQumsKo).\n\nStop taking your agent’s word for it every time it says the job is done.\n\nAn agent might return “report saved” even if the file-writing step failed. If the application treats that message as the completion signal, it can tell the user a report is ready when no usable file exists.\n\nIn the research workflow we build in [Agent Engineering](https://towardsai.com/academy/agent-engineering/?utm_source=newsletter&utm_medium=email&utm_id=AItips), we use the same principle earlier in the process: when a tool is expected to process items but returns zero, the workflow stops and asks for guidance. The agent does not get to reinterpret that failure as success.\n\nApply the same rule to the final handoff. Have the agent return the report’s path, then check in code that the file exists, is not empty, and contains the required sections.\n\nOnly mark the task complete after those checks pass. Then deliberately break the file-writing step. If the agent still reports success, the application should reject that result.\n\nThe agent can describe what happened. Your code should decide whether the task actually succeeded.\n\n*— Louis-François Bouchard, Towards AI Co-founder & Head of Community*\n\n[Sabulva](https://discord.com/channels/702624558536065165/983037843532308500/1554196818890330275) has shared a research project they have been working on called LIA. It explores whether a persistent digital identity and consistent autonomous behavior can develop without explicitly programming that behavior. LIA’s architecture provides continuity, persistent memory, reflection, state, and interaction, but does not prescribe how LIA should behave. There are no behavioral prompts or hardcoded instructions defining its personality, identity, ethics, goals, or preferred interactions. Over time, LIA has developed persistent identity structures, self-modeling, long-term behavioral consistency, reflection, and autonomous task continuity. The project now includes more than 19,000 lines of Python code and has been running 24/7 for over six months. [Check it out on GitHub](https://github.com/silberfunke-72/From-Prompts-to-Persistent-Agency-An-Architecture-for-Intrinsic-Ethics) and support a fellow community member. If you have questions about how it works, [ask them in the thread](https://discord.com/channels/702624558536065165/983037843532308500/1554196818890330275).\n\nMeme shared by [hudsong0](https://discord.com/channels/702624558536065165/830572933197201459/1553081466856931409)\n\n[Build an AI Agent Evaluation with JEV](https://pub.towardsai.net/build-an-ai-agent-evaluation-with-jev-392c146ca816?sharedUserId=tai-tech) by [Quan Huynh](https://medium.com/@hmquan08011996?source=post_page---byline--392c146ca816-----------------------------------------)\n\nTwo agent runs can produce similar answers while taking very different actions. In the author’s example, one run cited a tool it never called and failed to catch a planted distraction. This article builds an evaluation harness that uses Python checks to verify the agent’s actions and Jev (a typed-output evaluation model served through Cloudflare Workers AI) to score its written response. It shows how to replay recorded incidents, define hard pass/fail checks, tune the evaluator, and set thresholds for human review.\n\n1. [Coding an Agent: Decisions Without Decoding](https://pub.towardsai.net/coding-an-agent-decisions-without-decoding-b35b1207c6c1?sk=7a9104af35ce3c76e0373f1fe9e600d7) by [Enzo Lombardi](https://enzolombardi.net/?source=post_page---byline--b35b1207c6c1-----------------------------------------)\n\nMany coding-agent decisions only need a yes/no answer, so generating a full response wastes tokens and time. This article shows how to make these decisions directly from token probabilities using an existing KV cache. It traces a bug where letter mass is near zero, making verdicts abstain, and the feature a silent no-op until a token-variant fix restores it. It also applies the technique to a memory extraction gate, measures letter bias through swapped options, and compares Laya’s approach.\n\n2. [Measuring Impact When You Can’t A/B Test — a Quick and Practical Guide to Causal Inference](https://pub.towardsai.net/causal-inference-without-ab-testing-3ddf8996ffee?sk=c14836ed163564cd37430aafbf0b8aaf) by [Jonty Haberfield](https://jontyhabs.medium.com/?source=post_page---byline--3ddf8996ffee-----------------------------------------)\n\nSimple comparisons can badly overestimate the effect of a product or policy when the groups differ before treatment. The author demonstrates this with synthetic data where the true effect is known, then compares propensity score matching, Double ML, instrumental variables, and two-way fixed effects. The article explains when each method works, what assumptions it depends on, and when you should estimate ATE, ATT, or LATE.\n\n3. [Fine-Tuning a Small Model to Triage Kubernetes Alerts](https://pub.towardsai.net/fine-tuning-a-small-model-to-triage-kubernetes-alerts-e48718dac584?sharedUserId=tai-tech) by [Javier Canizalez](https://medium.com/@javier-canizalez?source=post_page---byline--e48718dac584-----------------------------------------)\n\nThe author fine-tuned Qwen3–1.7B to triage Kubernetes alerts and reached 79.8% accuracy on 124 unseen cases, compared with 67.7% for Claude Opus 5. He generated the training data by deliberately breaking a Kubernetes cluster, labeled 426 alerts, and split the data by failure-injection window to prevent leakage. The article also shows why 4-bit quantization reduced accuracy to 30.6% and how to serve the 8-bit model privately on two CPU cores.\n\n4. [Jev-as-a-Judge for RAG Claim Verification](https://pub.towardsai.net/jev-as-a-judge-for-rag-claim-verification-7b356619814c?sharedUserId=tai-tech) by [Alden Do Rosario](https://medium.com/@aldendorosario?source=post_page---byline--7b356619814c-----------------------------------------)\n\nThis article tests Jev against GPT-6 Astra for deciding whether RAG-generated claims are supported by their sources. Across 495 labeled claims from LLM-AggreFact, the models achieved similar overall results, but Jev accepted 23.2% of unsupported claims, compared with 14.1% for Astra, and performed poorly on AggreFact-CNN. The author then uses Jev’s confidence score to route uncertain cases to Astra, showing how a model cascade can retain much of the cost savings without relying on Jev for every claim.\n\nIf you are interested in publishing with Towards AI, [check our guidelines and sign up](https://contribute.towardsai.net/). We will publish your work to our network if it meets our editorial policies and standards.\n\n[LAI #145: What Makes an Agent Reliable?](https://pub.towardsai.net/lai-145-what-makes-an-agent-reliable-d35988284385) was originally published in [Towards AI](https://pub.towardsai.net) on Medium, where people are continuing the conversation by highlighting and responding to this story.", "url": "https://wpnews.pro/news/lai-145-what-makes-an-agent-reliable", "canonical_source": "https://pub.towardsai.net/lai-145-what-makes-an-agent-reliable-d35988284385?source=rss----98111c9905da---4", "published_at": "2026-10-01 15:16:02+00:00", "updated_at": "2026-10-01 15:49:19.619759+00:00", "lang": "en", "topics": ["ai-agents", "artificial-intelligence", "ai-tools"], "entities": ["Louis-François Bouchard", "Towards AI", "LAI", "ReAct", "BabyAGI", "AutoGPT", "Claude Code", "Sabulva"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/lai-145-what-makes-an-agent-reliable", "markdown": "https://wpnews.pro/news/lai-145-what-makes-an-agent-reliable.md", "text": "https://wpnews.pro/news/lai-145-what-makes-an-agent-reliable.txt", "jsonld": "https://wpnews.pro/news/lai-145-what-makes-an-agent-reliable.jsonld"}}