{"slug": "the-agent-was-wrong-all-tuesday", "title": "The Agent Was Wrong All Tuesday", "summary": "An AI support agent handling roughly 4,000 reports a month kept giving a confidently wrong answer to every customer for an entire Tuesday because it had no way to recognize that a new bug was new, according to a first-person essay by Amit Kvint. Kvint's team fixed the problem not with a model change but with an operational override that pushes new information into the system immediately, measured in minutes rather than update cycles, plus a human whose job is to notice the world changed and overrule the agent before lunch. Kvint argues the lasting damage was not the wrong answers themselves but the loss of customer and support-staff trust, which does not return to neutral once the underlying bug is fixed.", "body_md": "# The Agent Was Wrong All Tuesday\n\nWant to talk about this essay? Email me: amitkvint@gmail.comcopied· [or message me on LinkedIn](https://www.linkedin.com/in/amitkvint/)\n\nEvery AI support demo runs in a world the system already knows.\n\nThe documentation is up to date. The bugs are ones someone has already reported. The questions are ones customers have already asked.\n\nIn that world, an agent can look excellent. And, to be fair, it usually is.\n\nProduction is different.\n\nProduction is Tuesday morning.\n\n## What Tuesday looks like\n\nA new bug goes out with a release. Something changes underneath you. Or two products that have never interacted before start behaving badly together on a customer's site.\n\nWhatever the cause, at some point on Tuesday morning there is a real problem that nobody has seen before.\n\nThe first customer writes in.\n\nThe report looks completely ordinary. There is nothing about the way someone describes a brand-new bug that tells you it is new. They describe what they are seeing. The symptoms look similar to something already documented. The system picks up the existing answer and goes with it.\n\nSo the agent answers.\n\nQuickly. Politely. Confidently.\n\nAnd incorrectly.\n\nThen it does the same thing for the next customer, and the one after that.\n\nThat consistency is one of the reasons to have an agent handling a queue of around 4,000 reports a month. It doesn't get tired, and it doesn't forget how to do something because it has had a difficult morning.\n\nOn Tuesday, though, that becomes the problem.\n\nA human support team has an early-warning system that nobody really designed. Someone on the second tier says, \"That's the third one of these today. Something's going on.\"\n\nThe agent doesn't have a third ticket.\n\nEvery ticket is the first ticket.\n\n## The fix was a door, not a model\n\nA handful of decisions mattered more to that rollout than any prompt we wrote. I've written about some of them already, but this one is important.\n\nWe built a way to get new information into the system immediately, instead of waiting for the normal update cycle.\n\nThat's it.\n\nIt wasn't particularly clever. It was the operational equivalent of putting a note on someone's desk:\n\n**As of this morning, this isn't the old problem. It's a new one. Here's what you need to know.**\n\nThe important part wasn't building the mechanism. Any competent team can do that.\n\nThe important part was deciding beforehand that the gap between \"a human knows\" and \"the system knows\" was something we needed to measure.\n\nAnd that we wanted it measured in minutes, not in update cycles.\n\nBecause the alternative isn't that the agent goes quiet until the documentation catches up.\n\nThe alternative is that it keeps giving the old answer. The one that was correct until this morning. All day. To everyone.\n\n## Someone has to notice\n\nA door is no use if nobody walks through it.\n\nAnd that part isn't a technology decision at all.\n\nThe override only works if someone is reading what the agent is actually telling customers, and if that person is allowed to overrule it without booking a meeting first. In our case that came out of a weekly habit of going through the agent's answers, which is a subject of its own and one I'll come back to.\n\nI think this is where \"human in the loop\" usually gets sold short.\n\nIt isn't a person approving replies one at a time. That doesn't scale, and it throws away most of what you just bought.\n\nIt's a person whose job is to notice that the world changed this morning, and who has a way to tell the system before lunch.\n\n## What it costs when you're slow\n\nThe early problems were mostly fixable from a technical point of view. A wrong answer is a bug, and bugs get fixed.\n\nWhat took much longer to repair was how customers and supporters felt about the agent afterwards.\n\nA customer who was told something wrong, confidently, doesn't go back to neutral when you fix the underlying issue. Neither does a supporter who spent Tuesday afternoon apologising for something they didn't write.\n\nThat's the real reason the override mattered, and it doesn't show up in any of the numbers.\n\nAfter the full rollout our average resolution time went from roughly 24 hours to around 10, and satisfaction stayed above 95%.\n\nThose numbers describe ordinary weeks. They survive the other kind of week only if the other kind of week gets caught on the morning it starts.\n\n## The part I keep coming back to\n\nA demo tests the model.\n\nProduction tests the distance between the moment a human knows something and the moment the system does.\n\nSo if you're evaluating an agent for your own queue, I'd ask about that distance before I asked about accuracy. How does new information get in? Who is allowed to put it there? How long does it take?\n\nAnd what does the thing say in the meantime, while it doesn't know?\n\n\"It'll be in the next update\" is a perfectly good answer for documentation.\n\nIt's a bad answer for Tuesday.", "url": "https://wpnews.pro/news/the-agent-was-wrong-all-tuesday", "canonical_source": "https://amitkvint.com/writing/the-agent-was-wrong-all-tuesday/", "published_at": "2026-09-11 00:00:00+00:00", "updated_at": "2026-09-28 05:16:30.621819+00:00", "lang": "en", "topics": ["ai-agents", "ai-safety", "ai-products"], "entities": ["Amit Kvint"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/the-agent-was-wrong-all-tuesday", "markdown": "https://wpnews.pro/news/the-agent-was-wrong-all-tuesday.md", "text": "https://wpnews.pro/news/the-agent-was-wrong-all-tuesday.txt", "jsonld": "https://wpnews.pro/news/the-agent-was-wrong-all-tuesday.jsonld"}}