cd /news/ai-agents/the-weekly-reading · home › topics › ai-agents › article
[ARTICLE · art-146711] src=amitkvint.com ↗ pub= topic=ai-agents verified=true sentiment=· neutral

The Weekly Reading

Amit Kvint kept an AI agent handling roughly 4,000 support reports a month honest through a weekly one-hour review in which he and an experienced supporter read the agent's actual customer replies and asked whether they would have sent them. The three recurring failure modes were citing the wrong documentation, being too confident on an unknown bug, and using the wrong tone, and every finding fed back into documentation, classification, or tuning. Kvint argues the loop must stay recurring because whole-queue metrics show when something is wrong but never what is wrong.

read5 min views1 publishedOct 7, 2026
The Weekly Reading
Image: Amitkvint (auto-discovered)

Want to talk about this essay? Email me: amitkvint@gmail.comcopied· or message me on LinkedIn

The AI agent we put on a queue of around 4,000 reports a month had a dashboard, an escalation path and a classification step in front of it. None of that is what kept it honest.

What kept it honest was an hour a week where two people sat down and read what it had actually told customers.

I have mentioned this loop in passing before, in the piece about scaling support without hiring, and I promised it its own piece when I wrote about what should never be automated. Here it is.

Who was reading #

Me, because I owned the quality of the whole thing and the call on whether the agent could take more of the queue. And an experienced supporter, because you need someone who would catch what you would miss.

That pairing was not an accident, and I would not do it with one person.

I was reading as the owner. I wanted to know whether the agent was good enough to widen, and when you want something to be good enough, you read a little generously. The supporter was reading as someone who had answered thousands of these reports by hand, and knew what a correct answer looked like, what a technically correct but useless answer looked like, and what a confident answer to a problem nobody has solved yet looked like. Those three read very differently when you have lived them, and almost identically when you have not.

What we were reading for #

Not for whether the conversation had closed. The dashboard already knew that, and I have written elsewhere about why a closed ticket proves very little.

We were reading the replies themselves, as a customer would have read them, and asking a simpler question: would I have sent this?

There were three general ways the AI failed: it cited the wrong documentation, it was too confident on an unknown bug, it used the wrong tone.

The useful finds were never the obvious failures. A reply that was plainly wrong gets escalated by the customer and shows up in the numbers on its own. The ones worth the hour were the replies that were plausible. The right document linked for the wrong version of the problem. A reassuring answer to a report that, to a human who had seen it before, was clearly the first sign of a bug we did not have a fix for yet. A reply that was accurate and arrived with a tone the situation did not deserve.

None of those bounce back cleanly. The customer tries the steps, they do not work, and a week later a different ticket with a different subject line arrives. The agent's scorecard, if it had one, would have shown a success.

What we did with it #

Every finding went back into the system. Some were documentation gaps: the agent was answering from what it had, and what it had was thin in that corner. Some were classification problems, reports that had been marked AI Ready and should not have been, which is how the line between agent and human moved over the months, in both directions. Some were tuning.

The point is not which bucket. The point is that the agent only got better in the places where somebody had read a reply and said, out loud, that it was not good enough. Nothing else told us. The whole-queue numbers, which is what we used to decide whether the agent could take more, told us when something was wrong. They never told us what.

Why it has to be a recurring hour #

It is tempting to treat this as a launch activity. Read a lot in the first weeks, fix what you find, and then let the numbers take over.

We did read most in the first weeks, and it was the most intense period. But the loop never came off, and if you stop reading you make a huge mistake.

Two reasons.

The reports change. A plugin ships a new version, a hosting company changes something, a new kind of site becomes common, and a category of question that the agent handled well in March starts drifting in September. The numbers will show the drift eventually. The reading shows it the week it starts.

And the people reading change. The supporter who sits with you learns, from the inside, what the agent is good at and where it bluffs. That knowledge goes back to the team in a way no report does. When an escalation lands on their desk with the agent's history attached, they already know which parts of that history to trust.

The part I keep coming back to #

Every discussion I hear about running AI in support is about the model, the prompts, the integrations and the metrics. All of those matter, and all of them can be bought.

The hour cannot. It needs someone senior enough to know what a good answer is, with enough authority to say the agent is not ready for more, and with a boss who treats that as the most important meeting of the week rather than a nice-to-have.

If you are about to run an agent on your own queue, I would put that hour in the calendar before the agent goes live, and I would be very suspicious of anyone, including a vendor, who tells you that after the first month you will not need it anymore. You will. The queue has not stopped changing. Neither has the agent, and somebody has to be reading. :)

── more in #ai-agents 4 stories · sorted by recency
── more on @amit kvint 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/the-weekly-reading] indexed:0 read:5min 2026-10-07 · —