When we rolled out an AI support agent to a queue of around 5,000 tickets a month, we never built a deflection dashboard. Not because we had some strong opinion against it. We just couldn’t see what it would actually tell us.
Deflection measures the conversations that never reach a human. It doesn’t tell you whether the customer’s problem was solved. Someone who gives up is not a successful deflection. They might be a refund next week, or another ticket with a completely different subject line.
This August, I noticed three very different parts of the industry making essentially the same point. Automattic’s CX lead argued that deflection is not a support strategy. A vendor-side column in Smart Customer Service argued that resolution rate should replace it. And China’s consumer-protection coverage quoted a legal expert saying AI support should be judged on resolution and escalation efficiency, with a regulator behind the idea.
So, problem solved? Not quite.
The problem is that a resolution rate can be gamed too.
Close-on-silence is one common example: if a customer doesn’t reply within 72 hours, the ticket is marked as resolved. Then there are bot-confirmed resolutions, where the agent asks “did that help?” and treats the customer disappearing as a yes. And if reopened tickets are logged as new tickets, one failure can conveniently turn into two “successful” resolutions.
Basically, any metric that only looks at the bot’s part of the queue can be made to look good by carefully choosing what counts.
For me, the answer is to not give the AI its own scoreboard at all. You look at one queue and one set of numbers. AI-handled and human-handled tickets should be measured together, using average resolution time and customer satisfaction across the whole queue. If the AI said it had solved something when it hadn’t, the customer would come back, and that would show up in our numbers. There should be no separate AI success metric to hide behind.
That also gives you a much safer way to scale it. You start the agent on 10% of the queue and only increase its share when the overall numbers hold up. Six months later, it'll be covering the entire queue. In our case average resolution time had dropped from around 24 hours to around 10, while satisfaction stayed above 95%.
And those numbers were for everyone, not just the tickets where the AI happened to perform well.
I understand why deflection dashboards are attractive. They’re easy to set up, easy to explain, and, conveniently, they usually go up.
But I think the metric you use to judge a support team should be something the customer would recognize as a success too.
Nobody ever wrote to us to say, “Thanks for deflecting me.” :)