Exactly how many support and related issues can AI Agents really, truly resolve today? It’s a good test of just exactly where the latest models and agents are.
Now $100M ARR AI CX for ecommerce leader Gorgias (where SaaStr Fund led the seed) has now published a public benchmark that tests 13 vendors on 212 live, mid-market ecommerce stores carrying real inventory, with no sandboxes, synthetic catalogs, or vendor demos. Every agent gets the same customer messages, and a judge scores each response blind to the vendor against 26 binary checks. Factual claims like price, policy, and SKU are checked programmatically against the live store.
The TL;DR: the best AI agents fully resolve about 70% of conversations. The typical vendor resolves under half.
The Top Five Resolve 64% to 75%. The Median Vendor Resolves 48%. #
Automation rate here is the share of engaged conversations the AI resolved with no human involved. Each number pools the last four weeks, with conversation counts in parentheses:
- **Rep AI:** 75% (120)
- **Decagon:** 73% (34)
- **Yuma:** 71% (118)
- **Ada:** 68% (420)
- **Gorgias:** 64% (388)
- **DigitalGenius:** 50% (561)
- **Sierra:** 48% (408)
- **Siena:** 46% (300)
- **Kodif:** 44% (188)
- **Intercom:** 42% (307)
- **Zendesk:** 37% (467)
- **Envive:** 22% (499)
- **Klaviyo:** 21% (309)
The top five average 70%, the median is 48%, and the field resolves 46% weighted by conversation volume.
Zendesk, Intercom, and Klaviyo are the platforms most brands already pay for, and none of them resolves more than 42% in this benchmark. On this data, the AI bundled with your existing helpdesk or email tool is one of the weakest options available.
Yuma, Decagon, and Gorgias Are the Only Vendors Above 64% Automation and 65 Quality #
Automation rate alone can point you to the wrong vendor. Here is each vendor’s rate next to its blind quality score (0 to 100):
Strong on both:
- Yuma: 71% automation, 72 quality
- Decagon: 73% automation, 70 quality
- Gorgias: 64% automation, 66 quality
High automation, weak answers:
-
Rep AI: 75% automation, 55 quality
-
Ada: 68% automation, 39 quality Strong answers, low automation:
-
Sierra: 48% automation, 72 quality
-
Intercom: 42% automation, 66 quality
-
Klaviyo: 21% automation, 65 quality
Ada closes more conversations than Gorgias but scores 27 points lower on the answers. Sierra ties Yuma for the best quality score and resolves under half of its conversations.
The benchmark’s case for weighting quality is practical: a ticket marked resolved that gives the customer the wrong return window creates a second contact that costs more than the first. A vendor scoring 39 on quality is generating some of its own future ticket volume.
9 of 10 Vendors Lost Automation in the Last Four Weeks #
Ten vendors have enough history to show a four-week trend. On automation, nine declined:
- Siena: -23 points
- Kodif: -13
- Klaviyo: -11
- Sierra: -11
- DigitalGenius: -10
- Ada: -9
- Gorgias: -6
- Intercom: -6
- Zendesk: -4
Envive was the only one to improve, at +8. Quality moved the same way, with nine of ten down: Ada -17, Zendesk -15, Sierra -14, Gorgias -12. Siena was the only vendor to improve on quality, at +2.
The benchmark doesn’t explain the declines. Model updates, store configuration drift, or a harder question set could all contribute. Whatever the cause, the rate you saw in the pilot is not the rate you’ll have next month. We run 20+ agents in production at SaaStr, and we re-check every one of them on a schedule, including the ones that seem to be working.
“Resolved” Here Means Zero Human Touch #
Part of the gap between homepage claims and this data comes from the definition. The benchmark counts a conversation as automated only when the AI handled it with zero human touch, no handover to a person, and no deflection out of the channel. “Email us,” a contact form, or “call us” all count against the agent. If more than half of an agent’s replies push the customer out of the channel, the conversation is scored as unresolved. The auditor never asks for a human, so every handoff is the agent’s own decision.
In-product analytics usually use looser rules. In Gorgias’s own customer-facing reporting, an interaction counts as automated once the AI resolves it and 72 hours pass without a human agent helping. A customer who gave up and never came back counts as a resolution there. In the benchmark, that customer counts as a failure.
This matters for your P&L because “resolved” is also the pricing unit. Gorgias AI Agent charges $0.90 per resolved conversation. Get each vendor’s definition of a billable resolution in writing before you sign.
Order Tracking Scores 17 Points Below Returns Policy #
The report breaks support quality down by topic, averaged across the field:
- Returns policy: 77.4
- Shipping policy: 74.4
- Modify or cancel: 63.5
- Order tracking: 60.1
- Damaged item: 59.7
- Shopper unsure: 49.3
The field drops 28 points from its easiest conversation type to its hardest.
Questions answered from a policy page score highest. Questions that need a live system lookup or a judgment call score lowest. Order tracking is usually the largest ticket category for an ecommerce brand, and it sits near the bottom of this list. Across the field, only about a third of post-sale conversations reach a finish.
Map your actual ticket mix onto that list before you choose a vendor. If most of your volume is “where’s my order,” the benchmark average overstates what you’ll get.
The Most Common Failures Were Auth Walls and Clarification Loops #
The benchmark found that the most common failures were authentication walls, repeated demands for an order number, clarification loops, and handing everything to a human. Judges rarely caught agents making up facts. They frequently caught agents that wouldn’t engage.
The same vendor can score well on one store and near zero on another, and the benchmark attributes the difference to onboarding and quality control more than to the model in the demo. Nearly a third of the “AI chat” widgets it found produced no real conversation at all.
Guardrails, identity checks, and escalation rules account for much of the gap between a 45% deployment and a 70% one, and the brand configures all three.
Yuma Is the Slowest Vendor and Resolves 71% #
Mean time to a complete reply runs from 5.4 seconds (Envive) to 16.5 seconds (Yuma). Envive is the fastest vendor on the board and resolves 22%. Yuma is the slowest and resolves 71%, with a top quality score.
For post-purchase support, speed matters much less than resolution. The benchmark weights the support composite 50% automation, 40% quality, and 10% speed, on the basis that customers tolerate moderate latency when the answer is accurate. For pre-sale shopping, speed gets 25% of the weight. The latency numbers also have a caveat. The benchmark only averages turns it could time cleanly. A vendor’s slowest turns are the ones most likely to break or time out and drop out of the average, which can make that vendor look faster than it is.
Plan for 70% on a Top Vendor and 50% on the Rest #
Build the business case on 65% to 70%, and only on a top-five vendor. The best tools in their best month still leave about 30% of conversations for a human or a second agent with deeper system access. On the AI bundled with Zendesk, Intercom, or Klaviyo, plan for 20% to 42%.
Shortlist on automation and quality together. Only three vendors cleared both bars in this window. Ask every vendor for both numbers on the same set of conversations.
Weight the results by your own ticket mix. Policy questions resolve well across the field. Order tracking and damaged items don’t.
Run a cold test on the vendor’s reference stores. Take 30 real tickets from last month, open fresh incognito sessions on their live customers’ sites, and count full resolutions yourself. The rubric is public, so you can grade your current agent the same way.
Re-test every month. Nine of ten vendors lost ground on automation in the last four weeks.
Account for who built the benchmark. Gorgias runs it, and Gorgias ranks fifth on automation on its own board, for now at least. The rubric, methodology, and transcripts are open to inspect, and anyone can add a store. None of the other twelve vendors has published anything comparable.