Every AI vendor in B2B says it’s the best, but too few show you the data. We’re all used to seeing “evals” of LLMs, but not enough post true detailed AI evals of their agents.
That’s a mistake. Buyers can now test AI agents side by side in an afternoon. Agents are starting to build the shortlists, and in some cases, make the vendor decisions (they do for us at SaaStr AI).
So publish the deepest, most direct competitive evals you can. Run them on live products, against named competitors, and include the categories where you lose.
Will there always be some bias? Yes. You wrote the rubric, picked the weights, and chose what to measure. Say that on the first page. Then make everything checkable.
$100M ARR AI CX leader Gorgias just did this the right way in ecommerce CX. We know the company well — SaaStrFund led the seed round. And in fact, we pushed them to do the most honest, detailed evals in the CX+ space. And they did.
What Gorgias Published #
Gorgias is at ~$100M ARR, and about 80% of that revenue comes from AI support for ecommerce brands. In just launched a public benchmark of its own agent against every major competitor:
- 8,356 live conversations
- 18 vendors
- 212+ live storefronts
- Refreshed weekly, with the same 10-turn conversation run against every vendor
- The full test harness open-sourced on GitHub
The results don’t all go Gorgias’s way, and it published them anyway.
- In support, Gorgias ranks #1, but Yuma resolves more conversations. Automation rate is the metric support buyers care about most, and a competitor leads on it, on Gorgias’s own page.
[SUPPORT: Yuma X% vs. Gorgias Y%; Gorgias wins the composite on quality.] - In pre-sale, Envive beats Gorgias (for now at least): a composite score of 72 to 65. Gorgias has the best answer quality in the lane (76), but it averages 18.4 seconds per answer to Envive’s 7.9. The report’s own latency breakdown shows 28% of Gorgias’s shopping answers taking longer than 20 seconds.
- Gorgias chose the weighting that cost it first place. The pre-sale composite gives speed 25% of the score; the support composite gives it 10%. Score pre-sale with the support weights and Gorgias comes out first at 74.3. It used the tougher weighting because shoppers leave when answers are slow.
Isn’t This Just “Benchmarking”? #
Yes, it’s a benchmark. It just goes much further than what most people picture when they hear the word.
Most B2B buyers know benchmarks as analyst quadrants, feature checklists, G2 grids, or a vendor’s “we’re 3x faster” slide. Those compare what vendors say their products do. They come from surveys, demos, and vendor-supplied answers, and they get updated once or twice a year.
An eval tests what the product actually does. It asks the AI real questions, captures every answer, and grades each one against a written rubric. In Gorgias’s case:
- It tests the live product. Each vendor’s agent is tested on real stores, through the same chat widget a shopper would use.
- Every answer gets graded. Each of the 8,356 conversations is scored on whether it was resolved, whether the answer was right, and how long it took. The results don’t come from a single score someone assigned after a demo.
- Every vendor gets the same test. The same 10-turn conversation runs against every vendor, from a simple first question through cart, shipping, and returns policy.
- It keeps running. The tests run daily and results are published weekly, so if a vendor ships a better model next month, it shows up in next month’s numbers.
- Anyone can rerun it. The code is public, so anyone who doubts the results can check them.
This matters more for AI than it did for traditional software. A CRM does the same thing every time you click the same button. An AI agent can give a great answer on Monday and a wrong one on Thursday, and it can get worse after a model update without anyone at the vendor noticing. A feature checklist can’t show any of that. Only continuous testing of real conversations can.
Why Publishing Your Evals Wins #
Buyers already discount your marketing. Every support leader has seen a dozen vendor charts where the vendor wins, and they ignore all of them. When a vendor shows competitors ahead in some categories, buyers believe the categories where it’s ahead. Gorgias’s #1 in support means more because Yuma’s automation lead is on the same page.
You’re doing the buyer’s diligence for them. No ecommerce brand is going to run 8,356 conversations across 13 vendors. Most sit through three demos and maybe run one pilot. When you run the full comparison and publish all of it, your report becomes what buyers rely on. You also get to decide what it measures, which is a real advantage, and it’s the reason being open about the method matters.
Agents are starting to do the shortlisting. More software evaluations now start with a question to Claude or ChatGPT. An agent can read, check, and cite a versioned rubric, public scoring weights, and open code. It can’t do much with a gated PDF. Vendors that publish structured, checkable evals will get cited more. The live Gorgias report currently blocks automated readers in its robots.txt, though the GitHub repo is open. That’s worth fixing.
Your team sees the gap every week. Inside Gorgias, Envive’s 7.9 seconds and Yuma’s resolution rate aren’t competitive intel buried in a deck anymore. They’re public, next to Gorgias’s own numbers, and updated weekly. The rubric is versioned in the repo, so nobody can quietly change the scoring to make a gap disappear.
How to Do It the Right Way #
The Gorgias approach is a good template:
- Test live products. Gorgias tested real storefronts, not sandboxes or demo environments.
- Name every competitor. An anonymized “Competitor A” tells a buyer nothing.
- Blind the judge. Gorgias strips vendor and store names before its AI judge scores a conversation. A second adversarial review pass has to agree with the first at least 90% of the time.
- Open-source the harness. Competitors who dispute the results can run the code themselves.
- Weight for the end customer. Gorgias weighted speed heavily enough in pre-sale that it lost the top spot.
- Publish the trend alongside the snapshot. A weekly chart is much harder to cherry-pick than a single number.
- State the bias up front. Gorgias’s repo describes the project as “competitive intelligence” and says to treat cross-vendor numbers as directional.
Sources: Gorgias AI Agent Benchmark, gorgias/ai-agent-benchmark on GitHub